A score is a proxy. Improving the proxy without improving the experience is possible and common, so measure enquiries alongside it.

What a score is

A single number summarising several measurements, weighted by somebody else's judgement about what matters.

That is useful as a rough indicator and it is not the thing you care about.

The thing you care about is whether visitors get what they came for quickly enough to stay and act.

Those two usually move together, and they can be separated, which is what this is about.

How a score improves without the experience

The second is the common one. A tool that delays scripts until a visitor interacts scores well, because the measurement finishes before the work starts, and the visitor still waits for it the moment they touch the page.

The test is what changed for a visitor

Which is a different question from what changed in the number.

Ask what a visitor now gets sooner than they did.

If the answer is that the text appears earlier, the page becomes usable sooner, or the menu responds faster, the improvement is real.

If the answer is that a measurement now runs before some work rather than after it, nothing has improved for anybody.

That question can be asked of any proposed change, and it filters most of the fake improvements without any technical knowledge.

Somebody who cannot answer it in plain terms has probably improved the score rather than the site.

A worked example

A business paid for performance work and their score rose substantially.

Enquiries over the following quarter were unchanged, as was the bounce rate.

Examining what had been done: images had been deferred including the one at the top, scripts had been set to load only on interaction, and a plugin had been installed that produced a better measurement.

The effect for a visitor was that the main image now appeared later than before and the menu took a moment to respond to the first tap.

The page scored well and was slightly worse to use.

Undoing two of those changes lowered the score and improved the actual experience, which was the correct trade.

Measure the things that are not proxies

Since the answer to a bad proxy is a better measurement rather than no measurement.

Time until the main text is readable, which you can see by watching a throttled load.

Time until a tap does something, which you feel on an old phone.

And the figures that actually matter to the business: enquiries, bounce rate on the pages you changed, and pages viewed per visit.

Record those before the work and compare afterwards, over whole weeks.

Where the score improved and none of those moved, that is the finding.

Field data settles it

Where you have enough traffic.

Measurements collected from real visitors describe what actually happened rather than what a test machine experienced.

They lag by weeks, which is inconvenient and is why people prefer the instant score.

But an improvement that shows in field data is real, and one that shows only in a test score may not be.

Check the field figures a month after any performance work, which is the honest verification.

For a small site with too little traffic for field data, the visitor-experience questions above are the substitute.

Diminishing returns are real

Worth saying, since the score encourages chasing the last points.

Going from eight seconds to three is transformative and going from three to two point five is not noticeable to anybody.

The score does not reflect that, since it treats the last twenty points much like the first twenty.

So a site already loading in about two seconds has very little available, whatever the number says.

Spend the effort on something else at that point, which is a decision the score will never suggest.

Ask what was actually done

Before and after paying for this.

A list of specific changes, in plain terms, with what each was expected to improve.

Deferred, minified, and optimised are categories rather than descriptions, and a competent supplier will name the actual files and settings.

Ask specifically whether anything was deferred that a visitor needs immediately, since that is where the fake improvement lives.

And ask what the before and after figures were for the visitor-facing measurements rather than for the score.

Two scores for the same page disagree

Worth expecting, since it causes unnecessary alarm.

The same page tested twice minutes apart can produce noticeably different scores, because the test involves a real network and a shared machine.

So a drop of a few points means nothing, and people have commissioned work on the strength of exactly that.

Run any test three times and take the middle result, and treat differences smaller than about ten points as noise.

Where a figure has genuinely moved, it will still be moved on a rerun tomorrow, which is the cheapest possible check.

Some of the recommendations do not apply to you

Worth knowing, since the list attached to a score is generic and produces wasted effort.

Tools flag opportunities by pattern rather than by whether they matter on your site, so a page with three small scripts gets the same advice as one with forty.

The saving shown beside each item is the useful part, and anything offering under about a tenth of a second is not worth anybody's attention.

Work down the list by that figure and stop when the numbers become small, rather than treating it as a set of tasks to complete.

A green result is not the goal, since the tool does not know what your site is for or who visits it.

The counter-case

Scores are not worthless.

They are a reasonable first indicator, they identify genuine problems, and the recommendations attached to them are frequently correct.

A site scoring very badly almost certainly has real problems, so a low score is more informative than a high one.

And for somebody with no other measurement, a score is considerably better than an impression.

Use it to find problems, verify improvements against visitor-facing measurements, and stop when the page is fast enough.

What to do

  1. Record enquiries before any work.
  2. Note when the text becomes readable.
  3. Note when a tap responds.
  4. Ask what a visitor now gets sooner.
  5. Check nothing needed was deferred.
  6. Compare field data after a month.
  7. Stop at fast enough.

Step four is the question that separates the two kinds of improvement, and anybody who cannot answer it plainly has probably improved the measurement instead.

Why two tests disagree is covered in why your desktop scores differ from mobile.


Frequently asked questions

What is a score?

A single number summarising several measurements, weighted by somebody else's judgement. Useful as a rough indicator and not the thing you actually care about.

How can a score improve without the experience?

By deferring things still needed, hiding work until after the measurement, improving a page nobody visits, or removing something that was doing a job.

What is the common trick?

Delaying scripts until a visitor interacts. The measurement finishes before the work starts, and the visitor still waits the moment they touch the page.

What question should I ask?

What does a visitor now get sooner. If the answer is that the text appears earlier or the page is usable sooner, it is real. If it is about measurement timing, it is not.

What should I measure instead?

Time until the main text is readable, time until a tap responds, and enquiries and bounce rate on the pages changed, recorded before and compared after.

When should I stop?

Around two seconds. Going from eight to three is transformative and three to two and a half is unnoticeable, which the score does not reflect.

West Coast Media Solutions Inc. provides web design, web development, hosting, digital marketing, and business consulting to organisations across Canada, drawing on more than twenty-five years in the field.

Paid for performance work and the score rose?

Ask what a visitor now gets sooner. If nobody can answer plainly, the measurement improved rather than the site.

Start a Conversation