Lab scores

I asked an AI to guess three speed scores. Then I measured them.

The guesses sounded right. One of them was 28 points away from the real number.

Three websites. What an AI guessed each one would score, next to what three measured runs actually gave.
Guessed on the left. Measured on the right.

Somebody told me last week that my real competition is not another speed tool. It is AI. Developers do not open a tool any more. They describe the problem to a model and take the answer.

He was right, so I ran a test on myself.

I picked three well known websites. Before loading a single page, I asked an AI two questions about each one. What will this page score on a mobile speed test, and how many points will that score move between three identical runs?

Then I measured all three and compared.

The rules

The guesses were written down first and not changed. That is the whole point. If you measure first, you are not testing anything.

The model got the address and nothing else. No page load, no tool, no network log. That is exactly what you get when you ask an AI instead of measuring.

Every measurement was three runs of the same page, on a mobile profile, on the same machine, with nothing changed in between. The machine reported itself steady on all three sites, so none of the movement below is the measuring box having a bad minute.

What happened

SiteAI guessedMeasuredOut by
gov.uk92997 points
bbc.com/news38479 points
smashingmagazine.com558328 points

The AI was low every single time. Not randomly wrong in both directions. Low, three times out of three.

The worst one is the interesting one. It put Smashing Magazine at 55, which is the sort of number that would start a rebuild. The real score is 83. If you had quoted 55 to a client you would have sold them work they do not need.

The second question, and where it got closer

I also asked how far each score would move between three identical runs. Here the AI did better than I expected, and it would be dishonest to pretend otherwise.

SiteThree real runsReal movementAI guessed
gov.uk99, 99, 990 points3
bbc.com/news50, 44, 476 points8
smashingmagazine.com89, 83, 836 points6

One exact hit. Two close. So the AI has a decent feel for which kinds of pages are jumpy.

But look at gov.uk again. The AI said it would move about 3 points. It moved zero. Three times, the same number.

That difference is not trivia. It changes what you are allowed to believe. On a page that never moves, a 2 point gain is real and you can say so. If you had trusted the guess of 3, you would have thrown that 2 point win away as noise.

Why a guess cannot get there

None of this means the model is bad. It means it is doing something else.

A model has read a great deal about how websites are built. It knows gov.uk is lean and BBC News carries a lot of advertising. That knowledge is real and it is why the guesses are in roughly the right area.

What it has not done is open your page on a slow phone, three times, and watch what happened. It has no measurement. So it produces the most reasonable sounding number, with no way to know it is 28 points out, and no way to tell you it is unsure.

That is the difference between advice and evidence. Advice is everywhere now and it is often good. Evidence still has to be taken.

What to do with this

Keep using AI. It is genuinely good at the part after the measurement. Give it a named file and a real number and it will help you fix the thing.

Just do not let it supply the number. Ask it what to change, not what your score is, and never let a guessed figure reach a client.

And before you celebrate a fix, find out how far the page moves on its own. On gov.uk that is zero, so every point counts. On BBC News it is six, so a five point improvement proves nothing at all.

You can do this by hand. Open Chrome, run Lighthouse three times on the same URL, change nothing, and subtract the lowest score from the highest. Fifteen minutes.

Or paste the address at pafcore.site and get the same three runs, the gap between them, and the name of the file costing you the most. It is free and it does not ask you to sign up. Everything in this article was measured with it, and you can re run all three sites yourself and check my numbers.

Repeat it on your own site

This test took about twenty minutes and you can run it on work you actually care about.

  1. Pick a page you know well.
  2. Ask your favourite model what it will score on mobile, and how much it moves between runs. Write both answers down.
  3. Measure it three times without changing anything.
  4. Compare.

If the guess lands close, good. You have learned that your page is ordinary enough to be predictable. If it lands 28 points away, you have learned something more useful, which is that nobody knew until somebody measured.