Speed
Why your scores bounce around so much
A single result is mostly noise. The interesting question is how many you need before it stops being.
Take the reaction test five times in ten minutes and you will get five different answers, sometimes spanning 60 milliseconds. Nothing about you changed in ten minutes. So what did?
Every score is a real number plus noise
Any measurement is the thing you are trying to measure plus everything else that got in. For a browser test, “everything else” is substantial:
- Your hardware, which contributes 30 to 40ms and varies slightly run to run
- Momentary attention — whether your mind wandered on round three
- Where the targets happened to fall, on the aim trainer
- Which words came up, on the memory tests
- How warmed up you were
None of that is you-today-versus-you-yesterday. It is the measurement being imprecise, and no amount of concentrating removes it.
Why the median is the headline everywhere on this site
One distracted round at 600ms moves a five-round average by about 70 milliseconds. It moves the median by nothing.
That is the whole reason the median is shown first. A lapse of attention is not what anyone means by “my reaction time”, and a statistic that a single lapse can move by 70ms is not measuring what it claims to.
The same logic drives the other design decisions here: five rounds rather than one, thirty targets rather than ten, twenty Stroop trials split exactly half and half. Each is a way of letting the noise cancel rather than dominate.
How many attempts a trend actually needs
More than people expect, and the number depends on how noisy the test is.
| Test | Roughly how many before a trend is readable |
|---|---|
| Reaction time | 10–15 sessions |
| Typing speed | 8–10 |
| Aim trainer | 15–20 |
| Sequence and visual memory | 20+ |
The memory tests need the most because a single lapse ends the run outright — there is no averaging within an attempt to soften it. A sequence run is pass/fail at every step, so its scores scatter much more widely than a timed test’s.
Three attempts tell you almost nothing on any of them. That is not a failure of the tests; it is what measuring a noisy thing a small number of times gets you.
Your best score is the least useful one
A personal best is, by definition, the attempt where the noise happened to run most in your favour. It is a record of your luckiest run, not your ability.
This has a specific consequence people run into: your best is very hard to beat, and that is expected. If your median improves by 5%, your best may not move for weeks, because it was already an outlier. Judging progress by your best is judging it by the one number designed not to move.
Watch the median. It responds to real change and ignores lucky rounds, which is exactly backwards from what feels satisfying and exactly right for what you want to know.
Regression to the mean, and why you should distrust an explanation
If you have an unusually bad run and then do something — stretch, get a coffee, change your grip — your next run will very probably be better.
Not because of the thing you did. Because an extreme result is usually followed by a more typical one, whatever happens in between. This is regression to the mean, and it is the most reliable manufacturer of false explanations in existence.
It is why people are confident about interventions that do nothing: you only reach for a remedy after a bad result, and a bad result is the moment when improvement is most likely regardless. The only way through it is to compare a fortnight with the change against a fortnight without, rather than one run against the one before it.
What genuinely moves the number
Once you can separate signal from noise, a few things move it enough to see:
Sleep, more than anything else on this list, and it shows up first as occasional very slow rounds rather than a uniformly slower average — so watch your worst rounds, not just your median. See what one bad night costs.
Time of day, by a consistent margin. Compare like with like or the afternoon dip will masquerade as a trend.
Practice, for the first fortnight only. Early gains are you learning the task — where to look, how the controls behave. It plateaus quickly, and the interesting measurement starts after it does.
Your equipment. Changing laptop or mouse resets the baseline entirely, and every comparison across that change is meaningless.
Reading a history strip honestly
The bars under each test are twenty attempts, oldest first, and there are two ways people misread them.
Seeing a trend in noise. Any random sequence contains runs that look like trends. Three improving attempts is not improvement; it is what three attempts do about a quarter of the time. If you can cover the last three bars and no longer believe in the trend, it was not there.
Missing a trend because of one bad day. The opposite error. A single terrible attempt in an otherwise improving run is a lapse, not a reversal, and the median is unmoved by it — which is precisely why the median is the number shown.
The useful habit is to look at the middle of the distribution across a fortnight rather than the ends of it. The best bar and the worst bar are the two least informative marks on the chart.
Why averaging more rounds does not fix everything
There is a limit to what repetition buys, and it is worth knowing where it sits.
Averaging reduces random noise — the lapses, the lucky target placements, the millisecond of scheduling jitter. Take enough attempts and those cancel out.
It does nothing about systematic error. If your monitor adds 16 milliseconds, it adds them to every round, and a thousand rounds will estimate your reaction-plus-monitor very precisely while telling you nothing more about the reaction. That is why the hardware caveat on the reaction test is not fixed by taking more attempts, and why comparing your number to someone else’s remains meaningless however diligent you both are.
The distinction is the whole reason this site compares you against yourself. Systematic error that is identical on both sides of a comparison cancels out. Your monitor is in Monday’s score and Tuesday’s equally, so the difference between them is clean even though neither number is.
What a plateau actually means
Most people improve for a fortnight and then stop, and read that as hitting their ceiling. It usually is not.
The early gain was task learning — where to look, how the controls feel, what the test expects. That is genuinely finite and it runs out. What follows is not your limit; it is the point at which the measurement finally starts being about you rather than about your familiarity with a web page.
The interesting phase begins at the plateau. From there the variation you see is mostly state: sleep, time of day, whether you are ill, how much you have had to drink. That is the signal the whole exercise exists to expose, and it was buried under the learning curve until now.
The one comparison that survives all of this
You, against your own history, on unchanged equipment, at the same time of day, after the first fortnight, using the median rather than the best.
That sounds like a lot of conditions. It is the difference between a number that tells you something and a number that tells you what the weather was doing in the measurement. The daily round exists to make that comparison the easy one to make: eight tests, one score, one line to watch.