Article
The Algorithm and the Randomized Trial
Two education systems built algorithms to help make decisions about individual students at scale. One imposed its algorithm nationally, on results day, with no pilot and no override. The other tested its tool for years, on a smaller population, before trusting it broadly, and built it to route stude
- Published
- September 8, 2026
- Updated
- August 19, 2026
- Reading time
- 9 min

The Global Signal
England's exam regulator, Ofqual, used a statistical model in 2020 to standardize grades after in-person exams were cancelled, weighting a school's historical performance heavily against individual teacher predictions. When results were released, nearly 36 percent of A-level grades came in one grade lower than teachers had predicted and 3 percent came in two grades lower, a pattern that disproportionately affected strong students at schools with weaker historical results. Public backlash forced a government reversal within days of the August 13, 2020 A-level results, and Ofqual's chief regulator, Sally Collier, resigned on August 25, 2020.
Georgia State University took a different path with a narrower tool. Its Pounce chatbot, piloted from 2016 to reduce "summer melt," the phenomenon of admitted students failing to enroll, was tested over multiple years before being trusted at scale. Reported outcomes include summer melt falling from roughly 19 percent to roughly 9 percent in the pilot population, with first-generation students reported to benefit disproportionately, plausibly because they are more reluctant to ask questions in person.
Share of UK A-level grades lowered by at least one grade from teacher predictions in 2020
Directly confirmed via Wikipedia's sourced account of the controversy; an additional 3 percent were lowered by two grades.
One trusted its model nationally on its first real use. The other spent years finding out, at a survivable scale, what the model got wrong.
What changes when a model is piloted before it is trusted
A model's aggregate accuracy says nothing about whether any one individual's case was decided fairly.
Piloting at a smaller, recoverable scale and building an explicit override path for the individual case are what separated these two outcomes.
Why the Visible Metric Misleads
A model's statistical soundness in aggregate, correctly predicting the national grade distribution, correctly reducing the average student's likelihood of a specific bad outcome, says nothing about whether any single decision made about a specific individual was fair or correct. Ofqual's approach was defensible at the level of overall distribution and indefensible at the level of any one student whose own performance did not match their school's history. The more revealing test for any algorithm applied to individual people is not whether the aggregate looks right; it is whether the system was piloted at a smaller scale first, and whether an individual's own circumstance can still override the model's output.
The Leadership Move
The right move is not avoiding statistical models for decisions that affect many people at once. It is piloting any such model at a smaller scale before trusting it broadly, and building an explicit override path for the individual case the aggregate model was never designed to get right.
- Ownership
A national regulator or a university's leadership owns the decision to trust a model at scale. The specific individuals affected, an exam candidate, an admitted student, have no equivalent power to challenge the model's specific output unless an override path was built in deliberately, before deployment, not improvised afterward under public pressure.
- Tradeoff
Piloting a tool for years before trusting it broadly, as Georgia State did, is slower than deploying nationally in a single results cycle, and a multi-year pilot costs real time during which the problem the tool is meant to solve keeps happening. Ofqual's faster, unpiloted approach shows the cost on the other side: a national reversal, a regulator's resignation, and a cohort of students whose results were publicly contested.
- Human consequence
An Ofqual student experienced a grade lower than their own teacher's assessment, with no path to argue their individual case before the reversal. A Georgia State student with a question the chatbot could not answer was, by design, routed to a person rather than left with an automated answer standing as final.
Implication for Operators
Any organization applying a statistical model to decisions about individual people should assume that the model's soundness in aggregate says nothing about its fairness in any one case, and that only a genuine pilot at smaller scale, plus an explicit override path to a person, closes that gap. Ofqual's national, unpiloted rollout and Georgia State's multi-year, human-routed pilot are the same category of tool built on opposite assumptions about what individual fairness requires.
Both Ofqual and Georgia State built a model meant to help make a decision about many people using limited human capacity. One trusted its model nationally on its first real use. The other spent years finding out, at a survivable scale, what the model got wrong before anyone let it touch outcomes broadly.
The real story here is not that one algorithm was better designed than the other. It is that one was tested at a scale where mistakes were recoverable, and the other was not.
Before automating a workflow or routing rule, how explicitly does your organization decide what the system should optimize for and who owns exceptions?
FAQ
Was Ofqual's model statistically inaccurate?
At the level of the national grade distribution, reporting on the case describes the model as broadly consistent with historical patterns. The failure was at the individual level: nearly 36 percent of A-level grades came in one grade lower than teacher predictions, with no override path for students whose own performance diverged from their school's history.
How much did Georgia State's chatbot actually change outcomes?
Reported figures include summer melt falling from roughly 19 percent to roughly 9 percent in the pilot population. This draft does not use the broader persistence-rate and graduation estimates as final claims until they are confirmed against a primary institutional source.
Does Pounce make decisions about students, or just answer questions?
Reporting describes it as designed to answer routine questions and route a student to a human advisor when a question exceeds what the tool can handle, not to make a final decision about the student's own case.
Could Ofqual have avoided its outcome with the same underlying model?
A smaller-scale pilot before national rollout, and an explicit override path for individual teacher assessment, are the two structural elements this case suggests were missing, not necessarily a different underlying statistical approach.
What is the single most transferable lesson from comparing these two cases?
Pilot before scaling, and build an explicit path for an individual case to override the model, before a model applied to many people at once is trusted with a decision that affects any one of them specifically.
Related intelligence
Article
The Report That Audited Robots and Hallucinated Its Own Citations
A firm whose entire service is verification can still fail to verify its own output, if nobody explicitly owns checking AI-generated content before a client sees it. The irony sharpens considerably when the report in question is reviewing a government's own automated penalty system.
Article
The Playbook Before the Rollout
Two governments built automated systems that made consequential decisions about their citizens. One never wrote down, before deployment, what pre-launch testing an AI system had to pass. The other built a public assurance playbook years earlier, and has kept updating that playbook as the technology
Article
The Algorithm That Toppled a Cabinet
A warning can exist, in an organization's own files, written by exactly the person whose job is to catch this kind of problem, and still change nothing at all. Writing a memo and routing it to someone with the authority to act on it are two separate acts, and a government can be entirely honest that