Skip to main content
Transformidy

Article

The Algorithm and the Randomized Trial

Two education systems built algorithms to help make decisions about individual students at scale. One imposed its algorithm nationally, on results day, with no pilot and no override. The other tested its tool for years, on a smaller population, before trusting it broadly, and built it to route stude

Published
September 8, 2026
Updated
August 19, 2026
Reading time
9 min
Paper-cut editorial illustration for The Algorithm and the Randomized Trial

The Global Signal

England's exam regulator, Ofqual, used a statistical model in 2020 to standardize grades after in-person exams were cancelled, weighting a school's historical performance heavily against individual teacher predictions. When results were released, nearly 36 percent of A-level grades came in one grade lower than teachers had predicted and 3 percent came in two grades lower, a pattern that disproportionately affected strong students at schools with weaker historical results. Public backlash forced a government reversal within days of the August 13, 2020 A-level results, and Ofqual's chief regulator, Sally Collier, resigned on August 25, 2020.

Georgia State University took a different path with a narrower tool. Its Pounce chatbot, piloted from 2016 to reduce "summer melt," the phenomenon of admitted students failing to enroll, was tested over multiple years before being trusted at scale. Reported outcomes include summer melt falling from roughly 19 percent to roughly 9 percent in the pilot population, with first-generation students reported to benefit disproportionately, plausibly because they are more reluctant to ask questions in person.

Visible cost
36%

Share of UK A-level grades lowered by at least one grade from teacher predictions in 2020

Directly confirmed via Wikipedia's sourced account of the controversy; an additional 3 percent were lowered by two grades.

One trusted its model nationally on its first real use. The other spent years finding out, at a survivable scale, what the model got wrong.

The Hidden Signal

Ofqual's model was applied to every affected student in the country simultaneously, with no smaller-scale test run to reveal how it behaved on real students before the entire system depended on it, and with no mechanism for an individual teacher's judgment to override a specific grade. Georgia State's chatbot was piloted for years on a single institution's incoming class before broader claims were made about it, and it was built specifically to route a student toward a human advisor when a question exceeded what the tool could answer, not to make a final decision about that student on its own. Consider what the alternative looks like at Georgia State's scale: a university deploys a similar tool nationally in its first year, skips the multi-year pilot, and discovers only after the fact whether the tool actually helps the students it was meant to help.

What changes

What changes when a model is piloted before it is trusted

A model's aggregate accuracy says nothing about whether any one individual's case was decided fairly.

Piloting at a smaller, recoverable scale and building an explicit override path for the individual case are what separated these two outcomes.

Why the Visible Metric Misleads

A model's statistical soundness in aggregate, correctly predicting the national grade distribution, correctly reducing the average student's likelihood of a specific bad outcome, says nothing about whether any single decision made about a specific individual was fair or correct. Ofqual's approach was defensible at the level of overall distribution and indefensible at the level of any one student whose own performance did not match their school's history. The more revealing test for any algorithm applied to individual people is not whether the aggregate looks right; it is whether the system was piloted at a smaller scale first, and whether an individual's own circumstance can still override the model's output.

The Leadership Move

The right move is not avoiding statistical models for decisions that affect many people at once. It is piloting any such model at a smaller scale before trusting it broadly, and building an explicit override path for the individual case the aggregate model was never designed to get right.

Ownership

A national regulator or a university's leadership owns the decision to trust a model at scale. The specific individuals affected, an exam candidate, an admitted student, have no equivalent power to challenge the model's specific output unless an override path was built in deliberately, before deployment, not improvised afterward under public pressure.

Tradeoff

Piloting a tool for years before trusting it broadly, as Georgia State did, is slower than deploying nationally in a single results cycle, and a multi-year pilot costs real time during which the problem the tool is meant to solve keeps happening. Ofqual's faster, unpiloted approach shows the cost on the other side: a national reversal, a regulator's resignation, and a cohort of students whose results were publicly contested.

Human consequence

An Ofqual student experienced a grade lower than their own teacher's assessment, with no path to argue their individual case before the reversal. A Georgia State student with a question the chatbot could not answer was, by design, routed to a person rather than left with an automated answer standing as final.

Implication for Operators

Any organization applying a statistical model to decisions about individual people should assume that the model's soundness in aggregate says nothing about its fairness in any one case, and that only a genuine pilot at smaller scale, plus an explicit override path to a person, closes that gap. Ofqual's national, unpiloted rollout and Georgia State's multi-year, human-routed pilot are the same category of tool built on opposite assumptions about what individual fairness requires.

Both Ofqual and Georgia State built a model meant to help make a decision about many people using limited human capacity. One trusted its model nationally on its first real use. The other spent years finding out, at a survivable scale, what the model got wrong before anyone let it touch outcomes broadly.

The real story here is not that one algorithm was better designed than the other. It is that one was tested at a scale where mistakes were recoverable, and the other was not.

Next Move

Reflection question

Name a model your organization applies to decisions about individual people at scale. Was it piloted at a smaller, recoverable scale first, and does an individual case have an explicit path to override its output?

Practical step

Before scaling any model that affects individual outcomes, run it at a small, recoverable scale first, and build a named override path before broader deployment, not after a public reversal forces one.

Soft invitation

Transformidy's decision-workflow review starts by asking whether your highest-stakes automated decision was ever piloted at a smaller scale before it reached everyone it now affects.

Signal checkDecision BlindnessRegistry-backed

Before automating a workflow or routing rule, how explicitly does your organization decide what the system should optimize for and who owns exceptions?

FAQ

Was Ofqual's model statistically inaccurate?

At the level of the national grade distribution, reporting on the case describes the model as broadly consistent with historical patterns. The failure was at the individual level: nearly 36 percent of A-level grades came in one grade lower than teacher predictions, with no override path for students whose own performance diverged from their school's history.

How much did Georgia State's chatbot actually change outcomes?

Reported figures include summer melt falling from roughly 19 percent to roughly 9 percent in the pilot population. This draft does not use the broader persistence-rate and graduation estimates as final claims until they are confirmed against a primary institutional source.

Does Pounce make decisions about students, or just answer questions?

Reporting describes it as designed to answer routine questions and route a student to a human advisor when a question exceeds what the tool can handle, not to make a final decision about the student's own case.

Could Ofqual have avoided its outcome with the same underlying model?

A smaller-scale pilot before national rollout, and an explicit override path for individual teacher assessment, are the two structural elements this case suggests were missing, not necessarily a different underlying statistical approach.

What is the single most transferable lesson from comparing these two cases?

Pilot before scaling, and build an explicit path for an individual case to override the model, before a model applied to many people at once is trusted with a decision that affects any one of them specifically.