Article
When the Model's Confidence Became the Company's
A pricing model trained on years of stable data will keep producing confident numbers long after the market underneath it has moved. The model cannot tell the difference. The organization deploying it has to, and that requires someone whose specific job is asking whether the model's confidence still
- Published
- August 20, 2026
- Updated
- August 19, 2026
- Reading time
- 10 min

The Global Signal
Zillow announced on November 2, 2021, that it was shutting down Zillow Offers, its algorithmic home-buying business, after disclosing in its third-quarter results a $304 million inventory write-down within its home-buying segment, with an additional $240 million to $265 million in losses anticipated for the fourth quarter, a combined figure widely reported as exceeding $500 million. The company's own disclosure statement attributed the shutdown directly to "unpredictability in forecasting home prices" that "far exceed[ed] what we anticipated," and the wind-down included cutting approximately 25 percent of Zillow's workforce. Zillow had purchased thousands of homes at algorithm-generated prices during 2021's unusually volatile housing market and found itself unable to resell many of them without a loss.
Zillow's own Zestimate model has historically carried a published national median error rate in the high single digits, with a wider margin of error, commonly cited around 14 percent, for homes not currently listed for sale, a figure Zillow itself has disclosed as part of its standard model documentation. That baseline error rate is not new information; it existed before, during, and after the Offers program. What changed in 2021 was the cost of that error rate at scale, during a period when home values were moving faster than the historical data underlying the model could account for.
Inventory write-down Zillow disclosed for its algorithmic home-buying program in Q3 2021
Directly confirmed company disclosure figure; Zillow additionally anticipated $240-265 million in further Q4 losses, making the widely reported '$500 million-plus' total an aggregate of two confirmed, disclosed figures rather than a single reported number.
The model does not know the difference. The organization using it has to.
What changes when confidence is checked against real time
A model's historical accuracy describes its past. It cannot tell you whether current conditions still resemble that past.
A standing, separate check against real-time market signals, with authority to pause scale purchasing, catches a shift the model's own output cannot flag on its own.
Why the Visible Metric Misleads
A model's published historical accuracy answers the question of how it performed on average across the conditions it has already seen. It does not answer the question that matters most in a volatile market: is the current environment still inside that range, or has it moved outside it in a way the model cannot detect on its own. The more revealing practice is not distrusting the model's accuracy statistic, which Zillow disclosed honestly and consistently, but building a separate, standing check comparing the model's purchase prices against real-time, on-the-ground signals, exactly the kind local agents were already producing publicly, before committing capital at the model's stated confidence level.
The Leadership Move
The right move is not to avoid algorithmic pricing models, which can genuinely outperform manual estimation across stable conditions. It is to build an explicit, standing check, separate from the model itself, that compares its outputs against real-time market signals specifically during periods of unusual volatility, with a named authority to pause purchasing at scale when that check flags a divergence.
- Ownership
The data science team that builds a pricing model typically owns its statistical accuracy under normal conditions. Operations and finance leadership need to own the separate, standing question of whether current conditions still fall inside the range the model was built for, since the team that built the model has little structural reason to be the one that decides it should be trusted less right now.
- Tradeoff
Building a real-time volatility check that can pause an entire purchasing program costs speed and scale exactly when a business is trying to capture a fast-moving opportunity, which is precisely why it tends not to get built until after the fact. The alternative, continuing to execute on a model's stated confidence through a period local market participants were already signaling as unusual, cost Zillow several hundred million dollars in a matter of months.
- Human consequence
Zillow's own employees, roughly a quarter of its workforce, were laid off as part of the program's closure, and the shareholders and homeowners on the other side of thousands of individual transactions absorbed a pricing model's confidence that had quietly stopped matching the market it was operating in.
Implication for Operators
Any organization deploying a pricing or forecasting model at scale should assume that the model's published historical accuracy describes its past performance, not a guarantee that current conditions still resemble the data it was trained on. The practical shift is building a standing, separate check against real-time market signals, particularly during periods when external, on-the-ground observers are already reporting something the model's own aggregate output has not yet reflected, rather than waiting for the model's own error to surface in a quarterly financial disclosure.
Zillow's pricing model was not secretly broken. Its published accuracy was honest, consistent, and unchanged. What changed was the market it was operating in, and no standing mechanism existed to notice that shift before local agents' public observations became a several-hundred-million-dollar write-down.
A bad model is not what happened here. A model's confidence trusted past the exact point external, real-time signals had already started contradicting it, with no standing check assigned to make that comparison, is what happened.
How predictable is demand for your products or services—would a significant market shift surprise you or would you see it coming?
FAQ
Was Zillow's Zestimate algorithm simply inaccurate?
Zillow's published error rates, a national median in the high single digits and a wider margin, commonly cited around 14 percent, for off-market homes, were disclosed honestly and were not new in 2021. The failure was not that the model's baseline accuracy changed; it was that 2021's unusually volatile market moved conditions outside the range that accuracy rate was built for, and no separate mechanism caught that shift before it caused significant losses.
Did anyone flag the problem before Zillow's own November 2021 announcement?
Bank of America's own market analysis, reported around the time of the shutdown, found Zillow paying prices in markets including Phoenix that reflected annual increases exceeding 20 percent, a divergence an external analyst comparing purchase prices against broader market data could identify independently of Zillow's own reporting.
How much did the Zillow Offers shutdown cost the company?
Zillow's own third-quarter 2021 disclosures included a $304 million inventory write-down, with an additional $240 million to $265 million in losses anticipated for the fourth quarter, a combined figure widely reported as exceeding $500 million.
Could Zillow have avoided this by simply using a more accurate model?
Not entirely. Even a highly accurate model, by definition, has some error rate, and that error rate can behave unpredictably once market conditions move outside its training range. The more durable fix is a separate, standing check for exactly that condition, not a search for a model with zero error.
Who should own deciding when a pricing model can no longer be trusted at its stated confidence level?
A function separate from the team that built and maintains the model's statistical accuracy, typically operations or finance leadership, with explicit authority to pause purchasing or lending decisions at scale when real-time signals diverge from the model's output.
Related intelligence
Article
The Algorithm and the Randomized Trial
Two education systems built algorithms to help make decisions about individual students at scale. One imposed its algorithm nationally, on results day, with no pilot and no override. The other tested its tool for years, on a smaller population, before trusting it broadly, and built it to route stude
Article
The Report That Audited Robots and Hallucinated Its Own Citations
A firm whose entire service is verification can still fail to verify its own output, if nobody explicitly owns checking AI-generated content before a client sees it. The irony sharpens considerably when the report in question is reviewing a government's own automated penalty system.
Article
The Playbook Before the Rollout
Two governments built automated systems that made consequential decisions about their citizens. One never wrote down, before deployment, what pre-launch testing an AI system had to pass. The other built a public assurance playbook years earlier, and has kept updating that playbook as the technology