Skip to main content
Transformidy

Article

When the Model's Confidence Became the Company's

A pricing model trained on years of stable data will keep producing confident numbers long after the market underneath it has moved. The model cannot tell the difference. The organization deploying it has to, and that requires someone whose specific job is asking whether the model's confidence still

Published
August 20, 2026
Updated
August 19, 2026
Reading time
10 min
Paper-cut editorial illustration for When the Model's Confidence Became the Company's

The Global Signal

Zillow announced on November 2, 2021, that it was shutting down Zillow Offers, its algorithmic home-buying business, after disclosing in its third-quarter results a $304 million inventory write-down within its home-buying segment, with an additional $240 million to $265 million in losses anticipated for the fourth quarter, a combined figure widely reported as exceeding $500 million. The company's own disclosure statement attributed the shutdown directly to "unpredictability in forecasting home prices" that "far exceed[ed] what we anticipated," and the wind-down included cutting approximately 25 percent of Zillow's workforce. Zillow had purchased thousands of homes at algorithm-generated prices during 2021's unusually volatile housing market and found itself unable to resell many of them without a loss.

Zillow's own Zestimate model has historically carried a published national median error rate in the high single digits, with a wider margin of error, commonly cited around 14 percent, for homes not currently listed for sale, a figure Zillow itself has disclosed as part of its standard model documentation. That baseline error rate is not new information; it existed before, during, and after the Offers program. What changed in 2021 was the cost of that error rate at scale, during a period when home values were moving faster than the historical data underlying the model could account for.

Visible cost
$304M

Inventory write-down Zillow disclosed for its algorithmic home-buying program in Q3 2021

Directly confirmed company disclosure figure; Zillow additionally anticipated $240-265 million in further Q4 losses, making the widely reported '$500 million-plus' total an aggregate of two confirmed, disclosed figures rather than a single reported number.

The model does not know the difference. The organization using it has to.

The Hidden Signal

A pricing model's published error rate is a statement about its behavior under the conditions it was built and tested on. It says nothing about how that error rate behaves once market conditions move outside that historical range, and a model has no internal mechanism for announcing that it has left its own reliable zone. Bank of America's own market analysis, reported alongside coverage of the shutdown, found Zillow paying prices in markets including Phoenix that reflected annual price increases exceeding 20 percent, a pattern an external analyst could identify by comparing Zillow's purchase prices against broader market data, well before Zillow's own aggregate reporting reflected the program as anything other than a going concern.

What changes

What changes when confidence is checked against real time

A model's historical accuracy describes its past. It cannot tell you whether current conditions still resemble that past.

A standing, separate check against real-time market signals, with authority to pause scale purchasing, catches a shift the model's own output cannot flag on its own.

Why the Visible Metric Misleads

A model's published historical accuracy answers the question of how it performed on average across the conditions it has already seen. It does not answer the question that matters most in a volatile market: is the current environment still inside that range, or has it moved outside it in a way the model cannot detect on its own. The more revealing practice is not distrusting the model's accuracy statistic, which Zillow disclosed honestly and consistently, but building a separate, standing check comparing the model's purchase prices against real-time, on-the-ground signals, exactly the kind local agents were already producing publicly, before committing capital at the model's stated confidence level.

The Leadership Move

The right move is not to avoid algorithmic pricing models, which can genuinely outperform manual estimation across stable conditions. It is to build an explicit, standing check, separate from the model itself, that compares its outputs against real-time market signals specifically during periods of unusual volatility, with a named authority to pause purchasing at scale when that check flags a divergence.

Ownership

The data science team that builds a pricing model typically owns its statistical accuracy under normal conditions. Operations and finance leadership need to own the separate, standing question of whether current conditions still fall inside the range the model was built for, since the team that built the model has little structural reason to be the one that decides it should be trusted less right now.

Tradeoff

Building a real-time volatility check that can pause an entire purchasing program costs speed and scale exactly when a business is trying to capture a fast-moving opportunity, which is precisely why it tends not to get built until after the fact. The alternative, continuing to execute on a model's stated confidence through a period local market participants were already signaling as unusual, cost Zillow several hundred million dollars in a matter of months.

Human consequence

Zillow's own employees, roughly a quarter of its workforce, were laid off as part of the program's closure, and the shareholders and homeowners on the other side of thousands of individual transactions absorbed a pricing model's confidence that had quietly stopped matching the market it was operating in.

Implication for Operators

Any organization deploying a pricing or forecasting model at scale should assume that the model's published historical accuracy describes its past performance, not a guarantee that current conditions still resemble the data it was trained on. The practical shift is building a standing, separate check against real-time market signals, particularly during periods when external, on-the-ground observers are already reporting something the model's own aggregate output has not yet reflected, rather than waiting for the model's own error to surface in a quarterly financial disclosure.

Zillow's pricing model was not secretly broken. Its published accuracy was honest, consistent, and unchanged. What changed was the market it was operating in, and no standing mechanism existed to notice that shift before local agents' public observations became a several-hundred-million-dollar write-down.

A bad model is not what happened here. A model's confidence trusted past the exact point external, real-time signals had already started contradicting it, with no standing check assigned to make that comparison, is what happened.

Next Move

Reflection question

Name a pricing, forecasting, or risk model your organization relies on at scale. Is there a standing check comparing its output against real-time, on-the-ground signals, or would a shift in conditions only surface once a financial result reflects it?

Practical step

For your highest-volume automated pricing or forecasting decision, establish a real-time divergence check against external market signals, with a named authority to pause scale purchasing when that check flags a gap.

Soft invitation

The first step in a Transformidy decision-workflow review is naming, in writing, who is authorized to pause your highest-volume automated pricing or forecasting decision, and what specific real-world signal would trigger that pause. If no one can answer either question today, that is the finding.

Signal checkRevenue Unknown Self-DiagnosticRegistry-backed

How predictable is demand for your products or services—would a significant market shift surprise you or would you see it coming?

FAQ

Was Zillow's Zestimate algorithm simply inaccurate?

Zillow's published error rates, a national median in the high single digits and a wider margin, commonly cited around 14 percent, for off-market homes, were disclosed honestly and were not new in 2021. The failure was not that the model's baseline accuracy changed; it was that 2021's unusually volatile market moved conditions outside the range that accuracy rate was built for, and no separate mechanism caught that shift before it caused significant losses.

Did anyone flag the problem before Zillow's own November 2021 announcement?

Bank of America's own market analysis, reported around the time of the shutdown, found Zillow paying prices in markets including Phoenix that reflected annual increases exceeding 20 percent, a divergence an external analyst comparing purchase prices against broader market data could identify independently of Zillow's own reporting.

How much did the Zillow Offers shutdown cost the company?

Zillow's own third-quarter 2021 disclosures included a $304 million inventory write-down, with an additional $240 million to $265 million in losses anticipated for the fourth quarter, a combined figure widely reported as exceeding $500 million.

Could Zillow have avoided this by simply using a more accurate model?

Not entirely. Even a highly accurate model, by definition, has some error rate, and that error rate can behave unpredictably once market conditions move outside its training range. The more durable fix is a separate, standing check for exactly that condition, not a search for a model with zero error.

Who should own deciding when a pricing model can no longer be trusted at its stated confidence level?

A function separate from the team that built and maintains the model's statistical accuracy, typically operations or finance leadership, with explicit authority to pause purchasing or lending decisions at scale when real-time signals diverge from the model's output.