Why Online Home Value Estimates Get It Wrong
Estimates miss for a short list of repeatable reasons: invisible condition, thin comp density, non-disclosure states, submarket boundaries, stale records, and homes that are unusual for their area.

Online home value estimates get it wrong for a short list of repeatable reasons, and nearly all of them collapse to one thing: the model cannot see what a person standing in the driveway notices in five seconds. The two largest causes are condition, which is not recorded anywhere in public property data, and geography, where values change block to block in ways a radius search does not respect. None of these are exotic failures. They are the normal operating limits of estimating a physical asset from paperwork. What separates a good AVM from a bad one is not whether it has these problems, it is whether it tells you when it has them.
1. Condition is not a field in the data
A county assessor records that a house exists, how big it is, how many bedrooms and bathrooms it has, and roughly when it was built. Nothing in that record distinguishes a home with a new roof, a renovated kitchen, and refinished floors from the identical floor plan next door with a failing HVAC system and original 1978 finishes. Those two houses can differ by a third of their value and look identical in the data. Models compensate by implicitly assuming typical condition for the area, which is the correct guess on average and wrong on exactly the properties where the difference matters. This is the single biggest source of error in residential AVMs, and it is why we estimate condition from listing photos, an approach that is genuinely informative and still an estimate rather than an inspection.
2. There are not enough comparable sales
Comp density is the fuel an AVM runs on. In a dense subdivision of similar homes with steady turnover, a model can find a dozen genuinely comparable sales within a few months and half a mile, and the answer is well constrained. In a rural market, a low turnover neighborhood, or a small town, the model may have to reach three miles out and eighteen months back to find five loosely similar sales. Each of those relaxations adds error, and the errors compound. This is why the same model can be reliable in one county and shaky in the next one over, and why national accuracy figures conceal so much.
3. Sale prices are not public everywhere
The United States has non-disclosure states, where the price paid for a property is not recorded in public records. Texas is the largest and best known example, and roughly a dozen states are non-disclosure or partially disclosing, with practice varying by county. In these markets an AVM is not learning from actual transaction prices. It is inferring them from listing data, loan amounts, transfer taxes where they exist, and modeled relationships. That inference is often decent and it is never as good as knowing. Anyone quoting a single national accuracy figure without noting this distinction is glossing over a structural difference between markets.
4. Submarket boundaries are invisible to a radius search
Value does not decay smoothly with distance. It steps. A school attendance line, a municipal boundary, a highway, a rail corridor, a flood zone edge, a different HOA, or the line between two subdivisions built a decade apart can each produce a sharp price break across a single street. A comp search defined by distance will happily cross those lines, because from the model's perspective a sale 0.4 miles away is closer than one 0.7 miles away. If the near one is across a boundary into a cheaper submarket and the far one is in the subject's own subdivision, the distance rule picks the wrong comp. The estimate is then wrong for a reason that has nothing to do with the model's math.
5. The records are stale
Every input has latency. Deed recording can lag a closing by weeks. Assessor files update annually in many jurisdictions. Permit data, where it is available at all, is often months behind and inconsistently coded. In a market that is moving several percent a quarter, comps from nine months ago are describing a different market, and an estimate built on them is describing the past with present tense confidence.
6. The house is unusual for its area
Models are interpolators. They are reliable inside the cloud of properties they have seen and unreliable outside it. A 4,200 square foot house on a street of 1,600 square foot houses, a home on five acres in a neighborhood of quarter acre lots, a converted church, a house with an accessory dwelling unit in an area where those are rare: all of these force the model to extrapolate. Extrapolation in valuation almost always produces the same failure shape, which brings us to the next reason.
7. Models regress toward the mean at the extremes
Statistical models pull predictions toward the center of their training distribution. The practical consequence is that the cheapest homes in a market tend to be over-valued and the most expensive homes tend to be under-valued, in a characteristic U shaped error curve when you plot error against price. This matters most to investors, because distressed and low priced properties sit precisely in the over-valued tail. It is one of the reasons an automated ARV can look encouraging on a deal that does not work.
8. Renovated and unrenovated homes sit in the same comp pool
Two houses on the same street, same size, same year built. One was fully renovated last spring and sold for a premium. The other has never been updated. If the renovated sale enters the comp pool for the unrenovated subject without a condition adjustment, the subject inherits a value it cannot support. The reverse happens too, and is less discussed: many data pipelines filter distressed and as-is sales out of the comparable pool on the theory that they are not arm's length transactions. That filtering is defensible for valuing a normal home and quietly disastrous for valuing a fixer-upper, because it removes exactly the sales that describe what the subject is worth today. There is a full treatment in why estimates break on fixer-uppers.
What the model sees versus what it misses
| What the data contains | What actually moves the price |
|---|---|
| Square footage on record | Whether the square footage is finished, permitted, and usable |
| Bed and bath counts | Layout, flow, whether the fourth bedroom is a converted garage |
| Year built | Year of the last real renovation |
| Lot size | Slope, usability, easements, what backs onto the lot |
| Recent nearby sales | Whether those sales are in the same submarket |
| Property type code | Ownership form, HOA rules, whether the project is warrantable |
| Listing photos, sometimes | Condition, finish level, deferred maintenance |
| Nothing at all | Roof age, foundation, HVAC, mold, permits pulled without records |
What can actually be done about this?
Some of it is fixable with engineering. Comp selection can be constrained to real submarkets rather than raw distance. Size bands can scale with the subject rather than using a fixed tolerance. Condition can be estimated from photos instead of assumed. Distressed sales can be handled deliberately rather than filtered by default. Some of it is not fixable at all. No model will know the foundation is cracked. The correct response to that is not a better marketing claim, it is a published error bar and a visible comp set, so a user can see the reasoning and override it.
The test that matters
Every one of these failure modes is invisible if a provider grades its own homework after the fact. That is why we freeze every estimate at the moment a property lists and grade it against the real closing price when the sale records. The results, including the cases where the model does poorly, are published at /accuracy. If you would rather pressure test a number yourself, open the comps in the CMA report or run the deal with a rehab budget in the deal analyzer.
Frequently Asked Questions
Why is my home value estimate too low?
The most common causes are a recent renovation the data never recorded, comps pulled from a cheaper adjacent submarket, or a property that is larger or better than the homes around it, which pulls the estimate toward the neighborhood average. Check the comparable sales first; the reason is almost always visible there.
Do estimates work in Texas and other non-disclosure states?
They work, less well. In non-disclosure states the price paid for a property is not public record, so the model infers transaction prices from listing data and other signals rather than knowing them. That inference is often reasonable and it is never as good as observing the actual price, so treat estimates in these markets with wider error bars.
Do home value estimates account for renovations?
Only if the renovation left a trace the model can read, such as a listing that mentions it or photos showing new finishes. Permits are inconsistently available and often uncoded. A high quality renovation on a home that has not been listed since is effectively invisible, which is why owner-updated details can move an estimate meaningfully.
Ready to Start Investing Smarter?
Join 2,000+ investors using Resideline.
Start free with 3 reports a month.