Automated valuation models (AVMs) promise speed. University of Auckland researchers William Cheung and Edward Yiu argue the harder problem is trust — transparency, confidence intervals, bias correction and AI auditing in a market as geographically and culturally varied as New Zealand.
For licensed agents, that research lands next to something you already know: an AVM is not your appraisal, and it never replaces rule 10.2 judgement.
Primary source: University of Auckland — How AI is changing property valuations (also on The Conversation, October 2024).
New Zealand’s economy has been described, only half-jokingly, as a housing market with bits tacked on. Buying and selling property is national sport. Yet the wider public often has little understanding of how valuations and estimates are created — even when those figures influence bank lending conversations. AVMs made estimates faster. They did not automatically make them clearer.
This is a research-editor reading for NZ property professionals. It is not legal advice and not a critique of any single commercial AVM vendor.
PropertyLM take
Faster black boxes are still black boxes. Prefer tools that show their homework — and keep a licensed human owning the recommendation.
The black-box problem
Traditional valuation was slow and human. Valuers inspected, compared and applied judgement. That process is expensive and fallible — and it does not scale to every lending or screening use case. AVMs arrived as an efficiency answer.
Cheung and Yiu note AVMs started gaining traction in New Zealand in the early 2010s. Early versions leaned on limited sales records and council information. More advanced models now pull richer geo-spatial data from sources such as Land Information New Zealand. Efficiency improved. Transparency often did not.
Proprietary algorithms can leave homeowners and professionals with little insight into how a specific figure was produced. The authors call that opaqueness a real-world risk: without monitoring and correction, models can perpetuate imbalances — especially where regional, cultural and historical factors shape value in ways a national model may miss.
False precision is part of the cultural problem. A single mid-point looks decisive on a screen. A range with disclosed uncertainty looks less “salesy” — and more honest.
What “trust” requires (beyond a single number)
Their framework pushes past “faster estimate” into governance:
Disclosure of data sources, methods and error margins
Confidence intervals — a price range that shows how much the estimate might vary, so users see uncertainty instead of false precision
Bias correction — detecting and adjusting for systematically over- or undervalued segments (for example regional disparities or particular property types)
AI auditing — comparing model estimates against market-transacted prices for the same houses in the same period
Their method discussion includes resampling approaches for skewed property-value distributions so users can see variability around each estimate. Transparency alone, they argue, is not enough; trust has to be built into the system through accountability and bias correction.
Cheung & Yiu
“The rapid integration of AI into property valuation is no longer just about innovation and speed. It is about trust, transparency and a robust framework for accountability.”
Courts already treat AI output as something a human must check
The authors note New Zealand courts now require a qualified person to check AI-generated information used in tribunal proceedings. That is a useful cultural signal for agencies: if a tribunal expects a human checker on AI evidence, your vendor pack should not treat a black-box mid-point as gospel.
REA’s Gen AI guidance lands in the same neighbourhood for licensees: fact-check accuracy, relevance and completeness; Gen AI error is not a defence. The parallel is not accidental. Both point to human accountability as the non-negotiable layer.
How this maps to agent CMAs (rule 10.2)
Under the Real Estate Agents Act (Professional Conduct and Client Care) Rules 2012, rule 10.2 still requires an appraisal that is in writing, realistically reflects current market conditions, and is supported by comparable sales. Rule 10.3 requires a written explanation when comps are thin. REA’s appraisals guidance is blunt that you can’t solely rely on electronic appraisals or market estimates, and that viewing usually still matters under rule 5.1.
AVM vs CMA — keep the labels honest
AVM: statistical context. Useful. Never present it as your professional opinion.
CMA / appraisal (rule 10.2): your reasoned view, with comps you selected and can defend.
Registered valuation: a different professional product.
REINZ Companion Terms likewise state that Estimate outputs are not appraisals and must not be labelled as a CMA. If you show an automated estimate, call it automated. (Companion text refers to “rule 9.5”; the Code appraisal rule is 10.2 — keep your office materials accurate even when vendor tools use older shorthand.)
Practical takeaways for agencies
If you show an automated estimate, label it as automated and show a range where you can.
Prefer tools that expose why comps were included or excluded — not just a mid-point.
Assume anything AI-touched in a vendor pack needs the same human check you would give a junior’s draft.
Never let “the model said” become your defence under rule 10.2 or REA Gen AI guidance.
Train salespeople on the vocabulary: estimate, appraisal, valuation — three different products.
When a vendor asks “what does the computer say?”, answer with context and then return to your written comps-backed opinion.
PropertyLM’s stance matches the Auckland Uni direction of travel: traceable comps, sources you can open, and a human who owns the recommendation. Atlas and Newton are built to speed the evidence pack — not to claim an AVM replaces appraisal.
Why New Zealand’s diversity makes black boxes riskier
Cheung and Yiu emphasise geography and culture for a reason. A model that performs acceptably on a large metro sample can still misread lifestyle blocks, coastal hazards, character housing quirks, or thin rural evidence. National averages hide local truth. That is precisely why rule 10.2 insists on comparable sales in similar locations — and why rule 10.3 exists when the comps are not there.
Confidence intervals and bias correction are research language for a salesperson’s everyday caution: do not pretend certainty you do not have. Show a range. Explain the thin spots. Own the judgement.
Agent translation
When researchers ask for error margins and audits, they are asking for the same humility a good CMA already shows: “here’s what we know, here’s what we don’t, here’s why this range.”
Lending conversations vs listing conversations
AVMs often appear in bank and screening contexts as speed tools. Listing conversations have a different legal and relational shape: you are forming a professional opinion for a prospective client before an agency agreement. Importing lending culture (“the model says X”) into a vendor lounge is how labels get blurred.
Keep the products in their lanes. Use estimates as context if you must. Write appraisals as appraisals. Leave registered valuations to valuers.
A simple governance checklist for agency tech buys
Before you licence another estimate product or “AI pricing” add-on, ask vendors questions borrowed from the Auckland Uni agenda: What data sources feed the model? Can we see comps or drivers behind a figure? Is a range available, or only a mid-point? How is performance monitored against actual sales? How should outputs be labelled in client materials?
If a vendor cannot answer those, you still might use the tool internally as a research hint — but you should be twice as careful about what reaches a vendor. Your Code duties do not shrink because a sales deck was impressive.
Pair that vendor diligence with staff training on vocabulary. Most appraisal risk is social before it is statistical: someone wants a decisive number in a hurry, and the black box obliges.
Bottom line for busy licensees
Read the primary sources. Write the policy. Train the team. Label estimates honestly. Check every client-facing draft. Keep confidential files out of consumer prompts. Own the appraisal under rule 10.2. Tools can accelerate good practice; they cannot invent it after the fact when a complaint lands.
PropertyLM’s product bias is deliberate: evidence you can open, workflows that assume human sign-off, and language that stays accurate about what is an estimate, what is an appraisal, and what is still your professional call.
What “audit” looks like in a small agency
You do not need a university lab to borrow the spirit of AI auditing. Pick a handful of recent settled sales in your patch. Compare what your estimate tool showed close to listing time against the eventual unconditional price. Note the misses. Ask whether the misses cluster by property type, suburb fringe, or renovation status. That simple exercise trains staff to treat mid-points as hypotheses — and it gives managers a story for vendors who ask “how accurate is the computer?”
Document the exercise lightly: date, tool, sample size, direction of error. You are not claiming statistical significance. You are proving the office still thinks in evidence, not in screenshots.
Pair that with the Cheung and Yiu agenda in vendor conversations: ask for ranges, ask what data feeds the figure, and keep your written comps as the recommendation. Speed remains useful. Opacity does not become policy just because a dashboard looks modern.
— PropertyLM.
