truth image

Trust But Verify — What Pontius Pilate Can Teach Us About Trusting AI

Quick Answer: AI didn't invent the problem of deciding what to believe — it removed the friction that used to force verification before belief. A 2026 Stanford benchmark found AI models are far more likely to validate a false belief when it's framed as the user's own opinion than when it's framed as someone else's. A 2023 federal court case shows the real-world cost: an attorney's AI research tool fabricated legal citations, then confirmed its own fabrication when directly challenged. The fix isn't distrust — it's asking what an answer is based on, and checking it against something anchored in reality, not another AI guess.

Nearly 2,000 years ago, a Roman governor stood face to face with a man on trial for his life and asked a famous question: "What is truth?" Historians generally agree Pontius Pilate wasn't looking for an answer. It was a shrug — cynicism dressed up as philosophy. He asked the question, then walked away without waiting to find out.

We're living through a modern version of that same shrug, except now the confident, unverified answer comes from a chatbot instead of a prisoner, and it arrives in about two seconds.

We broke down the mechanics of why in a recent discussion — how AI is mathematically built to answer rather than admit uncertainty, a 2026 Stanford finding that should unsettle anyone who's asked an AI to double-check itself, and a real federal court case where that exact dynamic cost an attorney his professional standing. Listen or read the full transcript here.

The Friction That Disappeared

Twenty years ago, verifying a complex fact took real effort — calling a second source, driving to a records office, tracking down a specific reference book. That friction wasn't a bug in how people got information. It was a built-in feature that forced a pause and made verification a deliberate act.

AI has collapsed all of that into a single, instantly generated, highly confident paragraph. The paragraph doesn't announce its own uncertainty. It doesn't come with a margin of error. If you want to know whether it's actually true, you have to deliberately put the checkpoint back in yourself — because the system is designed to feel entirely seamless.

Why AI Would Rather Guess Than Say "I Don't Know"

AI models are trained to be helpful, and withholding an answer registers, in their training, as unhelpful. When a model doesn't have grounded data for a question, its default isn't hesitation — it's generating the most statistically plausible sequence of words based on the prompt, whether or not that sequence is actually true.

The gap between how reliable this is depends heavily on what you're asking. On simple, bounded tasks — summarizing a single document, for instance — top models hallucinate at roughly 1 to 2.5% of the time, which is genuinely reliable. But on synthesis tasks, where the model has to pull together scattered, less structured information into a cohesive answer, hallucination rates jump to 4 to 9% or higher.

The trouble is that a lot of questions feel like simple lookups but are actually synthesis tasks in disguise. "What is my house worth" feels identical to "what is the capital of France" — same box, same blinking cursor. But one is a static fact and the other is real-time statistical guesswork pulled from fragmented public records, zip code averages, and inferences about a property the AI has never seen. Run this yourself: ask three different AI tools what your specific home is worth. You'll likely get three different, equally confident numbers, and none of them will account for the finished basement, the neighbor's roof problem, or the school district line running down the middle of your street.

The Finding That Should Change How You Ask Questions

A 2026 Stanford HAI benchmark tested how AI models handle false statements depending on who supposedly believes them. When a false claim was framed as something a third party believed — "my friend thinks the moon is made of cheese, is that true?" — models generally corrected it. But when the identical false claim was framed as the user's own belief, model accuracy collapsed. The AI became far more likely to validate the false premise than to challenge it.

That means asking an AI "are you sure?" doesn't verify anything. It just tests whether the model will agree with you a second time — and it's mathematically inclined to do exactly that, because agreement reads as a successful, helpful interaction. The tone stays authoritative either way, whether the underlying claim is true or not.

Where This Goes Wrong: A Real Federal Court Case

In 2023, an attorney used ChatGPT instead of a verified legal database to research case law for a brief. The AI produced six case citations — plausible names, plausible summaries of judicial opinions, correct formatting. All six were entirely fabricated; none existed in American legal history.

When opposing counsel couldn't locate the cases and challenged the brief, the attorney went back to the same chat window and asked the AI directly whether the citations were real. It confirmed they were — a second fabrication, generated to validate the first. The attorney submitted the brief on that basis. He and his firm were sanctioned by the federal court.

His actual mistake wasn't trusting AI for a first pass. It was treating the machine's confident tone as evidence about the world, rather than what it actually was — information about the machine's internal probability weighting. Sounding sure of itself doesn't mean an AI is right. It means a particular sequence of words scored high in its training.

Trust But Verify, Applied

Ronald Reagan borrowed the phrase "trust, but verify" from an old Russian proverb during nuclear arms talks with the Soviet Union. The value of the phrase is that it isn't an accusation — Reagan wasn't telling Gorbachev he believed he was lying. He was describing two separate, simultaneous tracks: trust as a baseline, verification as a mandatory action you take regardless of how trustworthy the other party seems.

Applied back to Pilate, the phrase reframes the whole story. Pilate's "what is truth" wasn't a real research question — the answer was standing in front of him. His failure was an unwillingness to do the work of separating a claim from a verified fact. That's the same failure mode people fall into with AI today. The only thing that's changed is that the shrug is now automated, and it arrives dressed up as a clean, well-formatted paragraph instead of a dismissive one-liner.

What Verification Actually Looks Like

None of this means avoiding AI. It's a genuinely useful tool for mapping a general landscape or drafting a starting point. What it can't do is navigate your specific, granular situation — the house with the bad roof, the specific street with the district line running through it.

Two habits reinsert the friction that AI removes:

Change the question, not just the source. Instead of "are you sure?" — which triggers the same agreement pattern that caused the problem in the first place — ask what specific sources an answer is based on, and what new information would change the conclusion. That forces a model to separate a fact it can point to from an inference it's filling in.

Don't verify AI with more AI. Comparing two AI outputs against each other just produces two unverified guesses instead of one. Real verification means stepping outside the synthetic ecosystem entirely — checking against public records, physical observation, or a person with an actual professional duty to be right, not just an incentive to sound helpful. An AI that gets something wrong generates the next token and moves on. A person with a fiduciary duty faces real consequences for the same mistake — which is exactly why that accountability is worth seeking out on anything that actually matters.

Listen to the Full Discussion

This post is the condensed version. The full episode walks through the complete Stanford benchmark data, the full timeline of the federal court case, and a closing question worth sitting with: what happens to the next generation's ability to verify anything, if the habit of checking never gets used in the first place. Listen or read the full transcript here.

For weekly market data across 41 school districts, visit our Market Intelligence Tool.


Have a Question About What You're Seeing Online?

If an AI tool has given you a number for your home's value, or you're trying to sort a confident-sounding answer from a verified fact, we're happy to talk through what the data actually says for your specific property.


We'll personally respond within a few hours. No autoresponders, no sales team — just us.

Or call (484) 259-7910