Watercolor workspace with laptop, notebook, and overlapping circles labeled accuracy, trust, and decisions.
Studying the space between confidence and truth.

Field Guide

When AI Sounds Certain

Confidence is persuasive. A behavioral research study exploring how people evaluate AI, where trust comes from, and why confidence often matters more than correctness.

Illustration of a researcher observing how people evaluate AI outputs.
Observing decisions, not opinions

What We Found

People rarely evaluate AI objectively. Trust is a feeling, not a fact. it feels.

“We expected accuracy to predict trust. Confidence did. And prior domain knowledge played a significant role.”

Research synthesis

It is a given that AI hallucinates. Knowing this did not stop it from convincing users falsely. Knowing the risk and catching it in the moment are two different skills.

What predicted belief was not correctness, but tone, fluency, and confidence. The more certain AI sounded, the more likely people were to believe it—even when it was wrong. The real gap was not between correct and incorrect answers, but between actual accuracy and perceived confidence.

What we changed

The greatest risk is not inaccurate AI. It is confidently inaccurate AI.

Illustrated chart showing trust and accuracy, with a danger zone where confidence exceeds correctness.
Danger zone map

We Shipped two changes

The first was a source trace, a small expandable link under any AI generated answer that shows which documents or data the system pulled to build the response. We only surfaced it prominently in high stakes flows, the ones our research flagged as most likely to get rubber stamped.

The second was a deliberate pause, half a second longer than felt natural, before the system would let a user accept an AI generated answer in an unfamiliar domain. Not long enough to be annoying. Half a second is not a lot of friction. It was enough to change behavior.

When people accept confident answers without verification, design decisions become business decisions.

The Original Question

Can we predict when people will trust AI?

Most organizations evaluate AI by measuring accuracy.

What causes someone to believe an answer? Most teams building AI features test whether the answer is right. We tested whether people believe the answer is right, which is a different question entirely. You can build the most accurate system in the world and still lose the room if nobody trusts what it tells them. We wanted to know if we could predict when people would trust AI, and what factors influenced their decisions.

Illustration of Mhaire reviewing AI outputs with notes, highlighter, and coffee.
The question before the fieldwork

How We Got There

Behavior tells a more honest story than opinions.

We went looking for the moment trust actually forms. Not the moment someone fact checks an answer. The moment before that, when they decide whether checking is even necessary. Rather than asking what people trusted, we observed what they verified, ignored, accepted, and acted upon.

The research combined diary studies, structured evaluation tasks, blind accuracy assessments, behavioral follow-through, expert review, and synthesis. We also designed a controlled experiment to test whether confidence or accuracy was more persuasive.

Many outputs. One repeated pattern. Confidence moved faster than verification.

Illustrated diary study notes.
Diary study
Illustrated controlled AI evaluation task.
Controlled tasks
Illustrated blind accuracy ratings.
Blind ratings
Illustrated synthesis map showing accuracy, trust, and decisions.
Synthesis map

What I Learned

Technology changes quickly. Human psychology rarely does.

AI makes the tendency to mistake confidence for competence easy to observe. The instinct in product is almost always to remove friction. Here, removing it is how you get people to trust the wrong thing. Helping people make better decisions is to make it easier for people to tell when they should believe AI, and to catch it when they shouldn't.

Watercolor illustration of Mhaire reflecting with a laptop, notebook, coffee, and overlapping circles.
Reflecting after the synthesis