Hallucination is a design problem, not a model problem
Waiting for a model that never invents anything is not a plan. Building systems that assume it will is.
- July 8, 2026
- Published
- 9 min
- Read time
- AI
- Category
On This Page
The question is not whether it will be wrong
Every conversation about deploying AI eventually arrives at hallucination, and it usually arrives as a blocking objection: what if it makes something up. It is the right concern and usually the wrong framing.
Models produce confident output regardless of whether they have the grounds for it. That is not a defect to be patched, it is a property of how they work, and any system design that assumes it will be solved by the next release is building on sand.
The useful question is the same one you would ask about a new employee: not whether they will ever be wrong, but what happens when they are.
Ground the output in something checkable
Most fabrication happens when a model is asked to produce specifics it was never given. Ask it to write about a service in a town it knows nothing particular about and it will produce something plausible, because plausible is all it has.
Retrieval fixes a large share of this. Give the model the source material and require it to work from that, and the failure mode shifts from inventing facts to misreading supplied ones, which is a much easier problem to detect.
It also makes verification possible. Output grounded in a document can be checked against that document. Output from nowhere can only be checked by someone who already knows the answer.
Separate the checks
A single reviewer asked to check accuracy, tone, formatting, and completeness at once will do all four badly. That applies to automated checks exactly as it applies to people.
Splitting them works better. One pass asks only whether anything was invented and whether the claims match the source. A second asks only whether the structure and format are right. Each check is simpler, and a simple check is a reliable check.
When a check fails, the specific failure is also more actionable. Knowing the format was wrong tells you what to fix. Knowing it did not pass tells you nothing.
Put the human where the stakes are
Review is not free, so spending it uniformly is waste. The design question is where an error is expensive and irreversible, and that is where a person belongs.
A wrong word in an internal summary costs nothing. A wrong figure in a client deliverable costs trust. A wrong instruction in a regulated document costs considerably more. The same system can reasonably auto-publish the first and require sign-off on the third.
This is also what makes AI acceptable to sceptical teams. People object far less when they can see exactly which decisions the system is allowed to make alone.
Measure it before you trust it
Build an evaluation set from your own real examples before anything goes live. Fifty representative cases with known correct answers turns accuracy from an impression into a number.
Then set the bar against the process being replaced, not against perfection. Manual processes have error rates too, and they are usually higher than anyone assumes because nobody measures them.
Keep running it. A system that passed in March can quietly degrade when a model version changes underneath it, and without a standing evaluation you will discover that from a complaint.
Log enough to diagnose
When something does go wrong, you need to know what the system saw, what it was asked, and what it produced. Without that, a wrong answer is a mystery and the only available response is to lose confidence in the whole thing.
Good logs turn a failure into a fix. They are also what lets you spot that a class of input is consistently producing bad output, which is usually a scoping problem you can solve by routing that class to a human.
ReynoldsBuilt
AI, automation, and custom software
We audit an entire operation before building anything, then build what the business actually needs. Everything here comes out of real engagements.
About the studioKeep reading
More from the blog
Written for the person who has to make the call, not the person writing the spec.
Strategy · 8 min
The spreadsheet your team built is the best spec you have
Every business has a shadow spreadsheet holding the operation together. Most software projects throw it away. That is a mistake, and an expensive one.
ReadAutomation · 7 min
The worst automation failure is the one nobody notices
An automation that breaks loudly gets fixed the same day. One that fails quietly can corrupt six months of data before anyone asks a question.
ReadStrategy · 6 min
Before you buy more software, audit what you already own
Licensed features sitting unconfigured are the most common finding in an assessment, and the cheapest thing to act on.
ReadFind out what should actually be built.
Start with an assessment. We walk your business end to end and show you where automation and AI pay off, ranked by what they are worth.