Skip to main content

Engineered With AI

Reducing Hallucinations in Business-Critical Systems

A model that invents an answer is not malfunctioning. It is doing what it was built to do, which is produce plausible text. The engineering problem is to constrain where that behaviour can reach anything that matters, rather than to eliminate it.

That reframing is useful, because it moves the work from prompt wording toward system design, which is where the reliable gains are.

Ground the answer in retrieved source material

The single largest reduction comes from giving the model the relevant material at question time and instructing it to answer only from that. Fabrication rates fall sharply when the facts are in front of the model rather than being recalled from training.

Grounding is not a guarantee. A model handed five documents can still blend them into a conclusion none of them support, which is why the controls below matter alongside it.

Make refusal an acceptable answer

Most systems implicitly punish refusal: the prompt asks a question and any answer scores better than none. Explicitly permitting a statement that the information is not available, and testing that it does so, converts a class of silent errors into a visible gap somebody can act on.

  • Instruct the model to say when the source material does not cover the question.
  • Test the refusal path deliberately with questions you know are unanswerable.
  • Treat a refusal in production as a content gap to fill rather than as a system failure.

A system that says it does not know is more useful than one that is right most of the time, because the first one tells you which answers to check.

Sam Ortiz, Director of Engineering, Engineered With AI

Require citations, then verify them

Asking for a source alongside each claim gives a human something checkable and constrains the model toward the supplied material. Verify programmatically that the cited passage exists and actually contains the claim, because a fabricated citation is otherwise more convincing than a fabricated fact.

Constrain the output shape

Free text gives a model room to elaborate. A defined schema with specific fields removes much of that room, and it makes validation possible: a value outside an allowed range or a missing required field is caught before anything downstream consumes it.

Put the human where the cost sits

Full autonomy is a decision rather than an objective. Where an error is cheap and reversible, let it run. Where an error is expensive, hard to reverse or lands in front of a customer, a review step keeps most of the time saving and nearly all of the safety. Systems get switched off after one visible failure, regardless of how many correct outputs preceded it.

Measure it rather than assuming it improved

Build a test set of real questions with known correct answers, including ones the system should refuse. Run it after every prompt or model change. Without that, every adjustment is judged on a handful of examples somebody happened to try, and prompt changes that fix one case commonly break another.

For where these controls are applied most strictly, our support agent index covers clinical, financial and education settings where the boundary is enforced by design.

Design the interface to invite checking

How an answer is presented changes how carefully it is read. An answer shown alongside the passage it came from gets checked. The same answer presented alone, in a confident paragraph, does not. This is an interface decision rather than a model one, and it is among the cheapest safeguards available.

The same applies to uncertainty. Where a system has a confidence signal, surfacing it plainly is more useful than hiding it, because it tells the reader which answers deserve a second look. Systems that present every answer with identical assurance train their users to stop checking any of them.

Watch what happens after launch

Test sets cover what you thought to test. Production surfaces the questions nobody anticipated, and the useful signal is in the ones where users rephrased, escalated or abandoned. Logging those and reviewing them weekly for the first couple of months finds gaps far faster than expanding the test set in the abstract.

It also tells you whether the refusal path is working. A system that never refuses in production is not confidently correct; it is almost certainly answering things it should not.

Need answers you can stand behind?

We build systems that cite their source and refuse when they cannot, so the output is checkable rather than merely confident.

Share this :

Leave a Reply

Your email address will not be published. Required fields are marked *