Skip to content
Determos
7 min read

Deterministic AI: how to take AI to production without hallucinations

Most AI projects impress in a demo and fall apart in daily use. You don't fix that by asking the model not to invent things: you fix it by closing the paths where it can.

Written byGuillermo GómezWorks with Artificial Intelligence at Airbus

Almost every company has already tried AI. Very few have it genuinely running inside their operation: an estimated 2% have moved from «trialling» to production. The reason is almost always the same. A language model is brilliant in a meeting and fragile in the real world: it hallucinates, forgets context, ignores your brand and gives different answers to the same question. For quoting a price, booking an appointment or generating a contract, that isn't help — it's a risk.

What deterministic AI is

It's worth starting with what it isn't, because the name misleads. Nobody makes a language model deterministic: an LLM is probabilistic by construction and will stay that way. What you can make deterministic is everything around it — what information reaches the model, what it's allowed to do with it, and what counts as a valid answer. That layer is what we build, and that's where the name comes from.

Put another way: the job isn't getting the model to be right every time, it's closing off, one by one, the paths where it could invent something. The fewer paths left open, the less room it has to hallucinate — and the less the underlying model matters. A small, cheap model with the space properly narrowed performs like a frontier one, without its cost or its unpredictability.

Why generic AI fails in production

Generic chatbots fail at the same three points every time, and they are exactly the ones a business can't afford:

  • You can't rely on it: ask the same thing twice and you get two answers.
  • It doesn't sound like you: it talks like anyone, or worse, says things your company never would.
  • It doesn't do the work: chatting doesn't sell; the value is booking, updating the CRM or generating the document.

Where hallucinations come from

A model hallucinates when it's left with a gap to fill. Hand it your whole catalogue, your pricing history and a customer's question, and expect it to work things out, and you've left it plenty of gaps: which price applies here, whether that discount is still live, whether the product even exists any more. Every gap is a path where it can invent.

The usual approach tries to plug them by writing instructions: «don't make up prices», «if you don't know, say so». Those are explicit rules, inside the prompt, and therefore suggestions: the model interprets them, loses track of them as the conversation runs long, and someone can talk it out of them.

The pillar: don't ask it not to invent — remove the gap

We do the opposite: the rules aren't asked of the model, they're applied before it's ever called, and they never appear anywhere it could interpret them. The algorithm looks up the price, checks the product is active, and hands the model a single figure with a single job: write it in your tone. There's no price to invent, because the model has only ever seen one.

Don't hand it everything and hope it picks well. Hand it only what you've already decided: what never reaches the model can't come out wrong.

That's the underlying shift, and it's what separates a pretty demo from a system you can trust. The usual approach gives the AI access to your calendar, CRM and data, hands it the rules and hopes it follows them. Flip it and the AI no longer interprets your rules: it can't even see what's out of scope, and it can't run any action you haven't granted. It's hardened by design against attempts to «talk it round», because there's nothing left to talk round.

How the paths get closed, step by step

  1. Isolate data per client: one company's information can't surface in another's answer, because it never sits in front of the model.
  2. Retrieve before answering (RAG): instead of your whole knowledge base, the model gets the handful of paragraphs this question needs. What isn't in them, it can't assert.
  3. Verify against the real system: availability, stock and data are checked in your calendar or ERP before anything is confirmed. The AI proposes; the system confirms.
  4. Constrain the output: tone, format and limits are enforced outside the model, and what doesn't fit never reaches the customer.
  5. Measure where it slips out: every real failure points at a path still left open. You fix it there, in the system, not by rewriting the prompt.

What it costs and how long it takes

Because the model only receives what's needed, the job left to it is small — and a small job doesn't need a giant model. Resolving 1,000 conversations with a «big» AI costs about $35–40; with this engineering, $4–8: up to 10× cheaper, with equal or better reliability. And there's no need for an endless project: you start with the process that hurts most and go live from there.

Data sovereignty: an advantage, not a constraint

European rules make where your data lives genuinely matter. With deterministic AI, information stays on European servers or runs 100% inside your company (on-premise). What slows the big tech firms in Europe — GDPR and the EU AI Act — protects you and becomes a selling point.

This isn't theory: companies are already generating tender contracts 5× faster, handling 4× more queries and cutting order errors from 30% to 1%. They're all in case studies. The difference isn't the model — it's the engineering around it.

Frequently asked questions

What is deterministic AI?

It isn't a model that stops being probabilistic — no such thing exists. It's the engineering layer around the model — data retrieval, rules and verification — and that layer is deterministic: it decides what information reaches the model, what it can do with it, and what counts as an answer. The model still improvises, but inside a very small space you have already narrowed.

How are AI hallucinations prevented?

By closing off the paths where it can invent, rather than asking it not to. A model hallucinates when it's left with a gap to fill: give it your whole catalogue and a question, and it has to decide which price applies. Give it one price, already looked up and verified, and there's nothing left to invent. The rules are applied before the model is called, not inside the prompt.

Isn't it enough to tell it in the prompt not to make things up?

No. An instruction in the prompt is a suggestion: the model interprets it, loses track of it as the conversation runs long, and someone can talk it out of it. A rule applied outside the model — filtering what it sees, verifying before confirming, limiting what it can execute — isn't interpreted: it always holds.

Is it cheaper than using a large model directly?

Yes. Because the model only receives the information it needs, a small, cheap model performs like a large one. Resolving 1,000 conversations can drop from $35–40 with generic AI to $4–8 with this engineering: up to 10× cheaper, with equal or better reliability.

Does my data leave Europe?

It doesn't have to. Information can be hosted on European servers or run 100% inside your company (on-premise), with nothing leaving your network. The approach is designed to comply with GDPR and the EU AI Act from the start.

Want to take your AI to production?

Let's talk