Book a call

What Surrounds the Model: Engineering Generative AI Systems That Hold Up in Production

What Surrounds the Model: Engineering Generative AI Systems That Hold Up in Production

A demo is one line: a prompt goes in, the model answers, the answer goes out. It works on the first try, in front of the people who built it, on the questions they thought to ask.

Production is everything that has to surround that line once real users, real data and real money are involved. The model needs knowledge it was never trained on. Its behaviour has to be measured, because it changes between runs. Every failure has to be traceable to a cause. Costs have to come down. Users have to be heard without being taken literally. And some rules have to hold no matter what the model says.

This piece walks through the seven layers that do that work, what goes wrong in each, and the loop that turns them into a system that improves over time.

Requests flow along the top. Production signals flow along the bottom and come back as fixes.

The short version

  • Retrieval gives the model the right information at the moment it answers.
  • Fine-tuning changes how the model behaves. It is a poor way to teach it facts.
  • Evals tell you whether the system works, and whether it works reliably.
  • Observability tells you why a particular answer came out the way it did.
  • Caching makes the system faster and cheaper, and becomes a correctness risk when it is wrong.
  • Feedback shows you the failures your tests missed.
  • Guardrails keep the system inside rules it is not allowed to break.

The layers matter less on their own than in the loop that connects them: every failure found in production becomes a test, and every change has to pass the tests before it ships.

Retrieval: the right knowledge at the right moment

A model only knows what was in its training data. It has never seen your contracts, your support tickets or anything written after its cutoff. Retrieval-augmented generation (RAG) looks up relevant text when a question arrives and places it in the prompt.

It runs in two stages. Offline, documents are split into chunks, turned into embeddings and stored in an index. Online, each query is embedded, the closest chunks are retrieved and reranked, and the best few are passed to the model with the question.

Two stages, and the six places retrieval goes wrong.

Most RAG failures happen before the model sees anything:

  • The right document is never retrieved, because the query and the document use different words for the same thing.
  • The wrong document ranks first, and the model trusts the top result.
  • The content is stale, because the index was built last quarter.
  • The chunking is poor, so a table is split from its header or a clause from its exceptions.
  • Too much context is passed in, and the relevant passage is buried.
  • The model hallucinates anyway, even with the correct passage in front of it.

Only the last of these is a model problem. The rest are search problems, and they are fixed with better indexing, better ranking and fresher data. Retraining the model does not fix them.

Fine-tuning: changing how the model behaves

Fine-tuning continues training a pretrained model on examples of the task you care about. It is the right tool when the model needs to behave differently: classify into your categories, follow a strict output format, write in a consistent voice, or handle a specialised task it does poorly out of the box.

Parameter-efficient methods have made this cheap. LoRA freezes the original weights and trains small adapter matrices alongside them. In the original paper, this cut the trainable parameters for GPT-3 by roughly 10,000 times and the GPU memory needed by about three times, with quality comparable to full fine-tuning.

The common mistake is fine-tuning to add facts. Facts that change need to be updatable in minutes, attributable to a source and removable on request. Weights offer none of that. A useful rule:

  • The model doesn't know something: use retrieval.
  • It knows it but answers in the wrong shape, tone or format: improve the prompt first, then fine-tune.
  • Both: retrieval for the facts, fine-tuning for the behaviour.

Evals: unit tests for systems that don't repeat themselves

Ordinary software gives the same output for the same input. A model does not. Outputs vary between runs, quality is often a judgement call, and an agent may take a dozen steps before producing anything you can check. Testing it needs its own tools.

An evaluation harness runs each task several times, records what the system did and where it ended up, and grades the result. Grading the outcome matters more than grading the exact steps, because a capable system will often find a valid path nobody wrote down.

The same task runs k times. Each run is recorded and graded.

There are three kinds of grader, and most suites use all of them:

  • Code graders check exact matches, patterns, schemas or the final state of a database. They are fast, cheap and reproducible, and they break on any valid answer phrased differently.
  • Model graders use another model with a rubric or a side-by-side comparison. They are flexible and can judge open-ended answers, and they can be inconsistent or biased toward certain styles.
  • Human graders are the reference standard. They are slow and expensive, so their main job is to calibrate the other two.

Suites also serve two purposes. Capability evals are hard tasks the system cannot yet do well, and they measure progress. Regression evals cover behaviour that already works, should pass close to 100% of the time, and exist to catch anything that breaks it.

Solving a task once is not the same as solving it reliably

Because outputs vary, a single run says little. Two metrics answer two different questions. pass@k asks whether at least one of k attempts succeeds. pass^k asks whether all k succeed.

The same 80% success rate, read two ways.

Take a task the system gets right 80% of the time. Over five attempts, the chance that at least one succeeds is 99.97%. The chance that all five succeed is 33%. Over ten attempts, it falls to 11%. A system that "can do" a task may still fail at it for most users who try it repeatedly. For anything customer-facing, pass^k is usually the number that matters.

Building the first suite

Start small and start from reality. Anthropic's engineering guidance suggests 20 to 50 tasks drawn from failures you have already seen, which is enough to be useful and small enough to maintain. Each task should be unambiguous, solvable, reproducible and easy to grade.

Test both directions. For every "the system should do X", add a case where it should not. A suite that only checks for the presence of a behaviour rewards a system that does it everywhere.

Then put the suite in the path of every change. A new prompt, a new model version, a new retrieval setting or a new tool is a candidate. If regression evals drop against the current baseline, it does not ship until someone has traced why.

Every change runs the suite and is compared with the baseline before it ships.

Observability: why it worked, or why it didn't

Evals tell you whether the system works on the tasks you wrote down. A trace tells you what happened on a specific request, including the ones you never anticipated.

A useful trace records every stage of a request: what was retrieved and how it scored, which prompt and model version were used, what the model returned, which tools were called, how long each step took and what it cost. Latency is best tracked as p50, p95 and p99, because averages hide the slow tail that users actually notice. Every trace should carry the versions of the prompt, the model and the index, so that a change in behaviour can be tied to a change in the system.

The goal is a system you can debug by reading what happened, instead of guessing. We go deeper on this layer in When Nothing Errors: Observability for Non-Deterministic Agentic Systems.

Caching: faster and cheaper, until it is wrong

Model calls are usually the slowest and most expensive step, so production systems avoid repeating them. Caches sit at several layers, and each can answer a request before it reaches the next:

Each layer can answer before the next one runs.

  • Exact-match response cache. The same request has been answered before, so return the stored answer.
  • Semantic cache. A request close enough in meaning has been answered before. This saves more and is riskier, because "close enough" is a threshold someone has to choose.
  • Retrieval and embedding cache. Skip re-embedding the query or re-running the search.
  • Prompt (prefix) cache. Many model providers can reuse the processed form of a long, repeated prefix such as a system prompt or a large document, cutting latency and cost on the model call itself.

Deciding whether something can be cached is the easy part. The hard part is knowing when the cached value stops being true. A cached answer about a refund policy is wrong the day the policy changes. A cache keyed without the user's identity or permissions can serve one customer's answer to another, which turns a performance feature into a data leak. Every cache needs a clear rule for expiry and a clear rule for who it is allowed to serve.

Feedback: what your tests missed

Offline evals only cover the failures someone thought to write down. Users find the rest. They leave signals, both explicit and implicit: thumbs up or down, corrections, rephrased questions, retries, escalations to a human and workflows abandoned halfway.

These signals are evidence. They are not ground truth. A thumbs-down might mean the answer was wrong, retrieval missed the right document, the data did not exist, the interface was confusing, or the user misunderstood the question. Treating every negative signal as a model error leads to fixing the wrong thing.

The discipline is to triage each important failure to its root cause using the trace, fix the cause, and then add the case to the regression suite so the same failure cannot return unnoticed.

Guardrails: the boundaries that hold regardless

Some rules are too important to leave to a model's judgement.

Input guardrails check what comes in: attempts at prompt injection, personal data that should not be processed, malformed input, and whether the user is authenticated and authorised for what they are asking.

Output guardrails check what goes out: that it matches the expected schema, that it is safe, that it does not leak personal data, and that it follows policy and business rules.

The principle that matters most is simple: business-critical rules belong in deterministic code. If refunds above a limit need human approval, that limit belongs in the function that issues refunds. A prompt can make a violation less likely. Code makes it impossible. The model proposes an action, and the system decides whether it runs.

The model proposes a refund. A rule in code decides whether it executes or waits for a person.

The loop that makes it production

None of these layers is finished on its own. What turns them into a production system is the loop along the bottom of Figure 1.

A failure surfaces through feedback or monitoring. The trace shows where it started. The cause is fixed in whichever layer it lives in: retrieval, prompt, model, cache or guardrail. The case is added to the regression suite. The next change has to pass that suite before it ships. Over time, the set of failures the system can repeat gets smaller, and the evidence that it works gets larger.

Where to start

Teams rarely build all seven layers at once. If we had to order them for a team moving a prototype toward production, it would be roughly this:

  1. Trace from the first day. It costs little early and is painful to add later, and every other layer depends on it.
  2. Write the first 20 to 50 evals from failures you have already seen, and run them on every change.
  3. Put hard guardrails on anything that touches money, permissions or personal data.
  4. Improve retrieval quality, using the traces to see which failures are search failures.
  5. Add caching once answers are correct, with explicit expiry and scoping rules.
  6. Fine-tune last, and only for behaviour that prompting cannot fix.

Why this matters to us

Cygnux Labs works on making AI systems verifiable: records of what a system did that can be checked later, and methods for finding out why it acted. Each layer in this piece produces part of that evidence. Evals show what a system can be trusted to do, traces show what it actually did, and guardrails define what it was never allowed to do. As these systems take on more consequential work, that evidence stops being an engineering convenience and becomes the basis for trusting them at all.

← Back to Research