Ordinary software is deterministic. Give it the same input and it does the same thing, which is why you can write a test once and trust it forever. A language model is not like that. The same prompt can produce a different answer tomorrow, a slightly different input can produce a very different answer, and the model will say something wrong with complete confidence. Owners feel this instinctively, and it is the real reason so many AI pilots never reach production. Nobody could say with a straight face that it was tested.
Here is how we do it, for every automation, assistant, or agent we put into a client’s business.
Start with what “correct” means
Before any model is involved, we write down what a correct result is for the task, in terms the business already uses. For an intake automation, correct means the right fields extracted from the right documents, and unclear ones flagged rather than guessed. For a drafting assistant, correct means a draft the person would have sent with light edits. If we cannot define correct, we do not build it yet; we go back to the workflow and find the piece where we can.
Build an evaluation set from real cases
Then we collect real examples: fifty to a few hundred actual documents, emails, records, or requests from your business, with the correct answer for each, agreed with the person who does the job today. This set is the test. Every change to the prompt, the model, or the surrounding code runs against it, and we track how many it gets right, how many it gets wrong, and how it fails. When a model update arrives, the set tells us in an hour whether anything broke.
The set also includes the ugly cases on purpose: the handwritten note, the email that asks three things, the record with a missing field. Those are where an automation earns its keep or embarrasses you.
Put a person where the cost of a mistake is real
Not every step needs approval, and an automation that asks a human about everything is not an automation. So we sort the decisions by what a wrong one costs. Routine, low-stakes, easily reversed: the system acts and logs it. Consequential, customer-facing, or hard to undo: the system prepares and a person approves. The line is drawn with you, in writing, before launch, and it can move as trust builds. A first version often approves more than it needs to; the evaluation data tells us when it is safe to let go.
Give the model less room to be wrong
Most of the reliability in a production system comes from what is around the model, not the model:
- Narrow tasks. One job per step. A model asked to extract a date is far more reliable than one asked to “handle the email.”
- Structured output. The model fills a defined shape, and code validates it before anything downstream sees it. A missing or malformed field stops the process instead of propagating.
- Grounding. For anything factual, the model reads from your data and cites what it read, rather than answering from memory.
- Rules where rules work. If a decision can be written as a rule, it is a rule, not a prompt. Cheaper, testable, and it never has a bad day.
- Limits. Budgets on retries, on spend, on how many records a run can touch. A runaway process is capped by design.
Log everything
Every run records what came in, what the model produced, what the code did with it, and who approved what. This is how a problem gets diagnosed in minutes rather than argued about for a week, how the evaluation set grows from production cases, and how you can answer a customer, an auditor, or your own curiosity about why the system did what it did.
Test the whole thing, not just the model
The model is one component. The integration with your CRM, the queue, the retry logic, the email that goes out: those get the same tests any software gets. We run the automation end to end in a staging environment against copies of real data, then shadow it in production, where it produces results that a person compares to what they would have done before it is allowed to act.
Launch small and measure
The first weeks in production run at reduced scope, one team or one location, with the KPI we agreed on tracked from day one. We look at the exceptions every week, add them to the evaluation set, and widen the scope when the numbers hold. An automation that saves twelve hours a week is not proven by a demo. It is proven by the timesheet.
What this costs you
Less than it sounds. The evaluation set takes a few hours of the person who knows the job. The checkpoints are decisions you would want to make anyway. What it buys is the difference between a pilot that everyone quietly stops using and an automation that is still running, unremarked, two years later, which is the only kind worth paying for.