Here's a pattern we've seen in almost every AI project that comes to us mid-crisis: the team shipped a demo that impressed stakeholders, users arrived, quality drifted, and nobody can say when it started or why. There's no baseline. There's no "it used to work" to compare against, because "used to work" was never measured.
Evals are the fix, and they're less glamorous than they sound. An eval is just a set of inputs with expected outputs, run automatically every time you change the prompt, the model, or the retrieval layer. Start embarrassingly small: 20 examples pulled from real user sessions, each with a human-judged good answer. That's a weekend of work and it will save you months.
What we actually run in production has three layers. Golden-set evals run on every change — fast, deterministic checks on the cases that must never regress. Sampled human review runs weekly: a human grades a random slice of real traffic, which catches the drift that golden sets miss. And we log everything with the full trace — prompt, retrieved context, tool calls, final output — so when something looks wrong we can replay it exactly.
The trace logging matters more than people expect. LLM behavior is non-deterministic; without traces, debugging is astrology. With traces, it's engineering: you find the exact step where the context went stale or the tool returned junk, and you fix that step.
One more thing teams get wrong: they eval the model but not the system. A RAG app's quality is mostly retrieval quality. Eval your retriever separately — precision and recall on your actual corpus — before you blame the model for bad answers. Half the "model is hallucinating" tickets we've investigated were retrieval failures wearing a model costume.
Observability isn't overhead. It's the difference between "our AI feels worse lately" and "retrieval recall dropped 12% after Tuesday's index rebuild." One of those is actionable.
This is why every AI product we ship has evals and tracing baked in from week one, not bolted on after the first incident.
Field note: The "Tuesday's index rebuild" example is real — a client's RAG answers degraded and nobody noticed for two weeks because there was no baseline. When we finally built the golden set (47 questions from actual support tickets), we found recall had dropped 12% after an embedding config change nobody had flagged as risky. With evals in CI, that change would have been caught in the pull request. The total cost of building those evals: about three engineer-days. The cost of the two silent weeks: a support backlog and a dented client relationship. Evals are the cheapest insurance in AI engineering, and almost nobody buys it until after the first fire.
