When teams see their first real LLM bill, the instinct is to blame the provider's pricing or hunt for a cheaper model. Sometimes that helps. But in our experience, most of the bill is architecture — and architecture is fixable without switching models.
Start with measurement. Break spend down by feature, by user, and by operation type: input tokens, output tokens, embeddings, reranking. Almost every team we work with discovers one surprise — usually that a single feature (often something with an agent loop or a chatty RAG pipeline) burns the majority of the budget. You can't optimize what you haven't attributed.
Then work the levers in order of effort. First, kill waste: cache repeated queries (system prompts and common questions don't need regenerating), dedupe retrieval (don't re-embed what hasn't changed), and set hard step budgets on agent loops — an unbounded loop is a blank check. Second, route by difficulty: a small model handles classification, extraction, and simple Q&A; reserve the flagship model for the genuinely hard reasoning. A router that sends 70% of traffic to a model at a tenth of the price is the single biggest cost win available to most teams. Third, shrink the context: retrieve less but better (reranking beats stuffing), summarize conversation history instead of replaying it, and strip the verbose system prompt nobody's audited since the prototype.
The lever people resist: charging the cost back to the feature. When each feature team sees its own token spend, optimization happens organically. When it's one shared bill, nobody owns it.
One warning: don't optimize cost before you've nailed quality. A cheap model producing garbage that users regenerate three times costs more than the expensive model that got it right once — in tokens and in trust. Get the evals green first, then squeeze.
Cost control is part of how we build production AI systems — designed for the bill you'll have at scale, not the bill you have in the demo.
Field note: The surprise-burn audit on a recent client: one "smart reply" feature — an agent loop drafting responses with three chained model calls — consumed 68% of the total token budget while serving 12% of users. The fix was a router: simple replies went to a small model in a single call, complex ones kept the full pipeline. Same user-facing quality on evals, 60% cost reduction overnight. The lesson generalizes: you don't have a model-cost problem, you have a routing problem. Measure per-feature spend first; the answer is usually one feature, and the fix is usually routing, caching, or a step budget.


