← All articles
AI Engineering6 min read

Shipping an LLM app is 20% model, 80% plumbing

Demos ship in a weekend. Production LLM apps are 80% plumbing: feature flags, live evals, latency budgets, and a rollback plan for intelligence.

Shipping an LLM app is 20% model, 80% plumbing

The demo is the easy part. A notebook, an API key, a clever prompt — you've got something impressive by Friday. Then someone asks "can users use it?" and the real project begins.

Production LLM systems are mostly infrastructure the demo never needed. Authentication and rate limiting per tenant. Request tracing that captures the full prompt, retrieved context, and tool calls for every single generation. PII redaction before anything hits a third-party API. Retry logic with backoff for the inevitable provider outages. Caching for the repeated queries that are quietly doubling your bill. None of this is AI research; it's the unglamorous engineering that determines whether the product survives contact with users.

Deployment strategy matters more than model choice. We ship behind feature flags so a new prompt or model version rolls out to 5% of traffic first, with evals running against the live slice. If quality dips, the flag flips back in seconds — no redeploy, no incident. Model versions are pinned and logged with every output, so when a provider silently updates their model (it happens), we can see exactly which responses came from which version.

Monitoring is its own discipline. Token usage per user, per feature, per day — with alerts, because a runaway agent loop can burn a month's budget overnight. Latency percentiles, not averages; p99 is where your users live. And quality signals: thumbs-down rates, human-review samples, regeneration rates. A spike in "user asked again differently" is your earliest warning that something degraded.

The part teams skip: a rollback plan for intelligence. If the new model version is worse at your specific task, you need to revert to the old one instantly — which means keeping old prompts, old configs, and old eval baselines around. "Just use the latest" is not a deployment strategy.

This plumbing-first mindset is how we deploy AI systems at QuantaloomAI — boring infrastructure, exciting reliability.

Field note: The overnight budget burn is real: a client's agent loop hit a retry storm against a flaky vendor API and spent a month's token budget before morning standup. What saved them the second time: per-run cost caps that halt the loop and page a human, plus per-feature spend alerts at 50/80/100% of budget. On the deployment side, the feature-flag rollback earned its keep when a provider's silent model update degraded extraction accuracy — the flag flipped in seconds, the pinned previous version served traffic, and nobody outside engineering noticed. Boring infrastructure, exciting reliability — it's not a slogan, it's the incident log.

Building something worth shipping?

We take on a small number of AI product engagements. Tell us what you are building — we reply within 48 hours.