← All articles
Production AI5 min read

Prompts are code. Version them like it.

Somewhere there's a prompt running in production that nobody versioned. Here's the discipline: prompts in git, evals in CI, composable prompt layers, and rollbacks one flag-flip away.

Prompts are code. Version them like it.

Somewhere in your company there's a prompt running in production that nobody can find, nobody versioned, and everybody's afraid to touch. It's the most load-bearing string in your product and it lives in a Slack thread.

Treat prompts as code and most production prompt problems evaporate. That means: prompts live in the repository, not in a dashboard. Changes go through review. Every version is tagged, and every model output is logged with the exact prompt version that produced it. When quality regresses, you diff the prompt like you'd diff anything else — because "someone tweaked the wording on Tuesday" is a root cause you should be able to find in minutes, not a mystery you investigate for days.

It also means prompts get tests. Not vibe checks — actual evals that run in CI. We keep a set of representative inputs per prompt with expected behaviors, and a prompt change that degrades them doesn't merge. This sounds heavy until the first time it catches a "small wording tweak" that silently broke the output format your parser depends on.

A few practices that pay off fast. Separate the prompt's roles into composable parts: the immutable system contract (what the model must always do), the task instructions, and the dynamic context. When something goes wrong you know which layer to fix. Parameterize ruthlessly — model name, temperature, max tokens are config, not prose. And write the failure instructions explicitly: what to output when context is missing, when tools fail, when the input is malformed. The default behavior of a confused model is confident nonsense; your prompt is the only thing standing between that and your users.

The mindset shift: a prompt in production is an API contract with a nondeterministic service. You wouldn't deploy an API integration without versioning, tests, and rollback. Don't do it with prompts either.

We run our AI product builds with prompts in git, evals in CI, and rollbacks one flag-flip away — because production prompts deserve production discipline.

Field note: the "Tuesday tweak" incident

Field note: The "Tuesday tweak" incident: someone edited a production prompt in a vendor dashboard to "improve tone." Extraction accuracy dropped 9% — discovered three weeks later during a routine eval run, after hundreds of malformed outputs had flowed downstream. With prompts in git, that edit would have been a pull request: diff visible, evals run in CI, reviewer asking "did you check the extraction suite?" The fix took ten minutes once found; finding it took three weeks. Version control for prompts isn't bureaucracy — it's the difference between a ten-minute revert and a three-week mystery.

Building something worth shipping?

We take on a small number of AI product engagements. Tell us what you are building — we reply within 48 hours.