"Should we build or buy?" is the wrong question for AI. The right question is: "What's our differentiating layer, and what's plumbing?" Buy the plumbing, build the differentiation — and be brutally honest about which is which.
Buy when the capability is generic and the vendor's scale beats yours: embeddings, transcription, moderation, base models. You will not build a better speech-to-text model than the dedicated providers, and trying is an excellent way to burn a year. Buy when speed matters more than control: a pilot that needs to run next month should ride vendor APIs, not a training cluster.
Build when the capability touches your proprietary data or your core workflow. A RAG system over your internal knowledge, an agent wired into your specific toolchain, a model fine-tuned on your domain's language — that's where the moat is, and no vendor sells it off the shelf because it's yours. Build when the economics demand it: at high volume, per-token API costs eventually exceed the cost of running your own inference, and the crossover point arrives sooner than finance expects.
The trap is the middle: "customize the vendor platform." Heavily customized SaaS AI tools give you the worst of both — vendor lock-in plus maintenance burden, without real ownership. If you need deep customization, that's a build signal wearing a buy costume.
There's also a sequencing answer most teams miss: buy first, build later. Launch the pilot on vendor APIs in weeks, learn what actually matters from real usage, then selectively bring the expensive or differentiating pieces in-house. The teams that try to build everything on day one spend a year on infrastructure before learning anything about their users.
We help clients make this call explicitly in every engagement — our AI consulting work starts with mapping what's commodity and what's moat, because the answer determines the entire architecture.
Field note: five weeks vs. eight months
Field note: A recent engagement made the sequencing concrete: transcription and embeddings bought from vendors (weeks to integrate, best-in-class quality), the RAG layer over internal policy docs built custom (the moat — nobody sells their policy corpus), and inference still on vendor APIs with a planned migration to self-hosted at a defined volume threshold. Total time to pilot: five weeks. The build-everything alternative was quoted at eight months. Buying the plumbing didn't just save time — it bought learning: real usage data told us exactly which pieces were worth bringing in-house and which never would be.


