Why AI projects fail after the demo
2 min read
Updated
- Production
- Evaluation
- Deployment
AI projects rarely fail at the demo. They fail months later in production, for reasons that were visible on day one: no evaluations, no owner, and releases that can't be undone.
The demo is rarely the risk. AI projects fail later, in production, because evaluation, ownership and safe releases were left for later, and later arrived with users watching.
The model call is treated as the system
The model is one box in the diagram. What decides success surrounds it: where the data comes from, queues and retries, logs, review screens, cost limits and deployment. A budget that covers only the box ships only the box.
Evaluation starts after the arguments
Without evaluations, every prompt change is an opinion and every regression stays invisible until a user reports it. An evaluation set is cheapest to write on the day the prototype first works, from real cases.
For the support agent I built, the evaluation set is a reviewed catalog of real tickets with their captured data. An independent model judges drafts against expected outcomes the drafting model never sees. Promoting a prompt change calls for two full passing runs on matching inputs; skipping that shows as a warning, never as a pass. When operators edit a draft, the edit can become a new case, so the catalog grows from real work.
Nobody owns the failure
A production AI workflow needs a named owner for wrong answers, provider outages, duplicate work and escalations. If the answer to "who gets alerted?" is a shrug, it isn't ready.
Releases can't be undone
A prompt, model or configuration change is a release like any other, and it can be wrong. Teams that test model quality carefully still push those changes with no health check and no rollback path.
A health gate proves a new version healthy before it takes traffic, and rolls back on its own when it isn't. Configuration changes behind an AI feature, such as a new prompt schema or a provider-routing update, are migrations and need the same discipline as a database schema change. Prompts stored as immutable versions make rollback a pointer change.
Fast delivery depends on all of this
Most of my code is written by AI coding agents. That only works because tests, evaluations, review and a deploy that can roll back decide what ships. With those in place, shipping several times a day is routine; without them, every release is a gamble however carefully it was read.
