Demos are optimists. Production is a realist. This post is a post-mortem of a Gemini feature that looked excellent in controlled tests and then failed when real traffic, real data, and real latency constraints arrived. The goal is not to blame the model. The goal is to make the failure modes visible so the next feature avoids the same traps.
I am writing this while the memory is still fresh. The feature is now stable, but the path from demo to stability taught me more about Gemini in production than any of the successful launches.
What the Feature Was Supposed to Do
We needed a lightweight classifier that could read short customer messages and route them to one of six internal queues. The demo used a clean set of twenty examples per class. Gemini Flash returned the correct label on nineteen of the twenty. Latency was under 400 ms. Cost was negligible. We shipped a prototype.

What Broke in Production
Three problems appeared within the first week:
Distribution shift – Real messages contained typos, mixed languages, and implied context that the clean demo set never had. Accuracy dropped from ~95 % to the low 70s. Messages that mixed English with another language were especially problematic; the model often defaulted to a catch-all category.
Latency tail – p95 latency climbed above 1.2 seconds when the prompt included a short conversation history. The interactive UI started to feel sluggish. Users began refreshing the page, which created duplicate tickets and made the accuracy problem harder to measure.
Silent schema drift – Occasionally the model returned a label that was not in the allowed set. The downstream router treated it as a hard failure. Because these events were rare, they did not show up in the original accuracy calculation; they only appeared as intermittent 500s in the logs.
None of these appeared in the original twenty-example test. That is the most important sentence in this post. The demo was not lying; it was simply answering a different question from the one production asked.
Root Causes We Identified
Problem | Root cause | Why the demo missed it |
|---|---|---|
Accuracy drop | Demo data too clean | No realistic noise or edge cases |
Latency tail | History tokens added in production | Demo used single-turn messages only |
Invalid labels | No output validation | Demo outputs were manually inspected |
The common thread was that the demo environment was friendlier than production in every dimension that mattered. We had optimized for the happy path and then been surprised when the unhappy path appeared. In hindsight the surprises were predictable; we simply had not allocated time to look for them before shipping.
I now treat any demo accuracy number as provisional until it has been measured against a set that includes the messy cases. That single change in attitude would have caught two of the three problems before they reached users.

What We Changed
We kept the model but changed the surrounding system:
Expanded the evaluation set to 300 real messages, including the messy ones
Added a strict allow-list check on every label before routing
Moved history summarization to a cheaper, faster step so the classifier prompt stayed short
Introduced a confidence threshold; low-confidence messages went to a human queue instead of being force-routed
After these changes the feature became stable enough to stay in production. The model itself was not the primary problem; the incomplete test harness was.
Lessons I Now Apply to Every New Gemini Feature
Build the evaluation set from production-like data before writing the prompt
Measure p95 latency with the actual prompt length you will use
Validate every structured output against an explicit schema or allow-list
Budget for a human fallback path from day one
Treat the demo accuracy number as an upper bound, not a forecast
I also changed how I present demos internally. I now include at least one slide that shows a realistic failure case and the mitigation we already have in place. That single slide has prevented more over-confidence than any amount of cautionary language.
The deeper lesson is that Gemini’s behavior is only one variable. The shape of the input distribution, the length of the prompt, the strictness of the output contract, and the presence of a human fallback all determine whether a feature survives contact with production. Optimizing only the model call is rarely enough.
I tried it first. The feature that worked in the demo taught me more about production than the features that worked on the first try. You can borrow the failure checklist even if your use case is different. The next feature I ship will still start with a clean demo—but the evaluation set will already contain the messy cases.
I also now treat every new Gemini feature as having two launch criteria: demo quality and production-readiness quality. Passing the first is necessary but not sufficient. The second requires the expanded evaluation set, the validation layer, the latency measurement under realistic prompt length, and the human fallback path. Building those into the definition of “done” has prevented more post-launch surprises than any single testing technique.
The broader lesson is that Gemini’s behavior is only one variable. The shape of the input distribution, the length of the prompt, the strictness of the output contract, and the presence of a human fallback all determine whether a feature survives contact with production. Optimizing only the model call is rarely enough.