Model comparison posts are easy to write and hard to trust. Most of them test the models on puzzles or coding challenges that do not resemble the work I actually do. This post takes a different approach: I gave Gemini, ChatGPT, and Claude the same realistic product task, used the same evaluation criteria, and recorded what each model did well and where it failed.
The task was deliberately ordinary. That is the point. Dramatic benchmark wins often disappear when the input is messy interview notes and the required output is something a product team would actually use. I wanted to see the models under those conditions.
The Task
I took a set of anonymized customer interview notes from a recent discovery cycle and asked each model to:
Identify the top three recurring pain points
Group supporting quotes under each pain point
Suggest one product opportunity for each pain point, grounded only in the notes
Flag any place where the notes were too thin to support a claim
The same cleaned notes, the same prompt structure, and the same temperature settings were used for all three models. I ran each model three times and kept the median-quality output for comparison. The notes themselves were deliberately imperfect: some themes appeared only once, some quotes were ambiguous, and a few important context points were buried in longer paragraphs. That messiness is closer to real discovery work than a clean synthetic dataset.
I also forbade the models from using any external knowledge. The instruction was explicit: base every claim on the provided notes only. This constraint made fidelity differences more visible.

Evaluation Criteria
I scored the outputs on four dimensions:
Fidelity to source (did it invent claims?)
Coverage of real themes (did it miss obvious patterns?)
Usefulness of opportunities (were they actionable and grounded?)
Consistency across the three runs
No model was perfect. The differences were informative. I deliberately avoided scoring “writing quality” or “insightfulness” as primary dimensions because those are harder to define consistently and easier to be swayed by fluent language. The four criteria above are closer to what actually matters when the output will influence a product decision.
Dimension | Gemini (median) | ChatGPT (median) | Claude (median) |
|---|---|---|---|
Fidelity to source | High | Medium-High | High |
Coverage of themes | Good | Strong | Strong |
Grounded opportunities | Good | Variable | Strong |
Run-to-run consistency | High | Medium | High |

What Stood Out
Gemini was the most consistent across runs and rarely invented details that were not in the notes. Its opportunity suggestions were conservative and usually stayed close to the source material. When the notes were thin, it was more likely to say so. That honesty about thin evidence is valuable when the output will be used to justify roadmap decisions.
ChatGPT produced fluent, readable output and sometimes surfaced themes I had under-weighted. It was also the most willing to stretch beyond the notes when proposing opportunities, which is useful for brainstorming and risky for decision documents. In one run it proposed an opportunity that sounded plausible but rested on a single ambiguous quote. That is exactly the kind of subtle overreach that is easy to miss when reading quickly.
Claude was strong on both fidelity and the quality of grounded opportunities. Its explanations of why a theme mattered were often the clearest of the three. It was also the most verbose, which required extra editing for internal use. The extra length sometimes included helpful nuance and sometimes included restatement that did not add signal.
What I Took Away for Real Work
For tasks that must stay tightly grounded in source material, I currently reach for Gemini or Claude first. For exploratory brainstorming where some extrapolation is acceptable, ChatGPT remains useful. The more important lesson is that the evaluation criteria matter more than the model ranking. A different task with different stakes would have produced a different ordering.
I also learned that run-to-run consistency is underrated. A model that produces a brilliant analysis once and a mediocre one the next time is harder to operationalize than a model that produces good-enough results reliably. Consistency affects how much human review you need to budget.
Finally, the act of writing explicit evaluation criteria before looking at any output changed the quality of the comparison. Without those criteria it is too easy to be impressed by fluent language and miss systematic gaps. I now write the scoring rubric first for any model evaluation that will influence a real decision.
I tried it first. You can run the same style of comparison on your own tasks. The value is in the disciplined setup, not in declaring a permanent winner. The models will keep changing. The habit of testing them on your actual work will remain useful.