Questions for an AI vendor whose demo went well
A good demo proves someone can build the happy path. These are the questions that tell you whether anyone has run the thing in production.
The demo will go well. It is built to. The data is clean, the question is one the team has asked a hundred times, and nothing in the room is running on a Tuesday afternoon with real traffic behind it.
That is not a criticism of the vendor. It is just the wrong thing to evaluate. The gap between a good demo and a system you can still operate in year two is where every AI procurement we have been asked to clean up went wrong. Four questions close most of it, and the useful part is less the answer than how fast it arrives.
Show me one you are still running
Not a case study, not a logo slide — a deployment that is live right now, with a date on when it went live. Then ask what has broken since, and what they changed because of it.
A team that has operated an AI system has incident stories, and they tell them easily, because those stories are the reason their current design looks the way it does. A team that has only built demos will answer this question in the future tense. Listen for the tense.
What happens when the model changes underneath you
The floor moves in this category in a way it does not in ordinary software. Providers deprecate versions, change defaults, and adjust behavior in ways that are improvements on average and regressions for your specific use case.
So: are you pinned to a version, and for how long? When a migration is forced, who re-runs the evaluation, and who pays for the work? A vendor with a real answer here has an evaluation set they can point at. A vendor without one will describe the new model as better and expect that to settle it.
Who is on call in year two
Ask what happens the first Saturday after the engagement ends and the output starts looking wrong. Someone has to notice, diagnose, and change something — usually a prompt, a retrieval index, or a threshold.
Whether that person works for you or for them is a decision, not a detail, and it should be priced either way. Ask it alongside the run cost, because the two are the same conversation: inference at real volume, the monitoring, and the salary of whoever reads it. Ask also what you keep if you leave — the prompts, the evaluation set, the labeled data, and in what format.
What does it do when it is wrong
Every one of these systems is wrong sometimes. The design question is what the wrong answer looks like from the outside. Does it arrive with the same confidence as a right one? Is there a path where the system declines instead of guessing? Does anything route a low-confidence case to a person, and does anyone find out that it happened?
Then ask which of those behaviors they will put in the contract. A vendor who will describe a failure mode out loud but not write it down has told you their own estimate of how often it happens. This is the buyer’s side of what an AI feature should report about itself — you cannot hold anyone to a behavior nobody measures.
None of these are gotchas. A good vendor enjoys them, because they are the questions their own engineers argued about, and answering them well is how they win against a prettier demo. The ones who stall are telling you something too.
If you want a second opinion on an evaluation already underway, that is part of how we do AI strategy work — including the cases where the honest recommendation is to buy nothing yet.
Written by Tommy Shrove, BluuAlpha Technologies.