Back to Blog Part 2 of 4 · The Missing Function
AI Strategy · 7 min read · Published August 31, 2026

Owning the AI Quality Bar vs. Owning the Proof

Matt Genovese
Matt Genovese
Founder & Product Strategy Lead
An analog precision gauge with a red line marked GOOD ENOUGH on its dial, the needle blurred in motion as it sweeps back and forth across the line without settling

Deciding whether a feature is good enough to ship has always been product’s call. With AI, answering it takes evidence most product teams have no way to produce, and that difference has very little to do with anyone being unreasonable.

Somewhere in the life of every feature, someone stops and asks whether it is good enough to ship. That question has always belonged to product, and I would argue it is about the most product question there is, because answering it means weighing what a user genuinely needs against what the thing will cost to build and to keep running, and then deciding, with some judgment and a little nerve, where the line falls.

The first article in this series was about a question that comes before that one, whether an item on the feature list is routine work or an open research problem, and how unreliable the usual instincts have become at telling those apart. This one picks up after that has been settled.

Put an AI feature in front of the good-enough question, though, and something breaks. Development asks for the requirements, the way any team would, and then a second request lands with no precedent: to know whether the feature clears the bar, someone has to assemble a ground truth dataset, from tens to hundreds of real examples with the correct answers already attached to each one. Without that, there is no honest way to test a good response from a mediocre-to-incorrect one, and so no way to answer the qualitative question you have always answered. And on many product teams, that job belongs to no one. With no data scientist in the room, progress can stall out.

Product really does own AI quality

It is worth conceding a point to the AI team properly, rather than through gritted teeth, because the literature is genuinely on their side. The field has spent a couple of years converging on the idea that product owns AI quality, and it does not phrase this timidly: Lovelaice calls quality “a product problem wearing an engineering costume,” and Architecture of Proof says an assertion about quality “is not an engineering artifact, it is the PM’s primary instrument of accountability.” I think that is right, and not only for AI. Great has always been the enemy of good enough, and knowing where that line falls, the point at which the objective is met without overburdening the build, the upkeep, or the cost to run it, is precisely the judgment product exists to make.

Implementation certainty, and what replaced it

What changed is what used to happen the moment after quality was defined. Once product said what good enough meant, engineering built it, and that it could be built was never really the question; the only open questions were how long it would take and what it would cost, and the thing would go on working for years afterward. I have started calling that implementation certainty, because naming it makes it easier to see how much of the product manager’s job had been resting on its two promises:

  1. Whatever product defined, it could be built
  2. Once it was built, it would keep working

Put a language model in the middle of the feature and neither promise holds. Whether it works at launch is no longer a given, and whether it keeps working is even less of one, so proving the bar has been met stops being a matter of writing test cases and becomes a matter of probability. So a team can sort the routine features from the research problems and still be stuck. Knowing a feature can be built is one thing; proving it works well enough is another.

Engineering could always build it; that certainty is exactly what AI takes away.

Hardware is the right analogy here. A sensor in the field fails in ways that have nothing to do with a manufacturing defect, so reliability is handled statistically and the exceptions get solved before it ships. An AI feature is more like that sensor than like a web form, with one added twist: the component you qualified can change underneath you the week the vendor ships a new model.

Two different jobs, and only one belongs to product

This is where the standoff actually lives, and both sides tend to argue right past it. Owning the definition of quality and manufacturing the evidence that it has been met are two different jobs. Say the feature is an AI router that sends each incoming support ticket to the right team. A product manager can state, with complete authority and no help from anyone, where the bar belongs: getting 8 of 10 tickets to the right team is not good enough, and the number to clear is 95%.

Building the ground truth dataset that would actually measure which of those the router hits is a separate discipline, and it is the foundation the rest rests on, since the evaluation set that later watches the model for drift in production is built from it. FutureAGI describes the people who assemble ground truth working from production traces, customer escalations, and adversarial seeds, and notes, tellingly, that the role “rarely overlaps cleanly with ML engineer or backend engineer responsibilities.” If it does not overlap cleanly with those two, it is certainly not going to overlap with a product manager’s Thursday afternoon.

Two gaps in that room, then, and only the first belongs to product.

How many dimensions the input varies along

The genuinely useful question, then, is not how many examples the ground truth dataset needs, which is the one everybody reaches for first. It is how many dimensions the input varies along. A feature that sorts tickets for a single team, in one language, inside one product area, is a contained problem, and a few hundred labeled examples is an afternoon of careful work that product can do on its own. A feature that pulls the line items out of supplier invoices, across hundreds of vendor layouts, currencies, tax regimes, and every idiosyncratic way a single charge can be worded, is a coverage problem along every one of those dimensions at once, and coverage problems of that shape are a data science project in their own right. It is the same request, very often the same sentence in the same ticket, and yet two completely different asks. The data decides which one you have been handed.

The two answers that don’t get you out of it

Two answers usually get offered at about this point, and neither gets you out of needing the skill:

  1. Just generate the ground truth with AI. This does not remove the skill requirement so much as move it somewhere less visible. Evidently’s guidance on synthetic test data is refreshingly blunt that doing it well still takes prompting proficiency, careful metric design, the infrastructure to run and track results, and a human checking every generated case by hand.
  2. Start smaller than they asked for. This is the more encouraging one, and the reason not to leave the meeting despairing: a smaller set of high-signal tests very often beats a sprawling one, so the honest first move is usually far narrower, and more achievable, than the request made it sound.

There is a further wrinkle worth conceding plainly, because it points at what this series is really about. Even the question I just called useful is, underneath, a statistics question. Working out how many labeled examples you need before a 95% score means anything, and making sure every dimension of variation is represented in something like the right proportion, is applied data science, and most product managers, through no fault of their own, were never trained in it. That is why data science has to be in the room. Product still owns the quality bar; what it does not own is the training to judge whether a dataset is large enough and balanced enough to trust, and that judgment is what the numbers depend on.

So before you agree to anything in that meeting, ask the one question that sorts the whole thing out: how many dimensions does the input vary along? The answer tells you whether you are being handed an afternoon of annotation or a project that genuinely needs someone who has done this before, and, conveniently, it is a question you can answer in the meeting where it comes up.

This is the second of four articles on the work an AI feature demands before anyone can write requirements for it. The next one goes underneath both of these questions, to why the specification could not have been written first, no matter who was holding the pen.

Matt Genovese
Matt Genovese
Founder & Product Strategy Lead

Matt founded Planorama Design after a career spanning semiconductor engineering and enterprise software, where he saw the same pattern over and over: features that shipped without the requirements and design work that would have made them succeed. He writes about the intersection of AI, product strategy, and the interaction design that carries them.

Let's meet.

Tell us what you're working on. We'll give our honest perspective, and share how we've helped similar teams address their challenges.

Schedule a Discovery Call