The executive workshop ends with forty-seven ideas on the wall. Each has a demo somewhere — a vendor’s, a consultancy’s, an enthusiastic colleague’s. The board wants to know which three to fund. “What can AI do for us?” is the wrong question; every idea on the wall answers it. The right question is which of these will still be running, paying for itself and reported in the numbers eighteen months from now.
The evidence says that most will not, and it also says why. The controlled field studies that found real gains are specific: customer-support agents resolving fourteen percent more issues per hour with an assistant, and novices thirty-four percent more; consultants completing twelve percent more tasks, a quarter faster and at markedly higher quality on work inside the model’s competence, and doing measurably worse on a task outside it. Next to those sit surveys in which the large majority of corporate pilots show no measurable effect on profit and loss at all — a much-quoted MIT NANDA study put it at ninety-five percent, and whatever one thinks of the method, the direction matches what I see. The difference between the two groups is not the technology. It is the method used to choose, measure and adopt.
Value has four mechanisms, and only one of them counts itself
Every use case creates value through one of four mechanisms, and they differ enormously in how much of the claimed value ever reaches the profit and loss.
| Mechanism | Shows up as | Reaches the P&L when … |
|---|---|---|
| Time | Hours saved per task, times volume | Capacity is redeployed, overtime drops, or a planned hire is avoided. Otherwise it is a nicer day — fine, but not a business case. |
| Quality | Fewer errors, less rework, fewer complaints, fewer audit findings | The error rate was measured before and after, and each error had a cost: rework hours, claims, penalties. |
| Revenue | A new offering, a shorter sales cycle, higher conversion, better retention | It is billed. The only mechanism that counts itself. |
| Risk | Avoided loss, better compliance evidence, a lower premium | Someone with authority accepts the estimate — a risk officer, an insurer, an auditor. The hardest; use sparingly. |
The slide-good use case almost always claims time savings that never convert. Two hours a week across eight hundred people is sixteen hundred hours a week — and zero euros, unless something changes in headcount, overtime or throughput. Write down which of the three will change and who owns making it happen. If nobody will own it, the value is real for the people and absent for the organisation, and the business case should say so.
A business case that survives scrutiny
The example is deliberately unglamorous. An accounts-payable team codes 2,400 invoices a week. Observation over three weeks — not an estimate from a manager — puts the average at four minutes per invoice. A pilot on real invoices, judged against an evaluation set of known-correct codings, brings it to one and a half minutes with the assistant proposing and a person confirming. That is the measured effect; a vendor’s “up to seventy percent” is not.
Then the adoption assumption: sixty percent of the volume goes through the assistant in year one. Not a hundred. Experienced clerks trust their own speed; the study on support agents found the gains concentrated in newer staff and near zero for the most experienced, who may adopt late or not at all. Sixty percent of a hundred hours a week is sixty hours, about one and a half full-time equivalents, roughly 95,000 euros a year fully loaded — if the capacity is redeployed or a planned hire is avoided. The cost, from the build-buy-partner analysis: about 45,000 to build with a partner and 2,500 a month to run, including model access, hosting, maintenance and the evaluation runs. Net, about 5,500 euros a month against 45,000 up front: payback in month nine.
Feasibility is mostly a data question
Feasibility conversations tend to be about models. They should be about data, in five questions.
- Does the data exist? Not “somewhere in the system”, but as records with the fields the use case needs, for the period it needs.
- May you use it? A lawful basis under the GDPR for personal data, and — in Germany — the works council’s co-determination right under § 87 of the Betriebsverfassungsgesetz for any system able to monitor performance or behaviour, which most assistants that log usage are. Involve the council early; a use case that surprises it is a use case that stalls.
- Is it good enough? A wiki is not a corpus. Documents without owners, dates or validity are the most common reason a promising retrieval use case fails on contact with real questions; the context-engineering guide covers what a corpus needs.
- Where does the output go? An answer that has to be copied into another system by hand loses most of its time saving on the way.
- Can you evaluate it? Can you write two hundred real questions with known-good answers? If not, you cannot know whether the pilot worked, and the jagged frontier will find you: the consultancy field experiment showed people with AI doing nineteen percentage points worse on the one task chosen to sit outside the model’s competence. Your tasks, your evaluation set.
A use case with high value and no usable data is a data project first. Fund it as one, honestly, and expect no payback in year one.
The AI Act draws lines through the portfolio
Some of the highest-value ideas on the wall are high-risk under the AI Act’s Annex III: screening applicants, evaluating employees, scoring creditworthiness, pricing risk in life and health insurance, deciding access to essential services. High-risk does not mean “don’t”. It means human oversight, logging, monitoring, information to affected workers and — for providers — conformity assessment and technical documentation, all of which change the payback calculation. The prohibited practices have applied since February 2025 and the transparency obligations since August 2026. For high-risk systems, the Digital Omnibus on AI that entered into force in July 2026 moved the application dates to December 2027 for Annex III systems and August 2028 for AI embedded in regulated products. That is time to design the oversight properly, not a reason to ignore it — and a reason to classify every use case in the portfolio now, so that the ones carrying obligations are budgeted with them.
Map the portfolio
Two things the map makes visible that a list hides. The slide-good case sits in the top-left — enormous claimed value, no data, and a high-risk classification on top — and it is the one the vendor demoed. And in the bottom-right sits a use case that is cheap to build and expensive to govern: applicant screening is feasible in an afternoon with any modern model, and it is a high-risk system with a works council, an oversight design and a documentation duty attached. The map does not say no to either. It says what each would actually cost.
Balance the portfolio like an investor
Forty-seven ideas become fourteen once each needs a named value mechanism, a data owner and a check against the prohibited practices. Fourteen become six once a baseline has been measured and a pilot effect exists on an evaluation set. Six become three once payback within twelve months at a realistic adoption rate is required — or the idea is honestly declared a capability builder and funded as one.
The balance matters as much as the selection. Quick wins fund the programme and buy the credibility to do the harder things. Capability builders — the document pipeline, the retrieval layer, the evaluation harness, the platform decision — make the second wave cheap; they rarely pay back alone within a year and should not be asked to. Options are small, time-boxed experiments on things that could become offerings, which brings us to the half of the portfolio most strategies leave out.
The business-development half: what you can now sell
Most AI strategies stop at efficiency. The more interesting half faces outward, and it follows three patterns.
An internal capability becomes a product feature. The machine builder whose service engineers use a knowledge assistant over twenty years of maintenance records can offer that assistant to customers as a support tier. The laboratory whose protocol assistant cut onboarding time can license it to partners. The value mechanism is revenue, the only one that counts itself, and the feasibility work has already been paid for.
Partnerships trade complementary capabilities. Your data and domain, their model operations; your distribution, their product. These deals fail on the terms that were treated as legal boilerplate: who owns the evaluation data, who may train on what, where inference runs, and how either side leaves. Sovereignty and exit are part of the deal, not an appendix.
Answering the buyer’s new questions is itself business development. Tenders in regulated sectors now ask where the data goes, which law can reach it, how the AI Act is handled and what happens when the model provider changes its terms. Organisations that answer with an architecture rather than a promise are winning those tenders. The temptation on the other side is to put “AI-powered” on an offering with no measured effect behind it. Regulated buyers’ due diligence has caught up with that, and it costs trust that is slow to earn back.
What this means for your next quarter
Take the wall of ideas and run the funnel once, fast. A value mechanism and a data owner for every idea: two weeks. Three measured baselines. One pilot with an evaluation set. Fund two quick wins and one capability builder, put one outward-facing option on the list, and report actuals — not projections — in the quarter after. The portfolio will be much smaller than the wall. It will also still be there in eighteen months, which is the whole point.
Sources and further reading
- Generative AI at Work (opens in a new tab) — Brynjolfsson, Li & Raymond, The Quarterly Journal of Economics, 2025
- Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality (opens in a new tab) — Dell’Acqua et al., Organization Science, 2025
- The GenAI Divide: State of AI in Business 2025 — MIT NANDA, 2025
- Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (AI Act), Annex III (opens in a new tab) — Official Journal of the European Union, 2024
- Regulation (EU) 2026/1744 amending Regulation (EU) 2024/1689 (Digital Omnibus on AI) (opens in a new tab) — Official Journal of the European Union, 2026
- Betriebsverfassungsgesetz § 87 — Mitbestimmungsrechte (opens in a new tab) — Bundesministerium der Justiz, 2024