Jakob Lange

← All insights AI strategy in practice

Which AI use cases pay for themselves within a year?

A portfolio method for AI use cases: four value mechanisms and how much of each reaches the profit and loss, a business case built on measured baselines and an honest adoption assumption, a value–feasibility map with the AI Act drawn on it — and the business-development half most strategies forget.

The executive workshop ends with forty-seven ideas on the wall. Each has a demo somewhere — a vendor’s, a consultancy’s, an enthusiastic colleague’s. The board wants to know which three to fund. “What can AI do for us?” is the wrong question; every idea on the wall answers it. The right question is which of these will still be running, paying for itself and reported in the numbers eighteen months from now.

The evidence says that most will not, and it also says why. The controlled field studies that found real gains are specific: customer-support agents resolving fourteen percent more issues per hour with an assistant, and novices thirty-four percent more; consultants completing twelve percent more tasks, a quarter faster and at markedly higher quality on work inside the model’s competence, and doing measurably worse on a task outside it. Next to those sit surveys in which the large majority of corporate pilots show no measurable effect on profit and loss at all — a much-quoted MIT NANDA study put it at ninety-five percent, and whatever one thinks of the method, the direction matches what I see. The difference between the two groups is not the technology. It is the method used to choose, measure and adopt.

Value has four mechanisms, and only one of them counts itself

Every use case creates value through one of four mechanisms, and they differ enormously in how much of the claimed value ever reaches the profit and loss.

Mechanism Shows up as Reaches the P&L when …
Time Hours saved per task, times volume Capacity is redeployed, overtime drops, or a planned hire is avoided. Otherwise it is a nicer day — fine, but not a business case.
Quality Fewer errors, less rework, fewer complaints, fewer audit findings The error rate was measured before and after, and each error had a cost: rework hours, claims, penalties.
Revenue A new offering, a shorter sales cycle, higher conversion, better retention It is billed. The only mechanism that counts itself.
Risk Avoided loss, better compliance evidence, a lower premium Someone with authority accepts the estimate — a risk officer, an insurer, an auditor. The hardest; use sparingly.

The slide-good use case almost always claims time savings that never convert. Two hours a week across eight hundred people is sixteen hundred hours a week — and zero euros, unless something changes in headcount, overtime or throughput. Write down which of the three will change and who owns making it happen. If nobody will own it, the value is real for the people and absent for the organisation, and the business case should say so.

A business case that survives scrutiny

The business case as arithmetic Six boxes. Baseline: 2,400 invoices a week at 4.0 minutes each, measured over three weeks. Measured effect: 4.0 to 1.5 minutes in a pilot judged against an evaluation set. Adoption: 60 percent of volume in year one. Gross value: 60 hours a week, about 1.6 full-time equivalents or 95 thousand euros a year, only if the capacity is redeployed or a hire avoided. Cost: 45 thousand to build plus 2.5 thousand a month to run. Payback: month nine. Baseline 2,400 invoices a week, 4.0 min each — measured over three weeks in the team, not estimated by a manager × Measured effect 4.0 → 1.5 min pilot on real invoices, judged against an evaluation set — not a vendor claim × Adoption 60 % of the volume in year one — the assumption that kills most business cases = Gross value 60 h / week ≈ 1.6 FTE ≈ €95k a year, but only if capacity is redeployed or a planned hire is avoided Cost €45k + €2.5k/mo build with a partner; hosting, model access, maintenance, evaluation runs Payback month 9 net ≈ €5.5k a month against €45k up front; from then on the actual replaces the plan
Fig. 01 The business case as arithmetic, for an invoice-coding assistant in an accounts-payable team. Every box is a number somebody measured or explicitly assumed. The payback month is what gets reported, and from the first quarter in production the actual replaces the projection.

The example is deliberately unglamorous. An accounts-payable team codes 2,400 invoices a week. Observation over three weeks — not an estimate from a manager — puts the average at four minutes per invoice. A pilot on real invoices, judged against an evaluation set of known-correct codings, brings it to one and a half minutes with the assistant proposing and a person confirming. That is the measured effect; a vendor’s “up to seventy percent” is not.

Then the adoption assumption: sixty percent of the volume goes through the assistant in year one. Not a hundred. Experienced clerks trust their own speed; the study on support agents found the gains concentrated in newer staff and near zero for the most experienced, who may adopt late or not at all. Sixty percent of a hundred hours a week is sixty hours, about one and a half full-time equivalents, roughly 95,000 euros a year fully loaded — if the capacity is redeployed or a planned hire is avoided. The cost, from the build-buy-partner analysis: about 45,000 to build with a partner and 2,500 a month to run, including model access, hosting, maintenance and the evaluation runs. Net, about 5,500 euros a month against 45,000 up front: payback in month nine.

Feasibility is mostly a data question

Feasibility conversations tend to be about models. They should be about data, in five questions.

  1. Does the data exist? Not “somewhere in the system”, but as records with the fields the use case needs, for the period it needs.
  2. May you use it? A lawful basis under the GDPR for personal data, and — in Germany — the works council’s co-determination right under § 87 of the Betriebsverfassungsgesetz for any system able to monitor performance or behaviour, which most assistants that log usage are. Involve the council early; a use case that surprises it is a use case that stalls.
  3. Is it good enough? A wiki is not a corpus. Documents without owners, dates or validity are the most common reason a promising retrieval use case fails on contact with real questions; the context-engineering guide covers what a corpus needs.
  4. Where does the output go? An answer that has to be copied into another system by hand loses most of its time saving on the way.
  5. Can you evaluate it? Can you write two hundred real questions with known-good answers? If not, you cannot know whether the pilot worked, and the jagged frontier will find you: the consultancy field experiment showed people with AI doing nineteen percentage points worse on the one task chosen to sit outside the model’s competence. Your tasks, your evaluation set.

A use case with high value and no usable data is a data project first. Fund it as one, honestly, and expect no payback in year one.

The AI Act draws lines through the portfolio

Some of the highest-value ideas on the wall are high-risk under the AI Act’s Annex III: screening applicants, evaluating employees, scoring creditworthiness, pricing risk in life and health insurance, deciding access to essential services. High-risk does not mean “don’t”. It means human oversight, logging, monitoring, information to affected workers and — for providers — conformity assessment and technical documentation, all of which change the payback calculation. The prohibited practices have applied since February 2025 and the transparency obligations since August 2026. For high-risk systems, the Digital Omnibus on AI that entered into force in July 2026 moved the application dates to December 2027 for Annex III systems and August 2028 for AI embedded in regulated products. That is time to design the oversight properly, not a reason to ignore it — and a reason to classify every use case in the portfolio now, so that the ones carrying obligations are budgeted with them.

Map the portfolio

Value–feasibility map with the AI Act drawn on it Scatter map: feasibility on the horizontal axis, value within twelve months on the vertical axis, four quadrants. Top right, build now: claims correspondence drafting, policy questions for agents, invoice coding. Top left, fix the data first: fully automated claims settlement, which looks great on a slide, and underwriting risk scoring, both marked high-risk. Bottom left, drop or defer: churn prediction. Bottom right, cheap experiments: applicant screening, marked high-risk, and meeting summaries. Contract clause extraction sits near the centre. size = investment high-risk under the AI Act (Annex III) the slide-good case Fix the data first Build now Drop or defer Cheap experiments Feasibility → data · integration · evaluability Value within 12 months → Claims correspondence drafting Invoice coding Policy Q&A for agents Fully automated claims settlement looks great on a slide Underwriting risk scoring Churn prediction Applicant screening Meeting summaries Contract clause extraction
Fig. 02 A value–feasibility map for the composite insurer from the build-buy-partner guide. Bubble size is the investment required; a dashed ring marks a high-risk category under the AI Act; the amber bubble is the one that looks best on a slide and has no usable data underneath it. Illustrative placement. Value is the twelve-month estimate after the adoption assumption; feasibility combines data readiness, integration effort and whether an evaluation set can be built.

Two things the map makes visible that a list hides. The slide-good case sits in the top-left — enormous claimed value, no data, and a high-risk classification on top — and it is the one the vendor demoed. And in the bottom-right sits a use case that is cheap to build and expensive to govern: applicant screening is feasible in an afternoon with any modern model, and it is a high-risk system with a works council, an oversight design and a documentation duty attached. The map does not say no to either. It says what each would actually cost.

Balance the portfolio like an investor

From forty-seven ideas to three in production, and the portfolio balance Left, a funnel of four stages: 47 ideas on the wall; 14 screened after requiring a named value mechanism, a data owner and no prohibited practice; 6 business-cased after a measured baseline and a pilot effect on an evaluation set; 3 in production after requiring payback within twelve months at realistic adoption or explicit funding as a capability builder. Right, a stacked bar: 60 percent quick wins, 30 percent capability builders, 10 percent options. 47 Ideas on the wall gate: a named value mechanism, a data owner, no prohibited practice under the AI Act 14 Screened gate: baseline measured; effect measured in a pilot on an evaluation set 6 Business-cased gate: payback within 12 months at realistic adoption — or honestly funded as a capability builder 3 In production measured quarterly; the number the board sees Portfolio balance 60 % quick wins payback within a year; they fund the programme 30 % capability builders data, retrieval, platform — they make the next wave cheap 10 % options small experiments on things that could become offerings
Fig. 03 From the wall to production: four stages with the gate each idea has to pass, and the balance that keeps a programme alive — quick wins that pay within a year, capability builders that make the second wave cheap, and a few options.

Forty-seven ideas become fourteen once each needs a named value mechanism, a data owner and a check against the prohibited practices. Fourteen become six once a baseline has been measured and a pilot effect exists on an evaluation set. Six become three once payback within twelve months at a realistic adoption rate is required — or the idea is honestly declared a capability builder and funded as one.

The balance matters as much as the selection. Quick wins fund the programme and buy the credibility to do the harder things. Capability builders — the document pipeline, the retrieval layer, the evaluation harness, the platform decision — make the second wave cheap; they rarely pay back alone within a year and should not be asked to. Options are small, time-boxed experiments on things that could become offerings, which brings us to the half of the portfolio most strategies leave out.

The business-development half: what you can now sell

Most AI strategies stop at efficiency. The more interesting half faces outward, and it follows three patterns.

An internal capability becomes a product feature. The machine builder whose service engineers use a knowledge assistant over twenty years of maintenance records can offer that assistant to customers as a support tier. The laboratory whose protocol assistant cut onboarding time can license it to partners. The value mechanism is revenue, the only one that counts itself, and the feasibility work has already been paid for.

Partnerships trade complementary capabilities. Your data and domain, their model operations; your distribution, their product. These deals fail on the terms that were treated as legal boilerplate: who owns the evaluation data, who may train on what, where inference runs, and how either side leaves. Sovereignty and exit are part of the deal, not an appendix.

Answering the buyer’s new questions is itself business development. Tenders in regulated sectors now ask where the data goes, which law can reach it, how the AI Act is handled and what happens when the model provider changes its terms. Organisations that answer with an architecture rather than a promise are winning those tenders. The temptation on the other side is to put “AI-powered” on an offering with no measured effect behind it. Regulated buyers’ due diligence has caught up with that, and it costs trust that is slow to earn back.

What this means for your next quarter

Take the wall of ideas and run the funnel once, fast. A value mechanism and a data owner for every idea: two weeks. Three measured baselines. One pilot with an evaluation set. Fund two quick wins and one capability builder, put one outward-facing option on the list, and report actuals — not projections — in the quarter after. The portfolio will be much smaller than the wall. It will also still be there in eighteen months, which is the whole point.

Sources and further reading

  1. Generative AI at Work (opens in a new tab) — Brynjolfsson, Li & Raymond, The Quarterly Journal of Economics, 2025
  2. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality (opens in a new tab) — Dell’Acqua et al., Organization Science, 2025
  3. The GenAI Divide: State of AI in Business 2025 — MIT NANDA, 2025
  4. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (AI Act), Annex III (opens in a new tab) — Official Journal of the European Union, 2024
  5. Regulation (EU) 2026/1744 amending Regulation (EU) 2024/1689 (Digital Omnibus on AI) (opens in a new tab) — Official Journal of the European Union, 2026
  6. Betriebsverfassungsgesetz § 87 — Mitbestimmungsrechte (opens in a new tab) — Bundesministerium der Justiz, 2024

Contact

Start with a conversation.

No forms, no funnels. Write me a short note about your situation — I answer personally, usually within two working days.

Mon – Fri, 18:00 – 20:00 CET