DuskByte

If someone still checks the output, what did the AI actually do?

September 17, 20268 min read

Most AI proposals turn on a single question nobody asks out loud: can a person check the output faster than they could have produced it themselves? The test, the four questions to ask before you spend anything, and the five numbers to write down first.

There is a question a lot of business owners are too polite to ask out loud, and it is the most useful question in the room.

You buy an AI tool. It reads your documents, or drafts your replies, or scores your leads. Then someone on your team checks everything it produced, because of course they do. Nobody is signing off a supplier invoice on trust.

So what did you buy? If a person still reads every line, where did the saving go?

We could not answer this cleanly for a while either. Here is the answer we settled on, and the test that comes out of it.

The mental model that causes the confusion

Most AI is sold as though it replaces the person making the decision. It does not, and the good implementations were never trying to.

What it actually changes is the cost of everything that happens before the decision.

Think about someone pulling twelve figures out of a forty page supplier agreement. Today they read the agreement, find each figure, and type it into a sheet. That is most of a morning. The decision at the end, is this contract acceptable, takes two minutes.

Now imagine the twelve figures arrive already extracted, each one showing the exact line it came from. Your person reads twelve highlighted lines, confirms them, and makes the same decision. Ten minutes instead of a morning.

Same person. Same judgement. Same accountability. The thing that changed is what they spent their attention on. They stopped producing and started verifying.

The human in the loop is not the failure. It is what makes the output safe to use at all.

The one test

That gives you a single question to put to any AI proposal, including the ones you are tempted by.

Can a competent person check the output faster than they could have produced it themselves?

If yes, it pays. If no, it does not, however impressive the demonstration was.

It sounds almost too simple. Watch how much it decides.

Extracting contract terms, with each value linked to its source line. Your person glances at the line and confirms. Checking is far cheaper than finding. This pays.

The same extraction, with no source line. To check anything, your person has to read the whole agreement. You have not removed the work, you have moved it, and added a new risk: a tired reviewer approving something they did not really check. This does not pay.

Drafting a reply to a routine customer question. Reading and adjusting a draft is faster than writing from a blank page. This pays.

"Write our three year strategy." Working out whether a strategy is any good is at least as hard as writing one. There is nothing to check it against. This does not pay, and no amount of prompt tuning will fix it.

A gate, can a person check the output faster than produce it. Two examples pass and two fail.

The consequence, and it is not what vendors emphasise

If the value lives in verification, then the valuable engineering is the verification layer, not the model.

That means the unglamorous parts. Every answer carrying a link to the source it came from. A signal when the system is unsure rather than a confident guess. Output structured so your team can compare this month's against last month's and see what moved. A defined route for the minority of cases the system should not attempt at all.

When a demonstration skips all of that and shows you a confident paragraph, it is showing you the cheap half of the problem.

Two panels with identical components. With a source line the reviewer checks one line; without it they re-read the whole document. The return arrow is the only difference.

Four questions before you spend anything

Ask these about the specific task, not about AI in general.

1. How often does this happen? Hundreds of times a month is a candidate. Five times a year is an afternoon of somebody's time, and you should just buy the afternoon.

2. Can a person check it faster than do it? The deciding question. If checking takes as long as doing, stop here.

3. When it gets something wrong, who notices? A wrong answer a reviewer catches is cheap. A wrong answer that flows straight into an invoice, a quote or a customer email is a different category of problem, and needs a different design.

4. Can anyone tell afterwards whether it was right? If there is no way to know, you cannot test it, you cannot improve it, and you will never be able to prove it paid for itself.

Two or more uncomfortable answers, and the honest recommendation is usually a set of rules rather than AI. A good part of what gets requested as artificial intelligence is a lookup table and a few conditions. We say so when that is the case. It is a shorter conversation and a better outcome.

Four screening questions in sequence, each with the reason to drop out. Question two, can a person check it faster than do it, is marked as the deciding one.

Write down the numbers before you build, not after

The reason so few businesses can say whether their AI spending worked is that nobody measured the starting point.

Before anything is built, capture five numbers:

That last one is the one people skip, and it is the one that decides whether the saving is real. If your reviewer is still at their desk for the same eight hours with nothing new to do, you have changed their day rather than your costs.

Afterwards, measure the same things, and include the cost of the AI and the review time in the new figure. The difference is your answer. Not a feeling about it. The actual difference.

Five numbers to record before building: volume, minutes per item, error rate and cost, cost per item, and what the person does with the freed time.

We learned this before AI existed

A European foodservice buying group came to us with a pricing problem. Fifty distributors, more than 10,000 products, and over 100 supplier sources arriving as REST, FTP, CSV and XML, all disagreeing with each other. Preparing a price list took more than five days.

The system we built did not try to take the decisions away from anyone. It rejected bad data at the boundary, and escalated the genuinely ambiguous cases to a person who could judge them.

Price preparation went from five days to under one. Errors fell by 95%. Promotions went out 50% faster.

The people still decided. Deciding just became cheap.

That is the same shape as every AI project that works, and we had built it years before we shipped our first generative AI system in January 2023. The technology changed. The principle did not.

Three outcome figures from a foodservice buying group pricing platform: price preparation from five days to under one, 95 percent fewer errors, 50 percent faster promotions.

What to do on Monday

Pick the one process in your business that happens most often and annoys people most.

Time it honestly for a week. Write down the five numbers above.

Then ask the fourth question: when this goes wrong, who notices? If the answer is nobody, fix that before you automate anything, because AI will simply make the same mistake faster.

You may well find the answer is not AI at all. That is a good outcome. It is certainly a cheaper one than finding out in month six.

If you want a second opinion on which of your processes would actually pay, we are happy to look. We will tell you which ones will not.

Want to talk through your own project?

Book a call. You'll talk to the person who'd actually architect it, not an account manager.