Our proactive agent did something worth telling in full.
It scans global flight data for anomalies every day. One day it flagged a partner operator: flight volume over the previous month had fallen off a cliff. No external news triggered the check — it hadn't read something and gone looking; it found this on its own, with nobody asking. The system ran several rounds of verification, ruled out the usual explanations (seasonality, fleet changes, under-reported data), judged it an operational anomaly, and pushed it to management.
Weeks later the business had a European flight to fly. Having worked with them before, procurement reasonably picked that operator. A deposit was paid.
At the final payment, the system stopped it — the risk flag was attached. On verification: the operator's AOC had been revoked and it was under regulatory investigation.
That interception avoided roughly half a million dollars in direct loss. It also avoided something harder to undo: had the money gone out and the flight not existed, the client would have found out at the airport, standing there while no aircraft arrived.
But the insight is not what did the work
The easy conclusion is "our proactive insights are accurate." I think that's the wrong one.
The same judgement, had it merely appeared on a management dashboard, would most likely have ended like this: someone read it, agreed with it, perhaps forwarded it. Three weeks later procurement ran its process, nobody connected that memory to this payment, and the money went out.
The insight created no value. What created value was that it sat somewhere it could stop a payment.
That isn't wordplay. It decides where your engineering should go.
Three levels, and most people stop at the second
How far proactive intelligence actually lands falls into three levels:
Level one: visible. The insight reaches a dashboard, a daily digest, a slide in the weekly review. Its fate depends on whether somebody happens to see the right item at the right moment. The overwhelming majority of "AI insight" products stop here, and their metric is usually volume — how many were generated today.
Level two: pushed. The insight goes looking for people: into a chat group, an inbox, a notification tray. Better than level one, but it still only nudges the probability that someone sees it. People can leave it unread, can see it while doing something else entirely, can forget it in three weeks.
Level three: able to stop things. The insight is no longer a message; it is a piece of state hanging on a business path. When a payment reaches this point, it must first ask whether this counterparty carries a risk flag. Nobody has to remember, because the process remembers.
The difference between level three and the other two isn't "more proactive." It is whether this judgement has the authority to say no.
The ceiling on levels one and two is whether someone happens to recall it. The ceiling on level three is the value flowing through that path.
The price: earn the authority before you get it
Level three isn't free. A judgement that can stop a payment is also a judgement that will block real business when it is wrong.
So the order matters. The work that looks like gold-plating — several rounds of verification, ruling out seasonality and fleet changes, marking the uncertain cases as "needs a human" rather than "blocked" — isn't there to make the insight prettier. It is the ticket of admission to the payment path.
A system that hasn't been through that verification should never be wired to a control point. It will produce enough false blocks within a fortnight to be demoted permanently back to level one — and this time a person will have switched it off deliberately.
I've written the same causality elsewhere: the real engineering in a proactive product isn't generating content, it's judging whether this moment is worth interrupting for. Once it's wired to a control point, that sentence gets heavier: whether this moment is worth halting for.
The boundary: not every judgement should hard-block
Making every insight a hard block is a different kind of failure.
Two tests decide: is the action reversible, and what does a wrong block cost?
- A non-refundable final payment → irreversible, large → hard block, released only after human review
- Routing a lead to a particular salesperson → reversible, cheap → surface it, don't stand in the way
- Sending a recommendation email to a client → semi-reversible (it can't be recalled, but the cost is bounded) → soft block, paused by default with a one-click override
One set of judgements, landed at different strengths according to the reversibility of what's downstream. This is the same rule as authorising an agent to write: high-risk actions require a human, low-risk actions shouldn't trouble one every time.
So
If you're building a proactive agent, there's a question worth asking before "are our insights accurate":
What can this judgement actually stop?
If the answer is "it will appear on a page", then however strong the model and however strict the verification, the ceiling is already set — it is a notification. And a notification's value depends on the reader's memory, which is the least reliable component in the whole system.