September 13, 2026
doctor

Every transformation programme has a favourite dashboard. Value identified, savings surfaced, opportunities captured, all trending pleasantly upward. Executives review it monthly, and the direction of travel is always the same, because the system generating it was built to find things that look like wins.

That is the blind spot, and one very large industry just paid a nine-figure price to demonstrate why it matters.

The case study

American health insurers covering more than thirty million older adults are paid by the government according to how ill their members’ documented conditions show them to be. To ensure nothing went unrecorded, the industry deployed exactly the kind of programme every transformation deck promises: automated review of years of clinical records, surfacing conditions that had been treated but never properly coded.

The technology worked. It found things. Dashboards filled with value identified.

What nobody instructed was the direction. The systems searched for conditions to add, because additions increased revenue, and there was no equivalent search for conditions that should be removed, because removals decreased it. For years, the automated review found errors almost exclusively in one direction, and every individual finding could be defended.

Federal auditors eventually asked the aggregate question. Reviews of three insurance plans this spring found 81 to 91 percent of certain sampled high-risk diagnosis codes unsupported by the records behind them. A major Medicare Advantage insurer settled federal claims for 117.7 million dollars, with prosecutors focused specifically on the one-directional design of the review programme rather than on individual errors.

Why this is a governance failure, not a technology failure

The software did what it was asked. The failure lived one level up, in the specification, and it is a failure that repeats across industries because it never looks like a failure while it is happening.

Consider the anatomy. A business case is written in terms of value recovered, so value recovered becomes the metric. The metric becomes the dashboard. The dashboard becomes the definition of whether the programme is working. At no point does anyone commission the mirror-image capability, finding errors that cost money, because it has no sponsor, no KPI, and no obvious return.

Then a regulator, an auditor, or a journalist runs a single query: across all corrections this system has made, what is the ratio of favourable to unfavourable? A result near a hundred to zero requires no proof of intent to be damning. The ratio is the finding.

The rebuild worth copying

The corrective architecture emerging in healthcare is instructive for any organisation running consequential automation. Modern ai risk adjustment systems pair neural components that read unstructured text with symbolic components that validate each finding against explicit rules, and they are built so the same pass that surfaces missed items also flags recorded items the evidence cannot support. Both streams route through human review, and every output carries its evidence trail.

Three governance moves follow directly, and none requires new technology.

Instrument direction as a first-class metric. For any system that modifies consequential data, report the ratio of favourable to unfavourable corrections to the risk committee, quarterly. It costs almost nothing and it is the exact query an examiner will eventually run.

Give the unprofitable direction to an owner. Symmetry does not survive on policy statements. If finding errors against your own interest has no headcount and no KPI, deadline pressure restores the asymmetry within two quarters, silently.

Attach evidence at output, not on request. Systems that emit conclusions with their reasoning are auditable by construction. Systems that store conclusions and reconstruct reasoning later are doing archaeology under deadline, which is where the worst outcomes happen.

The question for your next steering committee

Digital transformation is not slowing down, and it should not. But the healthcare precedent hands every leadership team a single diagnostic question worth institutionalising: for each automated system we run, if a hostile examiner reviewed its full correction history, what story would the direction tell?

Organisations that can answer “both ways, with evidence” have built analytics. Organisations that cannot have built a dashboard that only brings good news, which is a comfortable thing to own right up until somebody outside the building runs the query themselves.

Leave a Reply

Your email address will not be published. Required fields are marked *