Reporting
Green on the report, red underneathThe issue
A core AI service quietly stopped delivering for days, while every dashboard and status report still showed green.
The result
Status lights were replaced by outcome checks: the platform now proves real work is being produced and raises an alarm when it is not.
Lesson learned: Ask for evidence of outcomes, not status colours. It is the most common reason programs are caught by surprise.
The case study: Detect, fix, escalate →▶ Watch the answer: Who fixes it when something breaks overnight?
Business continuity
The recovery plan nobody had testedThe issue
Backups ran every night, but one had been failing quietly while another job made everything look fine, and none had ever been restored.
The result
Every backup now has one owner, failures reach a person, and restores are tested from the off-site copy.
Lesson learned: Disaster recovery is proven only by recovering. Test it before you need it.
The case study: Proven by restoring →▶ Watch the answer: Who fixes it when something breaks overnight?
Third-party risk
One supplier behind every safety netThe issue
The backup for the main AI service depended on the same supplier, so when one failed, both failed together.
The result
An independent provider was added, and the switch was tested by forcing a failure on purpose.
Lesson learned: Concentration risk hides inside contingency plans. Check that the backup really is independent.
The case study: No single point of supply →▶ Watch the answer: What if your AI vendor changes the rules?
Operations
Alerts nobody sawThe issue
The channel that carried every alert was removed, and warnings went nowhere for days without anyone noticing.
The result
Alerts now watch what a person would actually notice, and the alert path has its own heartbeat.
Lesson learned: Design alerting around the person who must act, then prove that they receive it.
The case study: Detect, fix, escalate →▶ Watch the answer: Who fixes it when something breaks overnight?
Governance
An audit that agreed with itselfThe issue
Two independent reviews reached the same wrong conclusion, because both trusted a check that had silently failed.
The result
Every review now includes a challenger whose job is to break the findings before they are accepted.
Lesson learned: Agreement is not proof. Build challenge into governance, not around it.
The case study: Challenge before acceptance →▶ Watch the answer: Can you prove what your AI did?
Problem management
Fixing the symptom, not the causeThe issue
A recurring failure was blamed on billing for weeks. The real cause was a design problem nobody had tested for.
The result
Causes are now proven by reproducing the problem before time or money is spent on a fix.
Lesson learned: Fund the fix only after the cause is proven.
The case study: Every lesson becomes a rule →▶ Watch the answer: Who fixes it when something breaks overnight?
Cost control
Runaway spend from hidden retriesThe issue
When a usage limit was reached, background jobs kept retrying and failing, while front-line users noticed nothing.
The result
Retries now back off and stop, and failure rates are watched alongside budgets.
Lesson learned: Watch failure rates as closely as spend. Hidden waste rarely shows up in the budget first.
The case study: One counted doorway →▶ Watch the answer: Why does the AI bill keep climbing?
Security
Too much access in one accountThe issue
A routine task run with administrator rights changed file ownership and took the main communication channel down for hours.
The result
Each service now runs under its own identity, and a restart must prove the service came back.
Lesson learned: Least privilege protects availability as much as it protects security.
The case study: Identity travels with every request →▶ Watch the answer: Who is your AI acting for?
Architecture
Customization debtThe issue
Most early incidents traced back to home-made fixes written to get things working quickly.
The result
Proven products replaced the custom scripts, and every change is now labelled standard or custom.
Lesson learned: Clean Core applies to AI: standard first, custom only by exception.
The case study: Clean Core for AI →▶ Watch the answer: Can you prove what your AI did?