Human-on-the-Loop — Rethinking Oversight in the Age of Autonomous Agents
It was 2:10 in the morning over the Atlantic Ocean, on the 1st of June 2009.
Air France Flight 447 was cruising at 35,000 feet between Rio de Janeiro and Paris. Two hundred and twenty-eight souls were on board — most of them asleep, trusting the quiet hum of the engines and an autopilot that had been flying the aircraft for hours. Two co-pilots were at the controls; the captain had stepped away for a rest break. Everything was, by every measure of modern aviation, safe.
Then, in a few seconds, three small pitot tubes on the outside of the aircraft iced over. The autopilot, suddenly blind to the airspeed, did the only thing it was designed to do — it disengaged. A soft chime sounded in the cockpit:
“Autopilot disconnected.”
What followed were the most studied four minutes in the history of commercial aviation. Two trained pilots, sitting inches apart, gave opposing inputs to the aircraft. One pulled back on the stick. The other pushed forward. The plane stalled. Warnings blared. Altitude bled away — 35,000 feet… 30,000… 20,000… 10,000. When the captain rushed back into the cockpit, it was too late. At 2:14 AM, Flight 447 hit the ocean. There were no survivors.
The investigation did not blame the autopilot. It did not blame the ice. It blamed something far more uncomfortable.
It blamed the handover.
The autopilot had been so good, so quiet, so reliable, that the pilots had drifted — not physically, but cognitively. When the machine handed back the controls, the humans were no longer in the loop. They were not even on the loop. They were spectators suddenly asked to fly.
• • •
I keep returning to Flight 447 these days — because those of us building the next generation of digital government, the advisors whispering in the ears of ministers and CIOs, are about to face the same moment in the operations cockpit.
We are about to deploy thousands of AI agents into our hospitals, courts, tax authorities, border systems, and citizen portals. These agents will not just suggest — they will act. They will approve permits, triage benefits, screen applications, route emergencies, and answer citizens at 3 AM.
One day, maybe not today, maybe not next year — the pitot tubes will freeze. A model will drift. A dataset will rot. An edge case will arrive that nobody trained for. And somewhere, in some control room, a soft chime will sound the same:
“Autopilot disconnected.”
The question that should be keeping every leader awake tonight is not “Is our AI good enough?”
“When the AI drifts — who is still running the operation, who has the authority to stop it, and who answers the citizen waiting on the other side of the screen?”
This is the question of human oversight in the age of autonomous agents. And the comfortable answers we have been giving ourselves — “don’t worry, there’s a human in the loop” — are no longer enough. The loop has changed. And the role of the human must change with it.
Why “Human-in-the-Loop” Is No Longer Enough
For the past decade, whenever an AI system was proposed — a credit scoring model, a fraud engine, a triage tool — someone in the room would ask the necessary question: “But will a human review the output?”
And someone else would lean forward, smile reassuringly, and say:
“Of course. There will always be a human in the loop.”
The room would relax. The project would move forward.
For a while, this answer was enough — because the AI of the last decade was a predictor, not an actor. It produced a score, a recommendation, a label. A human read it, decided, signed. The loop was small, the volume was manageable, and the human was always the last line of defense.
The AI we are living with now does not produce a score for a human to review. It opens a ticket, queries databases, calls APIs, drafts an email, books an appointment, updates a record, and notifies the citizen — all before the human has finished their first cup of coffee. Multiply that by ten thousand transactions a day in a tax authority, a hundred thousand in a national health system, a million in a citizen portal — and the comforting answer cracks.
“You cannot put a human in every loop when the loops are firing every two seconds. The math simply does not work.”
The uncomfortable truth is this: a human-in-the-loop that cannot keep up is not oversight. It is theatre.
It is the supervisor who approves five hundred decisions a day by lunch. The case worker who clicks “accept” on every recommendation because the queue never empties. The manager who signs off on an algorithmic decision they did not have the time or the information to actually understand.
We do not need fewer humans in the loop. We need humans in the right loop, at the right altitude, with the right authority.
A quiet but important shift is taking place across global AI governance — from Human-in-the-Loop to Human-on-the-Loop. You can see it explicitly named in:
- The EU AI Act, where Article 14 distinguishes between continuous oversight and exception-based oversight for high-risk systems.
- NIST’s AI Risk Management Framework (and its Generative AI Profile), where “appropriate human oversight” is framed as contextual, not universal.
- Singapore’s Model AI Governance Framework for Agentic AI, the first national framework to formally introduce Human-on-the-Loop as a distinct operating mode.
Three continents. Three regulatory traditions. One quiet convergence on the same idea:
The human does not need to be inside every decision. The human needs to be above the system — watching it, shaping it, and ready to land it when the chime sounds.
That is a fundamentally different posture. It demands a different design, a different skillset, and a different accountability model.
The Three Modes — A Mental Model for the New Cockpit
If “Human-in-the-Loop” is no longer the universal answer, what replaces it? Nothing replaces it. It becomes one of three modes — and the leadership job is to choose between them intentionally.
For too long, we have treated human oversight as a single setting — on or off, in or out. The reality is that oversight comes in three distinct modes, and the leadership question is no longer “do we have a human reviewing this?” — it is “which mode does this decision belong in, and have we designed for it?”
| Mode | The Human’s Role | Best Fit For | Where It Breaks |
| Human-in-the-Loop (In) | Reviews and approves every decision before it takes effect. The AI suggests; the human signs. | High-risk, low-volume decisions where each case carries weight — interviews, oncology treatment plans, judicial sentencing support, large procurement awards. | Becomes a bottleneck. At scale, it forces approval without review — the very thing it was designed to prevent. |
| Human-on-the-Loop (On) | Supervises the system, not each decision. Monitors patterns, dashboards, and exceptions. Intervenes when something drifts. | Medium-risk, high-volume decisions where speed matters and patterns reveal more than individual cases: benefits triage, tax return classification, visa renewals, fraud screening, citizen service routing. | Demands new skills, new tools, and real authority to stop the system. Without these, it collapses into alert fatigue. |
| Human-out-of-the-Loop (Out) | Designs the guardrails before deployment and audits after. The human is not in the live decision flow at all. | Low-risk, very-high-volume decisions where individual error is tolerable and patterns are the only meaningful unit of oversight: spam filtering, content moderation at first pass, ticket auto-classification, anomaly flagging. | Creates a false sense of safety. Without disciplined audit, small drifts compound silently into systemic harm. |
Three things are worth pausing on as you read this table.
- First: none of these modes is “better” than the others. A government that puts every decision in the loop will collapse under its own caution. One that puts every decision out of the loop will lose public trust. Matching the mode to the decision is a leadership decision, not a technical one.
- Second: most organizations today are using the wrong mode by default — either over-supervising (which produces fatigue) or under-supervising (assuming a quarterly audit counts as oversight for a system making thousands of decisions a day). Both failures come from the same root cause: nobody made an intentional choice.
- Third: the same AI system may operate in different modes for different decisions. A citizen service agent might be out of the loop when classifying inquiries, on the loop when routing complaints, and in the loop when triggering a payment above a threshold. The mode is attached to the decision, not the system.
The shift is from “is there a human?” to:
“For this decision, at this risk level, at this volume — where should the human stand?”
The Oversight Matrix — A Framework for Choosing the Right Mode
Knowing that three modes exist is not the same as knowing which to use, when, and for what. Most AI governance conversations stop at the awareness layer, never reaching the decision layer.
So let me offer a simple framework I have been using with public sector leaders and enterprise architects. I call it The Oversight Matrix. It rests on two questions every leader can answer about any AI-driven decision:
- What is the risk to the citizen, the institution, or the mission if this decision goes wrong?
- At what volume and velocity is this decision being made?
Plot those on a two-by-two, and you get four quadrants — each pointing to a clear oversight posture.

Quadrant 1 — High Risk, Low Volume → Strict Human-in-the-Loop
Example: A system supporting refugee interview scoring, sentencing-range recommendations for a judge, or tender scoring above a national threshold.
The volume is manageable. The stakes are enormous. A single bad decision can change a life, void a contract, or end up in front of a parliamentary inquiry. The human signs every decision — and ideally, two humans sign. Speed is not the priority here. Defensibility is.
If you find yourself trying to “speed up” a Q1 decision with more automation, pause. You are almost certainly optimizing the wrong variable.
Quadrant 2 — High Risk, High Volume → Human-on-the-Loop, with an AI Observer
Example: An agent that triages social benefit applications, flags suspicious tax returns, screens visa renewals, or routes emergency calls.
The volume makes Human-in-the-Loop impossible. The risk makes Human-out-of-the-Loop unacceptable. The only viable posture is Human-on-the-Loop — supervising the system, watching patterns, ready to intervene.
But human attention is a temporary defense. By month three, the dashboard is open in a background tab. This is why Q2 needs a second layer: an AI Observer watching the operational AI — detecting drift, flagging unusual decision distributions, surfacing fairness anomalies, and escalating only what truly requires judgment.
One agent acts. One agent observes. One human supervises both. That is what mature Q2 oversight looks like.
Quadrant 3 — Low Risk, Low Volume → Don’t Automate
Example: A minister’s foreword for an annual report. A bespoke quarterly newsletter a Director General writes personally for key stakeholders.
These decisions share a quiet characteristic: automation provides no real business value. The volume is too low to justify the investment, and the personal weight of each communication is precisely what makes it valuable. A quarterly newsletter drafted by an agent is not faster in any meaningful way, and the personal voice is precisely what makes it valuable to the people receiving it..
The right answer here is often: don’t automate. Spend the budget on Q2 instead. And if an agent is introduced anyway, the human stays firmly in the loop, as the author, not the approver.
Quadrant 4 — Low Risk, High Volume → Human-out-of-the-Loop, with Living Guardrails
Example: OCR and classification of incoming citizen letters. Auto-routing of standard service requests. First-pass redaction of personal data in FOI responses. Categorization of social media mentions.
The volume makes human review impossible. The risk per decision is low enough that statistical errors are tolerable. The human exits the live flow — but designs the guardrails, sets the audit cadence, and reviews aggregate performance monthly.
But guardrails decay. Training data shifts. Citizen populations change. The legal context moves. Q4 is not “set it and forget it.” It is “set it, audit it, and reset it” at least annually.
Out of the loop is not the same as out of the governance loop.
How to Use the Matrix
- Map every AI use case in your portfolio to one quadrant, today. You will be surprised how many systems are running in the wrong mode for their current quadrant.
- The same AI system can span multiple quadrants. A citizen service agent might be Q4 when classifying inquiries, Q2 when routing complaints, and Q1 when authorizing a payment above a threshold. The quadrant attaches to the decision, not the system.
- Quadrants migrate over time. What was safely Q4 last year may be Q2 today — and your oversight has not caught up. Review your matrix every six months.
The Oversight Matrix is a leadership tool — a way to make oversight a deliberate design choice, not a default assumption. Three failure modes are quietly waiting for any organization that skips this step.
Three Failure Modes No One Talks About
The Oversight Matrix tells you where the human should stand. It does not guarantee that the human will stand there well.
In every public sector deployment I have studied, the same three failure modes appear, quietly, again and again. They rarely make the headlines until something goes wrong — but they are the reason most AI oversight programs look strong on paper and weak in practice.
Failure 1 — Rubber-Stamping
A supervisor is given a queue of five hundred AI recommendations to review by the end of the day. In the morning, she reads each one carefully. By eleven, she is skimming. By lunch, she is approving the entire batch with a single keystroke. The signature is real. The review is not.
This is rubber-stamping — the bureaucratic reflex of approving without truly reviewing. It is the most common failure of Human-in-the-Loop systems at scale, and it is dangerous precisely because it looks like oversight. The audit trail is perfect. The accountability box is checked. And yet, in any practical sense, the AI is operating autonomously.
The Discipline: design forced reflection checkpoints. Limit batch sizes. Require written justification for any override, and for a random sample of approvals. If approving is free, approval is meaningless.
Failure 2 — Alert Fatigue
A monitoring dashboard lights up with 70 exceptions on Monday morning. The supervisor investigates each one. By Wednesday, the same dashboard shows 85 exceptions, and most of them turn out to be false alarms. By Friday, the supervisor has stopped looking. By the following month, the dashboard is open in a background tab and nobody is reading it.
This is alert fatigue — the slow decay of human attention under a flood of noise. It is the most common failure of Human-on-the-Loop systems, and it is what turns Q2 into a quiet drift toward Q4 without anyone signing off on the change.
The Discipline: tune alert thresholds quarterly, not at deployment. Use an AI Observer to filter routine variance from genuine anomalies. Track the supervisor’s alert response rate as a system metric — when it drops, the system is failing, not the human.
Failure 3 — The Accountability Donut
Something goes wrong. A citizen is wrongly denied a benefit. An algorithm misclassifies a case. A wrongful arrest is made on the basis of a flagged pattern. The investigation begins.
The case worker says: “I was following the system’s recommendation.” The vendor says: “The human approved the decision.” The agency head says: “We followed industry best practice.” The regulator says: “Where is the named owner of this agent?”
And the silence that follows is the most dangerous sound in modern AI governance.
This is the accountability donut — the empty space in the middle of a system where everyone is partially responsible and nobody is fully accountable. It is not a failure of any individual. It is a failure of design.
The Discipline: assign a named human owner for every AI agent in production — a single person, with a job title, whose KPIs include the agent’s performance, drift, and outcomes. “Who owns this agent on Monday morning?” should be a question with one answer.
The Common Thread
None of these failures is caused by a bad algorithm or malicious intent. None shows up in a technical risk register.
All three are caused by the same root error: we designed the AI, but we did not design the human’s role around it.
Which raises the question we have been circling: if oversight is not just a matter of having a human present, what does the human’s role actually need to become?
Redesigning the Human Role
We have to redesign the human role. Not replace it. Not remove it. Redesign it.
The supervisor of AI agents is not the same job as the supervisor of a team of clerks, even if the title on the business card has not changed. The instincts that made someone good at reviewing a hundred case files a week are not the same instincts that will make them good at watching ten thousand decisions a day flow through an autonomous system.
In my work with public sector leaders, four capabilities are needed.
1. Pattern Literacy
The old supervisor read individual cases. The new supervisor reads distributions. They do not look at decision number 6,217 and ask “was this right?” They look at all decisions and ask “is this shape right?”, “Is the approval rate drifting?”, “Are certain demographics suddenly disadvantaged?”, and “Is the model unusually confident this week, or unusually uncertain?”.
This is a different cognitive muscle. Most of our public sector workforce has never been trained to use it. Training programs need to start now, not after the first incident.
2. Intervention Authority
It is one thing to notice that something is wrong. It is another to be allowed to stop it.
Many oversight programs put a supervisor in front of a dashboard but do not give that supervisor the authority to pause the agent, override a decision, or escalate without three layers of approval. By the time the bureaucratic chain has been climbed, the harm has already happened.
The new role must come with a clear, pre-approved set of intervention powers — including the authority to pause the agent on the spot, with accountability for using that power and for failing to use it when it was warranted.
3. Investigative Instinct
A drifting metric is not an answer. It is a question.
The new supervisor must be trained to look at an unusual pattern and ask why — to pull the threads, interview affected citizens, challenge developers, demand explanations that satisfy a non-technical audience. This is the instinct of a journalist crossed with the instinct of an internal auditor.
We do not currently hire for this. We do not currently train for this. We will need to do both.
4. Moral Courage
This is the rarest, and the most important.
There will be moments when the dashboard is green, the algorithm is confident, the vendor is reassuring, the schedule is tight, and the political pressure is real — and the supervisor will still need to say:
“Stop the system. Something is wrong here.”
That sentence carries a cost. The supervisor may be wrong. They may be overruled. They may slow down a project leadership has staked its reputation on. And yet, when the moment comes, they must be willing to say it anyway — because the alternative is the accountability donut, the wrongful denial, the headline nobody wants.
This is not a technical skill. It is a character trait. You cannot install it with a training course. You can only hire for it, develop it, protect it — and most importantly, reward it when it shows up.
A New Job Description
If you take only one practical action from this section, let it be this:
For every AI agent currently in production in your organization, write a new job description for the human who oversees it.
Not an org chart entry. Not a paragraph in an existing role. A proper job description with pattern literacy as a required skill, intervention authority as a stated power, investigative instinct as a hiring criterion, and moral courage as a culture commitment.
If you cannot write that job description today, you do not have oversight today. You have the appearance of oversight — and the appearance of oversight is the most dangerous form of it.
The role exists. The cockpit needs a pilot. It is time we designed the seat, and named the people who will sit in it.
The Cockpit, Revisited
Let me end where we began, at 2:10 in the morning, somewhere over the Atlantic, on a quiet flight that did not know it was about to be remembered forever.
The lesson of Flight 447 was never that autopilots are dangerous. Autopilots are extraordinary. The lesson was that a brilliant machine and an unprepared human are a worse combination than either one alone.
The technology did exactly what it was designed to do. And when it could no longer fly the aircraft, it handed the controls back — politely, with a soft chime. It was the handover that failed. The humans had not been trained for the moment the chime would sound. Their authority had not been clarified. Their role had drifted — not because anyone made a bad decision, but because nobody had designed the role they would need to play.
The AI is not the risk. The unprepared handover is the risk.
We will deploy the agents. The benefits are too compelling, the citizen demand is too clear, the budgets are already moving. What remains in our hands is what kind of oversight we put around them.
Will we put one human in front of a dashboard and call it governance? Or will we do the harder, slower, more important work — mapping every AI decision to a quadrant, building observer agents alongside the operational ones, writing real job descriptions for the new role, and naming a guardian for every agent we put into production?
Because in the end, someone still has to look the citizen in the eye.
When the chime sounds in your organization, and it will sound — the three questions will arrive together:
Who is running it? Who can stop it? Who is accountable for it?
If you have good answers today, you are ready. If you do not, the time to write them is now — before the autopilot disconnects, not after.
• • •
The cockpit is yours. The chime will sound. May the people sitting in the seat be the ones we have prepared.
References
- Air France Flight 447 — Bureau d’Enquêtes et d’Analyses (BEA), Final Report on the accident on 1 June 2009 (July 2012). bea.aero · Overview (Wikipedia).
- European Union — Regulation (EU) 2024/1689 (AI Act), Article 14 — Human oversight. artificialintelligenceact.eu/article/14.
- NIST — AI Risk Management Framework (AI RMF 1.0) and Generative AI Profile (NIST AI 600-1). nist.gov/itl/ai-risk-management-framework.
- Infocomm Media Development Authority (Singapore) — Model AI Governance Framework for Agentic AI, Version 1.0 (January 2026). imda.gov.sg.



Let me know your thoughts