Human in the loop is a design in which a person participates in an AI system’s decision or execution process by reviewing, approving, correcting or providing input at a defined stage. Human involvement may occur before consequential actions, during execution or when reviewing results, depending on the system’s purpose and risk.
The phrase predates generative AI. It became widely used in debates about autonomous weapons, where a 2012 Human Rights Watch report distinguished systems that act only with a human command, systems that act under a human operator who can override them, and systems that act without any human input. Those three positions, in the loop, on the loop and out of the loop, still describe the choices agent builders face.
Why it matters more for agents
A chatbot that gets something wrong produces a wrong answer, and a person reading it is already in the loop. An agent that gets something wrong takes a wrong action, often before anyone reads anything. As agents take more steps on their own, the question of where a person participates stops being implicit and has to be designed.
Agent tooling now builds it in. The Model Context Protocol specification says there should always be a human in the loop with the ability to deny tool invocations, and that clients should present confirmation prompts for operations. OpenAI’s Agents SDK can pause a run until a person approves or rejects a sensitive tool call, either always or according to a rule decided per call. Claude Code’s permission settings let each tool be allowed, set to ask for confirmation, or denied.
Regulation points the same way. Article 14 of the EU AI Act requires that high-risk AI systems be designed so that people can effectively oversee them while in use, including the ability to override or reverse an output and to interrupt the system with a stop button or similar procedure.
The attention problem
The same Article 14 contains a warning that matters as much as the requirement. It asks that the people overseeing a system remain aware of the tendency to rely, or over-rely, on its output: automation bias.
The research behind that warning is long-standing. Parasuraman and Manzey’s review of automation complacency and bias found that people make both kinds of error when an automated aid is imperfect, missing problems it did not flag and accepting recommendations they should have questioned. They found it in experts as well as novices, and concluded it cannot be prevented by training or instructions alone.
Clinical medicine shows what happens at scale. The US Agency for Healthcare Research and Quality describes alert fatigue as clinicians becoming desensitized to safety alerts and ignoring or failing to respond to them, and notes that clinicians override the vast majority of medication-ordering warnings, even critical ones. One study it cites counted 187 warnings per patient per day in an intensive care unit.
An approval step that fires on every action stops being oversight. A person asked to approve hundreds of tool calls learns to approve them, and the one that mattered goes through with the rest.
Designing where the human goes
The practical question is not whether to have a human in the loop, but where their attention will actually change an outcome. A few principles follow from the evidence.
Put approval where consequence is. A request to read a file and a request to delete a production database should not reach a person the same way. Designating high-consequence operations, the costly or irreversible ones, tells a system where an approval is worth its cost, and lets everything else flow.
Give the person something to judge. An approval prompt that shows only a tool name and arguments asks a person to reconstruct the context. One that shows what the agent was asked to do, and how this action relates to it, asks for a judgment.
Use review after the fact for the rest. Not every action needs a decision before it runs. A person “on the loop” can review what happened, provided there is a record worth reviewing. Findings that separate what was noticed from which rules were broken, as flags and violations do, let that review start where it matters.
Measure the loop itself. Approval rates near one hundred percent, approval times of a second or two, and long queues are signs that the loop has become a reflex.
An example
The following is illustrative.
A team gives an infrastructure agent approval prompts on every command. In the first week, reviewers read each one. By the third week, they approve about 400 a day and spend under two seconds on each. The team changes the design: routine reads and in-scope changes run without approval and are recorded; destructive changes to production, and any action outside the agent’s declared scope, ask a person first and show the task alongside the command. Prompts fall to a handful a day, and reviewers start rejecting some of them again.
How Sentience Governor relates
Sentience Governor does not provide an approval step; it records without blocking. It supports the other half of human oversight: the record a person reviews. Each attempted action is evaluated against the session’s declared intent and scope and the operator’s governance profile, and the findings, including high-consequence designations, are kept as evidence. That gives a person on the loop a place to start, and gives anyone designing approval steps a measure of where they would actually be needed.
Where this is heading
Emerging approaches. Approval rules decided per call rather than per tool, as agent frameworks now allow. Consequence-tiered autonomy, where an agent works freely on reversible actions and asks on irreversible ones. And automated reviewers working alongside people: Claude Code, for example, documents a mode in which a classifier reviews actions instead of the user.
Open questions. How to tell whether a human reviewer is still exercising judgment. Who is accountable when an approved action causes harm. And how human oversight should work in multi-agent systems, where no single prompt shows the whole picture.
Often confused with
Human approval. Approval before an action runs is one form of human in the loop, the most visible one. Review, correction and supplying input are others.
Human on the loop. On the loop, a person supervises and can intervene, but actions do not wait for them. It trades prevention for scale, and depends on a good record and a working way to stop the system.
Where the concept stops
A human in the loop establishes that a person had an opportunity to intervene at a defined point. It does not establish that they exercised judgment, that they had the context to do so, or that their decision was right. The value of the loop lies in what the person sees and how often they are asked.