…
Skip to content
Topics
On this page

Human-in-the-Loop in Agentic AI

Human-in-the-loop (HITL) is a control pattern in agentic AI in which a person must review, approve, correct or reject specific agent actions before they take effect. It provides human oversight of AI agents where errors are costly and keeps an accountable human decision in every high-risk step.

  • Approval gate: A checkpoint that pauses the agent until an authorised person approves or rejects the proposed action.
  • Risk policy: A table that assigns every tool or action a risk level and a handling rule.
  • Escalation: The transfer of an unanswered or uncertain request to another person or a more senior role.
  • Audit log: An append-only record of every agent action, approval, rejection and escalation.
  • Human-on-the-loop: A lighter variant in which the agent acts on its own while a person monitors it and can intervene.
A risk-based approval gate for an incident agentThe agent proposes an action, such as rollback_deploy. A risk policy check sends low risk actions straight to Execute. Medium risk actions in production and all high risk actions go to an approval queue. The on-call approver approves, which leads to Execute, or rejects, which returns feedback to the agent. Requests that stay unanswered are escalated. Every step is written to an audit log. The flow is illustrative.Agent proposesrollback_deployRisk policylowExecutemedium, highApproval queueOn-call approverapprovereject: feedbackEscalateAudit log: every proposal, decision and escalation
A risk-based approval gate for an incident agent

For example, an incident agent may read checkout service logs automatically, but a production rollback of release v2.4.1 waits in an approval queue for the on-call engineer.

Key Characteristics of Human-in-the-Loop

  • Risk-based gating: Approval is required only for actions whose policy level demands it, so routine steps remain automatic.
  • Fail-closed defaults: Unknown or unclassified actions are treated as high risk instead of being executed.
  • Context for reviewers: Each request includes the proposed action, its arguments, the environment and supporting evidence, often gathered through agentic RAG.
  • Enforcement in code: The gate runs in the application layer, so the model cannot bypass it through generated text.
  • Traceability: Every decision is written to an audit log with the actor, time and reason.
  • Time limits: Requests that remain unanswered past a deadline are escalated instead of waiting indefinitely.

How Human-in-the-Loop Works

  1. Policy definition: The team classifies each tool as low, medium or high risk and assigns a handling rule to each level.
  2. Action proposal: The agent requests a tool call through tool calling, with arguments and a stated reason.
  3. Policy check: Application code looks up the risk level and either executes the action or places it in an approval queue.
  4. Human review: An authorised approver inspects the evidence and approves, rejects or modifies the request.
  5. Execution or feedback: Approved actions run, while rejections return to the agent as observations that inform its next plan.
  6. Escalation: Unanswered requests move to a secondary approver after a defined interval.
  7. Recording: Every step and decision is appended to the audit log.
Risk levelExample actionsRule
Lowread_logs, get_metricsExecute automatically
Mediumrestart_pod, scale_replicasAutomatic in staging, approval in production
Highrollback_deploy, change_db_configApproval in every environment

Example: An AI Agent Approval Workflow in Python

The program below applies the risk policy to five proposed incident actions, processes two approval decisions and escalates the remaining request.

Python
# Risk policy: each action has a risk level, and each level has a rule.
POLICY = {
    "read_logs": "low", "get_metrics": "low",
    "restart_pod": "medium", "scale_replicas": "medium",
    "rollback_deploy": "high", "change_db_config": "high",
}
RULES = {"low": "auto", "medium": "auto_if_staging", "high": "approval"}

queue, audit = [], []  # pending approvals and an append-only audit log

def request(step, action, env, reason):
    level = POLICY.get(action, "high")  # unknown actions are treated as high risk
    rule = RULES[level]
    if rule == "auto" or (rule == "auto_if_staging" and env == "staging"):
        decision = "executed"
    else:
        queue.append({"id": step, "action": action, "env": env, "level": level, "reason": reason})
        decision = "queued for approval"
    audit.append((step, "agent", action, env, level, decision))

def review(item_id, approver, approve, note=""):
    item = next(i for i in queue if i["id"] == item_id)
    queue.remove(item)
    decision = "approved, executed" if approve else f"rejected: {note}"
    audit.append((item_id, approver, item["action"], item["env"], item["level"], decision))

def escalate(to):
    for item in queue:  # still unanswered when the approval window closes
        audit.append((item["id"], "system", item["action"], item["env"], item["level"], f"escalated to {to}"))

request(1, "read_logs", "prod", "5xx spike on checkout")
request(2, "get_metrics", "prod", "confirm error rate")
request(3, "restart_pod", "prod", "checkout-7f9 not ready")
request(4, "rollback_deploy", "prod", "errors began after v2.4.1")
request(5, "purge_cache", "prod", "stale assets")
print("Pending approvals:")
for item in queue:
    print(f"  #{item['id']} {item['action']} ({item['env']}, {item['level']}): {item['reason']}")

review(4, "sre.primary", True)
review(5, "sre.primary", False, "not related to the 5xx errors")
escalate("sre.secondary")
print("Audit log:")
for step, actor, action, env, level, decision in audit:
    print(f"  #{step} {actor:<11} {action:<15} {env:<4} {level:<6} {decision}")

Output:

Example
Pending approvals:
  #3 restart_pod (prod, medium): checkout-7f9 not ready
  #4 rollback_deploy (prod, high): errors began after v2.4.1
  #5 purge_cache (prod, high): stale assets
Audit log:
  #1 agent       read_logs       prod low    executed
  #2 agent       get_metrics     prod low    executed
  #3 agent       restart_pod     prod medium queued for approval
  #4 agent       rollback_deploy prod high   queued for approval
  #5 agent       purge_cache     prod high   queued for approval
  #4 sre.primary rollback_deploy prod high   approved, executed
  #5 sre.primary purge_cache     prod high   rejected: not related to the 5xx errors
  #3 system      restart_pod     prod medium escalated to sre.secondary
  • Automatic steps: Read-only diagnostics execute immediately, so the agent gathers evidence without waiting for a person.
  • Fail-closed handling: The unlisted purge_cache action defaults to high risk and is later rejected by the approver.
  • Accountability: The audit log records the agent's proposal, the approver's decision and the escalation as separate entries.

Applications of Human-in-the-Loop

  • Production operations: Approving rollbacks, failovers and configuration changes proposed by incident agents.
  • Software engineering: Reviewing agent-generated pull requests before they are merged.
  • Financial operations: Confirming payments or refunds above a defined threshold.
  • Access management: Approving permission grants that an agent recommends.
  • Data management: Authorising deletions or schema migrations on production databases.
  • Security response: Confirming account lockouts or host isolation during an investigation.

Advantages

  • Reduced impact of errors: Irreversible actions require an explicit human decision before they execute.
  • Clear accountability: Each high-risk change has a named approver recorded in the audit log.
  • Defence in depth: Approval gates stop harmful actions triggered by prompt injection or faulty reasoning.
  • Gradual autonomy: Actions move to lower risk levels as AI agent evaluation shows reliable behaviour.

Limitations

  • Added latency: Gated actions wait for a person, which can extend an outage at night or on holidays.
  • Approval fatigue: Too many requests lead reviewers to approve without reading the evidence.
  • Policy maintenance: New tools must be classified, and outdated levels create gaps in coverage.
  • Reviewer skill: An approver without the relevant context can approve a harmful action.
  • Coordination cost: In multi-agent systems, several agents can compete for the same approver, which the queue design must handle.

Approval gates are one of the agentic design patterns that complement reflection, planning and tool use.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. In the Python example, why is purge_cache queued even though it is not in the policy table?

Frequently Asked Questions

What is human-in-the-loop AI?

Human-in-the-loop AI is a design in which a person reviews, approves, corrects or rejects an AI system's output before it takes effect. In agentic AI it usually applies to actions, such as a deployment or a data change, rather than to text alone.

What is the difference between human-in-the-loop and human-on-the-loop?

In human-in-the-loop designs, the agent waits for a person before a gated action runs. In human-on-the-loop designs, the agent acts on its own while a person monitors it and can stop or reverse its actions.

Which AI agent actions should require human approval?

Actions that are hard to reverse, affect production or customers, move money, change permissions or delete data should require approval. Read-only actions, such as querying logs or metrics, usually run automatically.

Does human-in-the-loop slow down an AI agent?

It adds waiting time to gated actions only, so a well-scoped policy keeps most steps automatic. The agent can prepare evidence while it waits, which shortens the human decision.

What should an AI agent audit log contain?

Each entry should record the step, the actor, the action and its arguments, the environment, the risk level and the decision. For approvals it should also record who approved or rejected the action and the reason given.