…
Skip to content
Topics
On this page

ReAct Agent Explained

A ReAct agent is an AI agent that alternates between reasoning and acting: it writes a short thought about the current situation, executes one tool action, and reads the resulting observation before deciding the next step. The name combines Reasoning and Acting, and the pattern grounds each decision in real evidence instead of a single unverified guess.

  • Thought: A brief reasoning statement that interprets the latest evidence and selects the next action.
  • Action: One tool invocation with specific arguments, produced through tool calling.
  • Observation: The tool output returned to the model, such as a test failure, a file excerpt or a commit message.
  • Trajectory: The accumulated sequence of thoughts, actions and observations that forms the agent's working record.
  • Stopping condition: A final answer, a completed action or a maximum iteration count that terminates the loop.
The ReAct loop of thought, action and observationThe goal, find why test_total fails, leads to Thought. Thought leads to Action, where the agent calls one of run_tests, read_file, git_log or open_ticket. Action leads to Observation. An arrow labelled next thought returns from Observation to Thought. When the evidence is sufficient, the loop exits to a final answer and ticket OPS-482. The ticket number is illustrative.Goalwhy test_total failsThoughtActionrun_testsread_filegit_logopen_ticketObservationnext thoughtFinal answerticket OPS-482
The ReAct loop of thought, action and observation

For example, a coding agent investigating a failing unit test reruns the test, reads the relevant source file, inspects the commit history and then opens a ticket for the responsible owner.

Key Characteristics of a ReAct Agent

  • Interleaved reasoning: Reasoning and execution alternate, unlike chain of thought prompting, which reasons completely before answering.
  • Evidence-driven: Every thought responds to the most recent observation, which reduces unsupported assumptions.
  • Incremental: The agent commits to only one action at a time, so it adapts immediately to unexpected results.
  • Transparent: The trajectory provides a readable audit trail that explains how the agent reached its conclusion.
  • Context-dependent: The growing trajectory occupies tokens, so long investigations need memory in AI agents to summarise earlier steps.
  • Published origin: The pattern was introduced in a 2022 research paper by Yao and colleagues as ReAct prompting, with example trajectories placed in the prompt.

How a ReAct Agent Works

  1. Receive the goal: The agent obtains a task, such as determining why a particular test fails in continuous integration.
  2. Generate a thought: The model analyses the goal and the trajectory so far, then describes what information is missing.
  3. Select an action: The model chooses one tool and its arguments, for example executing the failing test.
  4. Execute the action: Application code runs the tool, enforcing permissions and a timeout.
  5. Record the observation: The tool output is appended to the trajectory and returned to the model.
  6. Repeat or finish: The cycle continues until the model produces a final answer or reaches the iteration limit.

Example: A ReAct Agent Loop in Python

The program below prints the Thought, Action and Observation lines for a failing test investigation, with a rule-based policy standing in for the LLM.

Python
# Sample tool outputs, hard-coded for illustration (nothing is executed)
TOOLS = {
    "run_tests": lambda arg: "FAILED test_orders.py::test_total - assert 105.0 == 100.0",
    "read_file": lambda arg: "def total(items): return sum(i.price for i in items) * 1.05",
    "git_log": lambda arg: "a1f3c2 'Add 5% service fee to total()' by ravi, 2 hours ago",
    "open_ticket": lambda arg: f"Created ticket OPS-482: {arg}",
}

def policy(goal, history):
    """Rule-based stand-in for the LLM: picks the next thought and action."""
    last = history[-1][2] if history else ""
    if not history:
        return "Reproduce the failure first.", "run_tests", "tests/test_orders.py"
    if "assert" in last:
        return "The total is 5 percent too high. Check the code.", "read_file", "orders/pricing.py"
    if "1.05" in last:
        return "A multiplier was added. Find the change.", "git_log", "orders/pricing.py"
    if "service fee" in last:
        return "Intentional change; the test is stale. Ask the owner.", "open_ticket", "Update test_total for fee in a1f3c2"
    return "Done.", None, None

goal = "Find why test_total fails"
history = []
for step in range(1, 6):
    thought, action, arg = policy(goal, history)
    print(f"Thought {step}: {thought}")
    if action is None:
        break
    observation = TOOLS[action](arg)
    print(f"Action {step}: {action}({arg!r})")
    print(f"Observation {step}: {observation}")
    history.append((thought, action, observation))
    if action == "open_ticket":
        print("Final answer: test is stale after commit a1f3c2; ticket OPS-482 opened.")
        break
Output
Thought 1: Reproduce the failure first.
Action 1: run_tests('tests/test_orders.py')
Observation 1: FAILED test_orders.py::test_total - assert 105.0 == 100.0
Thought 2: The total is 5 percent too high. Check the code.
Action 2: read_file('orders/pricing.py')
Observation 2: def total(items): return sum(i.price for i in items) * 1.05
Thought 3: A multiplier was added. Find the change.
Action 3: git_log('orders/pricing.py')
Observation 3: a1f3c2 'Add 5% service fee to total()' by ravi, 2 hours ago
Thought 4: Intentional change; the test is stale. Ask the owner.
Action 4: open_ticket('Update test_total for fee in a1f3c2')
Observation 4: Created ticket OPS-482: Update test_total for fee in a1f3c2
Final answer: test is stale after commit a1f3c2; ticket OPS-482 opened.
  • Observation-driven choices: Each rule inspects only the latest observation, just as an LLM would condition its next thought on the newest evidence.
  • Iteration limit: The loop permits at most five steps, a simple safeguard against endless cycles.
  • Appropriate outcome: The agent identifies an intentional pricing change and requests a test update instead of reverting production code.

Applications of a ReAct Agent

  • Test failure triage: Reproducing failures, reading source code and identifying the responsible commit.
  • Log investigation: Querying logs and metrics iteratively until an anomaly is explained.
  • Documentation lookup: Searching API references repeatedly until a specific question is answered.
  • Infrastructure diagnostics: Inspecting container status, configuration values and recent deployments.
  • Framework implementations: Serving as the default agent loop in libraries covered in the LangGraph tutorial.

Advantages

  • Adaptability: The agent revises its direction after every observation instead of following a fixed script.
  • Reduced hallucination: Conclusions rest on tool outputs rather than on the model's unsupported recall.
  • Interpretability: Engineers can review each thought and action to diagnose incorrect behaviour.
  • Simplicity: The loop requires only a model, a tool registry and a stopping rule.

Limitations

  • Short horizon: Choosing one step at a time can miss an efficient overall strategy, which planning in AI agents addresses.
  • Latency and cost: Every iteration requires another model invocation and a larger prompt.
  • Repetitive loops: Agents can repeat the same unproductive action without explicit detection.
  • Evaluation difficulty: Many valid trajectories exist, which complicates AI agent evaluation.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. What is the correct order of one ReAct iteration?

Frequently Asked Questions

What does ReAct stand for in AI agents?

ReAct stands for Reasoning and Acting. The agent writes a reasoning step, performs one action with a tool and reads the result before reasoning again.

How is a ReAct agent different from chain of thought?

Chain of thought reasons through a problem in one response without external input. A ReAct agent interleaves reasoning with tool actions, so each step can use new evidence from the environment.

When should a ReAct agent stop?

It stops when the model returns a final answer, when a terminal action such as opening a ticket succeeds, or when an iteration limit is reached. A limit is necessary because a model can repeat actions indefinitely.

Is ReAct better than plan-and-execute?

Neither is better in every case. ReAct adapts well to exploratory tasks with uncertain paths, while plan-and-execute suits tasks whose steps and dependencies can be known in advance.