…
Skip to content
Topics
On this page

Planning and Reasoning in AI Agents

Planning and reasoning in AI agents is the process of decomposing a goal into an ordered set of smaller steps, identifying the dependencies between them, and executing them while monitoring each result. When a step fails or produces unexpected information, the agent revises the remaining plan instead of abandoning the goal or repeating the same unsuccessful action.

  • Task decomposition: Dividing a broad objective into concrete steps that a single tool or action can complete.
  • Dependencies: Relationships that specify which steps must finish before another step can begin.
  • Plan-and-execute: An architecture in which a planner produces the complete plan first and an executor performs each step.
  • Replanning: Modifying the remaining steps after a failure, a new observation or a changed constraint.
  • Reasoning: The model's analysis of the goal, the available tools and the evidence, often expressed as chain of thought prompting.
A plan with dependencies and one replanned stepThe goal is to fix a flaky integration test. The step collect leads to two steps, rerun and compare. The rerun step failed and is shown dashed; an arrow labelled replan leads from it to its replacement, rerun_local. Both rerun_local and compare lead to patch, and patch leads to open_pr. The steps are illustrative.collectrerunfailedreplanrerun_localcomparepatchopen_prGoal: fix flaky integration test
A plan with dependencies and one replanned step

For example, a coding agent assigned to fix a flaky integration test plans to collect CI history, rerun the test, compare logs, patch the code and open a pull request.

Key Characteristics of Planning and Reasoning in AI Agents

  • Goal-oriented: Every step exists to advance a clearly defined objective with a verifiable completion condition.
  • Structured: Plans are often represented as a list or a dependency graph, which application code can inspect and validate.
  • Separation of roles: A planner model designs the steps, while an executor performs them through tool calling.
  • Adaptive: Replanning replaces failed steps with alternatives that still satisfy the original dependencies.
  • Stateful: Completed steps and intermediate results are preserved, frequently through memory in AI agents.
  • Reviewable: A written plan can be approved by an engineer before any risky operation is executed.

How Planning and Reasoning in AI Agents Works

  1. Interpret the goal: The planner analyses the objective, the constraints and the available tools.
  2. Decompose the task: The planner generates discrete steps, each with a description and a list of dependencies.
  3. Validate the plan: Application code confirms that every tool exists and that the dependencies contain no cycles.
  4. Execute ready steps: The executor runs each step whose dependencies are complete and records the outcome.
  5. Detect failures: A failed step, an unexpected result or an exceeded budget triggers a replanning decision.
  6. Replan: The planner substitutes an alternative step or reorders the remaining work, then execution resumes.
  7. Verify completion: The agent confirms the goal is satisfied, for example by checking that the test passes repeatedly.

Example: Plan-and-Execute with Replanning in Python

The program below executes a dependency-ordered plan for a flaky integration test and replaces a step that fails.

Python
# Plan for "fix flaky integration test": step -> (description, dependencies)
plan = {
    "collect": ("Collect the last 20 CI runs of test_checkout_flow", []),
    "rerun": ("Rerun the test 50 times in isolation", ["collect"]),
    "compare": ("Compare logs of passing and failing runs", ["collect"]),
    "patch": ("Add an explicit wait for the payment stub", ["rerun", "compare"]),
    "open_pr": ("Open a pull request with the fix", ["patch"]),
}

# Sample step results, hard-coded for illustration (a real agent runs tools)
OUTCOMES = {"rerun": "fail: CI runner quota exceeded"}

def replan(plan, failed):
    """Replace the failed step with an alternative that reaches the same goal."""
    desc, deps = plan.pop(failed)
    plan["rerun_local"] = ("Rerun the test 50 times in a local container", deps)
    for name, (d, step_deps) in plan.items():
        plan[name] = (d, ["rerun_local" if x == failed else x for x in step_deps])
    return f"replanned: {failed} -> rerun_local"

done = []
while len(done) < len(plan):
    ready = [s for s, (_, deps) in plan.items() if s not in done and all(d in done for d in deps)]
    step = ready[0]
    result = OUTCOMES.get(step, "ok")
    print(f"{step:12} {plan[step][0]:50} {result}")
    if result.startswith("fail"):
        print("  ->", replan(plan, step))
        continue
    done.append(step)
print("Completed order:", " > ".join(done))
Output
collect      Collect the last 20 CI runs of test_checkout_flow  ok
rerun        Rerun the test 50 times in isolation               fail: CI runner quota exceeded
  -> replanned: rerun -> rerun_local
compare      Compare logs of passing and failing runs           ok
rerun_local  Rerun the test 50 times in a local container       ok
patch        Add an explicit wait for the payment stub          ok
open_pr      Open a pull request with the fix                   ok
Completed order: collect > compare > rerun_local > patch > open_pr
  • Dependency ordering: The patch step waits until both the rerun and the comparison are complete.
  • Targeted replanning: Only the failed step is replaced, and every dependency on it is redirected to the new step.
  • Deterministic stand-in: The replan function is a fixed rule here, whereas a production agent would ask the planner model for an alternative.

Applications of Planning and Reasoning in AI Agents

  • Bug fixing: Reproducing a defect, locating the cause, writing a patch and opening a pull request.
  • Dependency upgrades: Updating a library, running tests across services and correcting breaking changes.
  • Incident follow-up: Converting a postmortem into ordered remediation tickets with owners.
  • Infrastructure changes: Sequencing configuration updates, validation checks and staged deployments.
  • Team coordination: Assigning plan steps to specialised workers in multi-agent systems.

Advantages

  • Efficiency: Planning ahead avoids redundant actions that a purely step-by-step ReAct agent might take.
  • Parallelism: Independent steps, such as rerunning tests and comparing logs, can execute concurrently.
  • Oversight: A visible plan allows human-in-the-loop approval before sensitive operations.
  • Resilience: Agent replanning recovers from failures without restarting the entire task.

Limitations

  • Incorrect plans: The planner can omit necessary steps or invent tools that do not exist.
  • Stale assumptions: A plan created early may rely on information that later observations contradict.
  • Additional cost: Planning and replanning require extra model invocations and tokens.
  • Replanning loops: Repeated failures can trigger endless revisions without a limit on replanning attempts.
  • Evaluation complexity: Judging plan quality requires comparing steps, order and outcomes, not only the final answer.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. What does task decomposition produce?

Frequently Asked Questions

What is planning in AI agents?

Planning is the step where an agent breaks a goal into smaller tasks, orders them by their dependencies and decides which tools to use. The agent then executes the plan and revises it when results differ from expectations.

What is the difference between ReAct and plan-and-execute?

A ReAct agent chooses one action at a time based on the latest observation. A plan-and-execute agent writes the full list of steps first, then carries them out and replans only when something fails.

What is replanning in an AI agent?

Replanning is the revision of the remaining steps after a failure or a new finding. It usually replaces or reorders steps while keeping the goal and the completed work unchanged.

Can LLMs plan reliably on their own?

LLMs can produce reasonable plans for familiar engineering tasks, but they sometimes skip steps or assume tools that do not exist. Validating the plan in code and limiting replanning attempts makes the result more dependable.