…
Skip to content
Topics
On this page

Chain of Thought Prompting

Chain of thought prompting (CoT prompting) is a prompting technique that instructs a large language model to write intermediate reasoning steps before stating its final answer. Generating the reasoning explicitly gives the model more tokens to work through a problem, which usually improves accuracy on multi-step tasks such as debugging and log analysis.

  • Intermediate steps: The model produces a sequence of reasoning statements, and each statement builds on the conclusions of the previous one.
  • Explicit trigger: A short instruction such as "think step by step", or a demonstration with visible reasoning, activates the behaviour, so the method is also called step by step prompting.
  • Separate final answer: The prompt usually requests a clearly labelled answer line so that application code can extract the result programmatically.
  • Zero-shot or few-shot: It works with instructions alone or with examples, two forms described in types of prompting.
  • Research origin: The technique was described in a 2022 research paper from Google on reasoning in language models.
Chain of thought prompting on an error logError log lines go to the LLM with a prompt that says to think step by step. The model writes three reasoning steps: the earliest error, the callers that time out, and the warning that follows. It then writes the final answer, payments-db. A direct prompt without steps goes straight from the log to an answer. The log and steps are illustrative.Inputerror.log+ think stepby stepLLMStep 1: earliest ERRORStep 2: callers time outStep 3: warning followsAnswer: payments-dbDirect prompt:no steps shown
Chain of thought prompting on an error log

For example, asking which service failed first in an error log produces a more reliable answer when the model first lists the timestamps and dependencies it observes.

Key Characteristics of Chain of Thought Prompting

  • Sequential generation: Because a large language model generates one token at a time, earlier reasoning tokens become context for the later answer.
  • Task dependency: It helps most on arithmetic, logic, debugging and multi-stage analysis, and offers little benefit for simple lookups.
  • Longer outputs: The reasoning adds output tokens, which increases latency and cost for every request.
  • Inspectable reasoning: Developers can examine the individual steps to diagnose exactly where an incorrect answer originated.
  • Imperfect faithfulness: The written explanation does not always reflect the internal computation that actually produced the answer.
  • Built-in reasoning models: Some newer models perform extended reasoning internally before responding, which reduces the need for an explicit instruction.

How Chain of Thought Prompting Works

  1. State the problem: Provide the task description and the complete input, such as the relevant log entries or the failing test output.
  2. Request reasoning: Add an explicit instruction to reason step by step, or include a demonstration whose answer displays each intermediate step.
  3. Define the step format: Request numbered reasoning lines so that the structure is consistent across responses and simple to review.
  4. Separate the answer: Require a single labelled answer line at the end, which application code can parse without interpreting the reasoning.
  5. Extract and verify: Parse the answer, validate it against the original input, and store the reasoning in logs for later debugging.

Example: Parsing a Chain of Thought Response in Python

The program below builds a chain of thought prompt for log triage, then extracts the steps and the answer from a hard-coded sample response.

Python
# Build a chain of thought prompt for log triage, then parse the reply
logs = """10:02:11 ERROR payments-db  connection refused on port 5432
10:02:14 ERROR checkout-api  timeout calling payments-db
10:02:15 WARN  web-frontend  checkout request took 5012 ms
10:02:19 ERROR checkout-api  timeout calling payments-db"""

prompt = (
    "Find the service that failed first in these logs.\n"
    "Think step by step. Write each step as 'Step N:'.\n"
    "End with one line: 'Answer: <service name>'.\n\n" + logs
)

# Sample model response, hard-coded for illustration (no model is called)
sample_response = """Step 1: The earliest ERROR is at 10:02:11 from payments-db.
Step 2: checkout-api errors start at 10:02:14 and name payments-db.
Step 3: The frontend warning is a slow request caused by those timeouts.
Answer: payments-db"""

lines = sample_response.splitlines()
steps = [line for line in lines if line.startswith("Step ")]
answers = [line.split(":", 1)[1].strip() for line in lines if line.startswith("Answer:")]

print("Prompt lines:", len(prompt.splitlines()))
print("Reasoning steps:", len(steps))
print("Final answer:", answers[0] if answers else "missing")
print("Answer appears in logs:", bool(answers) and answers[0] in logs)
Output
Prompt lines: 8
Reasoning steps: 3
Final answer: payments-db
Answer appears in logs: True
  • Structured reasoning: Numbered step lines make the reasoning countable, so an empty or missing chain can be detected automatically.
  • Parseable answer: The labelled answer line separates the result from the reasoning, so downstream code never needs to interpret unstructured text.
  • Basic verification: Checking that the answer appears in the logs catches a common failure, an invented service name that is a form of LLM hallucination.

Applications of Chain of Thought Prompting

  • Incident triage: Identifying the first failing service by correlating timestamps across logs from several components.
  • Debugging: Explaining why a failing unit test produces its error message before proposing a specific fix.
  • Code review: Evaluating boundary conditions and edge cases in a function before writing a review comment.
  • Query analysis: Explaining why an SQL query is slow before recommending an index.
  • Agent planning: Deciding the next tool call in ReAct agents, which alternate reasoning and actions.
  • Capacity estimation: Calculating infrastructure capacity or monthly cost from several numeric inputs.

Advantages

  • Higher accuracy: It improves results on problems that require several dependent reasoning steps.
  • Transparency: Reviewers can identify the exact step where the reasoning went wrong.
  • Simple implementation: It requires only an additional instruction or a single reasoning demonstration.
  • Compatibility: It combines easily with role prompts and structured output requirements.

Limitations

  • Additional cost: Reasoning tokens increase the latency and the price of every individual request.
  • Plausible errors: A fluent chain can still contain one incorrect step that produces a confident but incorrect answer.
  • Limited benefit: Straightforward classification or extraction tasks rarely show measurable improvement.
  • Parsing risk: The model occasionally omits the answer line, so application code must handle a missing answer.
  • Exposure risk: Reasoning text shown to end users can reveal internal instructions or sensitive context.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. What does chain of thought prompting ask the model to produce?

Frequently Asked Questions

What is chain of thought prompting in simple words?

It is asking a language model to show its reasoning step by step before giving the final answer. Writing out the steps helps the model handle problems that need several dependent decisions.

What is zero-shot chain of thought?

Zero-shot chain of thought adds an instruction such as "think step by step" without any worked examples. Few-shot chain of thought instead includes examples whose answers show the reasoning.

Does chain of thought prompting always improve accuracy?

No. It helps most on multi-step tasks such as debugging, calculations and log analysis. On simple lookups or classification it adds cost with little or no gain.

Is the reasoning in chain of thought reliable?

Not always. The written steps can look correct while containing an error, and they may not match how the model actually reached its answer. Important answers still need verification.