Chain of Thought Prompting
Chain of thought prompting (CoT prompting) is a prompting technique that instructs a large language model to write intermediate reasoning steps before stating its final answer. Generating the reasoning explicitly gives the model more tokens to work through a problem, which usually improves accuracy on multi-step tasks such as debugging and log analysis.
- Intermediate steps: The model produces a sequence of reasoning statements, and each statement builds on the conclusions of the previous one.
- Explicit trigger: A short instruction such as "think step by step", or a demonstration with visible reasoning, activates the behaviour, so the method is also called step by step prompting.
- Separate final answer: The prompt usually requests a clearly labelled answer line so that application code can extract the result programmatically.
- Zero-shot or few-shot: It works with instructions alone or with examples, two forms described in types of prompting.
- Research origin: The technique was described in a 2022 research paper from Google on reasoning in language models.
For example, asking which service failed first in an error log produces a more reliable answer when the model first lists the timestamps and dependencies it observes.
Key Characteristics of Chain of Thought Prompting
- Sequential generation: Because a large language model generates one token at a time, earlier reasoning tokens become context for the later answer.
- Task dependency: It helps most on arithmetic, logic, debugging and multi-stage analysis, and offers little benefit for simple lookups.
- Longer outputs: The reasoning adds output tokens, which increases latency and cost for every request.
- Inspectable reasoning: Developers can examine the individual steps to diagnose exactly where an incorrect answer originated.
- Imperfect faithfulness: The written explanation does not always reflect the internal computation that actually produced the answer.
- Built-in reasoning models: Some newer models perform extended reasoning internally before responding, which reduces the need for an explicit instruction.
How Chain of Thought Prompting Works
- State the problem: Provide the task description and the complete input, such as the relevant log entries or the failing test output.
- Request reasoning: Add an explicit instruction to reason step by step, or include a demonstration whose answer displays each intermediate step.
- Define the step format: Request numbered reasoning lines so that the structure is consistent across responses and simple to review.
- Separate the answer: Require a single labelled answer line at the end, which application code can parse without interpreting the reasoning.
- Extract and verify: Parse the answer, validate it against the original input, and store the reasoning in logs for later debugging.
Example: Parsing a Chain of Thought Response in Python
The program below builds a chain of thought prompt for log triage, then extracts the steps and the answer from a hard-coded sample response.
# Build a chain of thought prompt for log triage, then parse the reply
logs = """10:02:11 ERROR payments-db connection refused on port 5432
10:02:14 ERROR checkout-api timeout calling payments-db
10:02:15 WARN web-frontend checkout request took 5012 ms
10:02:19 ERROR checkout-api timeout calling payments-db"""
prompt = (
"Find the service that failed first in these logs.\n"
"Think step by step. Write each step as 'Step N:'.\n"
"End with one line: 'Answer: <service name>'.\n\n" + logs
)
# Sample model response, hard-coded for illustration (no model is called)
sample_response = """Step 1: The earliest ERROR is at 10:02:11 from payments-db.
Step 2: checkout-api errors start at 10:02:14 and name payments-db.
Step 3: The frontend warning is a slow request caused by those timeouts.
Answer: payments-db"""
lines = sample_response.splitlines()
steps = [line for line in lines if line.startswith("Step ")]
answers = [line.split(":", 1)[1].strip() for line in lines if line.startswith("Answer:")]
print("Prompt lines:", len(prompt.splitlines()))
print("Reasoning steps:", len(steps))
print("Final answer:", answers[0] if answers else "missing")
print("Answer appears in logs:", bool(answers) and answers[0] in logs)Prompt lines: 8
Reasoning steps: 3
Final answer: payments-db
Answer appears in logs: True- Structured reasoning: Numbered step lines make the reasoning countable, so an empty or missing chain can be detected automatically.
- Parseable answer: The labelled answer line separates the result from the reasoning, so downstream code never needs to interpret unstructured text.
- Basic verification: Checking that the answer appears in the logs catches a common failure, an invented service name that is a form of LLM hallucination.
Applications of Chain of Thought Prompting
- Incident triage: Identifying the first failing service by correlating timestamps across logs from several components.
- Debugging: Explaining why a failing unit test produces its error message before proposing a specific fix.
- Code review: Evaluating boundary conditions and edge cases in a function before writing a review comment.
- Query analysis: Explaining why an SQL query is slow before recommending an index.
- Agent planning: Deciding the next tool call in ReAct agents, which alternate reasoning and actions.
- Capacity estimation: Calculating infrastructure capacity or monthly cost from several numeric inputs.
Advantages
- Higher accuracy: It improves results on problems that require several dependent reasoning steps.
- Transparency: Reviewers can identify the exact step where the reasoning went wrong.
- Simple implementation: It requires only an additional instruction or a single reasoning demonstration.
- Compatibility: It combines easily with role prompts and structured output requirements.
Limitations
- Additional cost: Reasoning tokens increase the latency and the price of every individual request.
- Plausible errors: A fluent chain can still contain one incorrect step that produces a confident but incorrect answer.
- Limited benefit: Straightforward classification or extraction tasks rarely show measurable improvement.
- Parsing risk: The model occasionally omits the answer line, so application code must handle a missing answer.
- Exposure risk: Reasoning text shown to end users can reveal internal instructions or sensitive context.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. What does chain of thought prompting ask the model to produce?
Frequently Asked Questions
What is chain of thought prompting in simple words?
It is asking a language model to show its reasoning step by step before giving the final answer. Writing out the steps helps the model handle problems that need several dependent decisions.
What is zero-shot chain of thought?
Zero-shot chain of thought adds an instruction such as "think step by step" without any worked examples. Few-shot chain of thought instead includes examples whose answers show the reasoning.
Does chain of thought prompting always improve accuracy?
No. It helps most on multi-step tasks such as debugging, calculations and log analysis. On simple lookups or classification it adds cost with little or no gain.
Is the reasoning in chain of thought reliable?
Not always. The written steps can look correct while containing an error, and they may not match how the model actually reached its answer. Important answers still need verification.
Related Articles
- Types of Prompting (Zero-shot, Few-shot)Compare the types of prompting: zero-shot, one-shot, few-shot, chain of thought, role prompting and prompt chaining, with a Python few-shot prompt builder.
- What is Prompt EngineeringLearn what prompt engineering is, how to design a prompt step by step, its key characteristics, uses and limits, with a Python code review prompt example.
- Structured Output (JSON) from LLMsLearn how structured output gets JSON from an LLM: schemas, extraction prompts, constrained decoding and validation, with a Python bug report JSON check.
- ReAct Agent ExplainedLearn how a ReAct agent alternates thought, action and observation, with a Python example that traces a failing unit test to a commit and opens a ticket.
- Planning and Reasoning in AI AgentsLearn planning and reasoning in AI agents: task decomposition, dependencies, plan-and-execute and replanning, with a Python flaky test fix example.
- LLM Hallucination: Causes and FixesLearn what LLM hallucination is, why models invent facts and code APIs, and how to reduce it with RAG and validation, with a runnable Python checker.