What is Prompt Engineering
Prompt engineering is the practice of designing the instructions, context and examples given to a large language model so that it produces accurate, consistent and usable output. It treats the prompt as a specification that can be evaluated and refined systematically, in the same way that application code is tested and improved.
- Prompt: The complete input sent to a large language model, including instructions, input data and optional examples.
- Instruction design: It defines the task, the intended audience and the expected result in specific, unambiguous language.
- Context selection: It supplies the information the model needs, such as a code diff, a log extract or a team style guide.
- Output specification: It defines the response format, for example a table, a numbered list or structured output in JSON.
- Iterative evaluation: It modifies one component of the prompt at a time and measures the effect on a fixed set of test inputs.
For example, a prompt requesting code review comments on a diff, limited to three comments with a line number and a severity category, produces more usable output than the instruction "Review this code."
Key Characteristics of Prompt Engineering
- No model modification: It influences behaviour at inference time, so the model parameters remain unchanged, unlike fine-tuning.
- Model dependency: A prompt optimised for one model can behave differently on another model or on a newer version of the same model.
- Probabilistic results: Identical prompts can generate different outputs, particularly at a higher temperature setting.
- Token consumption: Every instruction and example consumes tokens, which increases the cost and latency of each request.
- Testability: A prompt can be evaluated against a fixed collection of inputs and expected outputs, similar to a unit test suite.
- Layered structure: Production applications separate a fixed system prompt from a variable user prompt, as described in system prompt vs user prompt.
How Prompt Engineering Works
The steps below show how to write prompts for LLM applications in a repeatable, testable way.
- Define the task: Document the input, the desired output and the criteria that identify a correct response.
- Assign a role: Give the model a role, such as "senior Python reviewer", to establish the vocabulary and the level of technical detail.
- Provide context: Include only the information the task requires, such as the diff, the relevant error log or the applicable section of a style guide.
- Specify the format: Describe the exact output structure, for example one line per comment in the form
line | severity | comment. - Add examples when necessary: Include input and output pairs when the format or the categories are difficult to describe, a technique covered in types of prompting.
- Evaluate and refine: Run the prompt on representative inputs, analyse the failures and modify one component at a time.
Example: Building a Code Review Prompt in Python
The program below assembles a code review prompt from named components, verifies that no required component is missing and compares its length with a vague prompt.
# Build a code review prompt from named parts, then check that none is missing
parts = {
"role": "You are a senior Python reviewer for a payments service.",
"task": "Review the diff below and write at most three review comments.",
"context": "The function issues customer refunds. The team follows PEP 8.",
"format": "Write each comment as: line number | severity | comment.",
"constraints": "Ignore formatting. Report bugs and security issues first.",
}
diff = """+def refund(order, amount):
+ if amount > order.total:
+ amount = order.total
+ db.execute(f"UPDATE orders SET refunded={amount} WHERE id={order.id}")"""
REQUIRED = ["role", "task", "context", "format"]
missing = [name for name in REQUIRED if not parts.get(name)]
prompt = "\n".join(parts.values()) + "\n\nDiff:\n" + diff
vague = "Review this code.\n\n" + diff
print("Missing parts:", missing or "none")
print("Vague prompt words:", len(vague.split()))
print("Engineered prompt words:", len(prompt.split()))
print("---")
print(prompt)Missing parts: none
Vague prompt words: 22
Engineered prompt words: 69
---
You are a senior Python reviewer for a payments service.
Review the diff below and write at most three review comments.
The function issues customer refunds. The team follows PEP 8.
Write each comment as: line number | severity | comment.
Ignore formatting. Report bugs and security issues first.
Diff:
+def refund(order, amount):
+ if amount > order.total:
+ amount = order.total
+ db.execute(f"UPDATE orders SET refunded={amount} WHERE id={order.id}")- Named components: Keeping the role, task, context, format and constraints separate makes each component easier to evaluate and modify independently.
- Specificity over brevity: The engineered prompt is longer, but it tells the model which issues to prioritise, how many comments to write and how to format each one.
- Machine-readable format: A fixed line, severity and comment structure allows a script to parse each comment and post it automatically to the pull request.
Applications of Prompt Engineering
Common prompt engineering examples in software teams include the following tasks.
- Code review: Generating review comments on pull requests with a consistent severity classification.
- Log summarisation: Condensing thousands of error log entries into a concise incident summary.
- Data extraction: Converting unstructured bug reports into JSON fields for an issue tracker.
- Test generation: Producing unit test cases from a function signature and its documentation.
- Documentation: Drafting API reference pages and changelog entries from source code.
- Retrieval applications: Instructing the model to answer only from passages supplied by RAG.
Advantages
- Rapid iteration: A modified prompt takes effect on the next request, without retraining or redeploying the model.
- Low infrastructure cost: It requires no training data, specialised hardware or model hosting.
- Transferable technique: The same principles apply across most instruction-following models from different providers.
- Improved reliability: Explicit formats and constraints reduce irrelevant and unparseable responses.
- Transparency: The prompt is readable text, so reviewers can audit exactly what the model received.
Limitations
- No new knowledge: A prompt cannot add information the model lacks, beyond the context supplied with each request.
- Persistent hallucination: Careful prompting reduces, but does not eliminate, LLM hallucination.
- Sensitivity: Minor wording changes or a model upgrade can alter results, so prompts require regression testing.
- Context limitations: Lengthy instructions and examples compete for capacity in the context window.
- Security vulnerability: Untrusted input inside a prompt can override instructions, an attack known as prompt injection.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. What does prompt engineering change?
Frequently Asked Questions
What is prompt engineering in simple terms?
Prompt engineering is writing and testing the instructions given to an AI model so that it produces the output a task needs. It covers the wording, the context supplied, the examples and the required output format.
Is prompt engineering a real job skill?
It is a common skill in roles that build LLM features, such as AI engineers and backend developers. It is usually combined with evaluation, retrieval and software engineering rather than practised on its own.
What is the difference between prompt engineering and fine-tuning?
Prompt engineering changes the input sent to a model and leaves the model unchanged. Fine-tuning changes the model's parameters by training it further on examples. Prompting is faster to change, while fine-tuning can fix behaviour that prompts cannot.
Do prompts work the same on every model?
No. Models differ in training and instruction following, so a prompt that works well on one model may need changes on another. Teams keep a test set of inputs and rerun it whenever the model or the prompt changes.
What is a prompt and what makes a good prompt?
A prompt is the complete input sent to a model, including instructions, context and examples. A good prompt states the task, gives the needed context, specifies the output format and sets clear limits.
Related Articles
- Types of Prompting (Zero-shot, Few-shot)Compare the types of prompting: zero-shot, one-shot, few-shot, chain of thought, role prompting and prompt chaining, with a Python few-shot prompt builder.
- Chain of Thought PromptingLearn chain of thought prompting: how step-by-step reasoning improves LLM accuracy and how to parse the final answer, with a Python log triage example.
- System Prompt vs User PromptSystem prompt vs user prompt: who writes each, what belongs in each, how models prioritise them, with a comparison table and a Python code review example.
- Structured Output (JSON) from LLMsLearn how structured output gets JSON from an LLM: schemas, extraction prompts, constrained decoding and validation, with a Python bug report JSON check.
- What is Context EngineeringLearn what context engineering is: selecting, ranking and fitting information into an LLM context window, with a Python token budget example and limits.
- What is a Large Language Model (LLM)Learn what a large language model is, how an LLM predicts the next token, its key characteristics, uses and limits, with a Python example and diagram.