← Back

Prompt Engineering for LLMs

2026-04-21 16:34:17

From instruction to optimization: when LLMs need prompt engineering, how effective prompts unlock capability, and a minimal few-shot prompting playbook.

![llm2](https://icuicuicu.icuicu.icu/shared-files/dfa90b1ed19559c8dd89e820aaade7a2.jpg)
## Part 1: Why Do We Need Prompt Engineering in LLM Application?
Large language models (LLMs) possess powerful language understanding and generation capabilities, but they cannot "read minds" — their outputs are highly dependent on the input prompts. Even state-of-the-art models can produce irrelevant, inaccurate, or low-quality results if the prompt is ambiguous, incomplete, or poorly structured. Prompt engineering — the process of designing and optimizing prompts to guide LLMs toward desired outputs — is where we bridge the gap between model capability and practical application.

In this section we'll talk about how zero-shot, few-shot, and chain-of-thought (CoT) prompting differ — and why structured prompts and context engineering make prompt engineering essential for unlocking LLM capability and reliability.

### Zero-Shot Prompting: Intuition
The simplest form of prompt engineering is zero-shot prompting. You provide a task description directly to the LLM, without any examples, and ask it to generate a response. The model relies solely on its pretrained knowledge to complete the task:

Prompt: "Summarize the following paragraph in one sentence: [Paragraph Content]"

This works when the task is simple, common, and aligns with the model’s pretraining data. But it assumes:
- The task is self-explanatory and requires no additional context.
- The model has sufficient pretrained knowledge about the task domain.
- Ambiguity in the task can be resolved by the model’s inherent reasoning.

### Few-Shot Prompting: Guiding with Examples
Few-shot prompting [1] changes the guidance signal. Instead of a single task description, you provide a small number of demonstration examples (input-output pairs) before the target task. The LLM learns the task pattern from these examples and applies it to the new input. A typical few-shot prompt structure is:

Example 1: Input → Output
Example 2: Input → Output
Target: Input → ?

This approach (closely tied to in-context learning; see In-Context Learning for LLMs):
- Pushes the model to follow the task pattern defined by examples.
- Allows adaptation to niche or specialized tasks not well-covered in pretraining.
- Balances guidance with flexibility, letting the model generalize beyond the examples.

For a more advanced variant that enhances few-shot performance, see: Chain-of-Thought Prompting.

### The Common Feeling: “Isn’t This Just Adding Examples?”
Many users ( myself included, at first) find prompt engineering underwhelming. If you squint, it looks like we just:
- Added a few example pairs to the prompt.
- Tweaked the wording of the task description.

Isn’t that just adding extra text, like giving the model a “cheat sheet”?

In practice, yes — the model just processes the entire prompt as input. But conceptually, there are two crucial differences:

**Unit of guidance:**
- Zero-shot: Task-level, the model only gets a description of what to do.
- Few-shot: Instance-level, the model gets concrete examples of how to do the task.

**Generalization:**
- Zero-shot relies entirely on pretrained knowledge, limiting generalization to new tasks.
- Few-shot teaches the model a task pattern on the fly, enabling generalization to unseen inputs.

It’s a subtle but important shift: from asking the model to “know” the task to teaching it to “learn” the task.

### Why the Distinction Matters in Complex Tasks
This difference is especially sharp for complex tasks like reasoning, translation, or structured output generation.

- Zero-shot setup: You ask the model to solve a complex problem with only a task description. If the task is niche (e.g., custom data formatting), the model will likely fail.
- Few-shot setup: You provide 2-5 examples of the complex task, showing the model the exact pattern. The model can then apply that pattern to new inputs, even if it has no pretrained knowledge of the specific task.

Formally, in-context learning allows the model to update its implicit task understanding based on the examples in the prompt [2], without changing its parameters:

f_θ(prompt + input) = f_θ(examples + task + input)

Where f_θ is the LLM’s forward pass, and the examples modify the model’s output distribution for the target input.

This is why few-shot prompting can train models to perform custom tasks without any fine-tuning — it leverages the model’s in-context learning capability.

### CoT Prompting: Unlocking Reasoning with Structure
When a task requires multi-step reasoning (e.g., math problems, logical deduction), we can use chain-of-thought (CoT) prompting [3].

Instead of just providing input-output examples, you add a step-by-step reasoning process to each example. The model learns to generate its own reasoning chain before producing the final answer.

- The model proposes a reasoning process; the quality of the reasoning is reflected in the correctness of the final answer.
- The guidance is structured, interpretable, and directly tied to task success.

This is a natural fit for reasoning tasks like math, logic, and problem-solving, where the path to the answer is as important as the answer itself. CoT prompting avoids the “black box” output of zero-shot/few-shot and makes the model’s reasoning process transparent.

### Intuition
- Zero-shot: “Do this task.”
- Few-shot: “Do this task like these examples.”
- CoT: “Do this task like these examples, and show your work.”

### Conclusion
Prompt engineering in LLM application can feel, at first, like we’re just adding extra text to the input. But the shift in guidance level and the introduction of structured reasoning are what make it different from zero-shot prompting. For simple tasks, zero-shot is sufficient. For niche tasks, few-shot unlocks generalization. For reasoning tasks, CoT shows the full potential: models can generate interpretable reasoning chains and produce more reliable, accurate outputs.

That’s why prompt engineering — in one form or another — remains central to shaping how models not only generate text, but also reason and solve complex tasks.

## Part 2: How to Design Effective Prompts for LLMs
This is a recap of Lecture 18 of Stanford CS336 (Spring 2025), which covered prompt engineering for language models with a focus on few-shot and chain-of-thought prompting. If you find this interesting, I highly recommend going through the full lecture materials:
- Course website
- YouTube lecture playlist

### Prompt Engineering Setup in the LLM Context
- Context (c): the prompt plus any examples or reasoning steps provided
- Task (t): the specific task the model is asked to perform (e.g., summarize, translate, solve math)
- Input (x): the target input the model needs to process
- Output (y): the desired response from the model
- Prompt (P): the combination of context, task description, and input (P = c + t + x)

Unlike fine-tuning, prompt engineering modifies the input (prompt) instead of the model parameters, which makes it flexible and low-cost.

### Naive Prompting: Learn from Basic Instructions
The simplest prompt engineering approach is naive prompting: provide a clear task description and the target input, with no examples. The core objective is:

max_y P(y | P = t + x; θ)

Interpretation: maximize the probability of the desired output y given the prompt (task + input) and the model parameters θ.

- If the task is simple (e.g., “Translate ‘hello’ to French”), this works.
- If the task is complex (e.g., “Format this unstructured data into a table”), this often fails.

This is like asking a human to do a task without any examples — they might misunderstand the requirements. But there are problems:
- Ambiguity: The task description may be unclear, leading to irrelevant outputs.
- Lack of structure: The model has no guidance on the format or style of the output.
- Poor generalization: The model cannot adapt to niche tasks outside its pretraining data.

As I noted in my personal lecture notes: we are asking the model to infer the task from a single description, with no examples to anchor its understanding.

### Few-Shot Prompting: Adding Examples to Guide Generalization
To improve reliability, we add a small number of examples (k) to the prompt. The prompt becomes:

P = (x₁→y₁) + (x₂→y₂) + ... + (x_k→y_k) + t + x

The objective remains the same, but the examples provide a pattern for the model to follow:

max_y P(y | P = examples + t + x; θ)

Key considerations for effective few-shot examples:
- Relevance: Examples should be closely related to the target task and input.
- Diversity: Examples should cover different variations of the task to encourage generalization.
- Conciseness: Examples should be simple and focused, avoiding unnecessary details.

Intuition: Don’t just tell the model what to do; show it what to do, and let it generalize to new inputs.

### CoT Prompting: Structuring Reasoning for Complex Tasks
CoT prompting is an extension of few-shot prompting, tailored for reasoning tasks: it adds a step-by-step reasoning chain to each example. The prompt becomes:

P = (x₁→reasoning₁→y₁) + (x₂→reasoning₂→y₂) + ... + (x_k→reasoning_k→y_k) + t + x

The model learns to generate a reasoning chain before producing the final answer, which improves accuracy and interpretability.

CoT Objective Function (simplified):

max_y, r P(r | P = examples + t + x; θ) * P(y | r; θ)

Where r is the reasoning chain, and the model first generates r, then generates y based on r.

### CoT Prompting Algorithm
Algorithm: CoT Prompting Workflow
1. Define the target task (t) and desired output format (y)
2. Select k relevant examples (x_i, y_i) for the task
3. For each example, add a step-by-step reasoning chain (reasoning_i) that leads from x_i to y_i
4. Construct the prompt P: examples (x_i→reasoning_i→y_i) + task description (t) + target input (x)
5. Generate output from the LLM: the model will first generate a reasoning chain (r), then the final answer (y)
6. Evaluate the reasoning chain for coherence and accuracy; refine the examples if needed

#### Key insights:
- Step 3: The reasoning chain should be clear, logical, and specific to the example — it should show the model “how to think” about the task.
- Step 6: Refining examples (e.g., adding more detailed reasoning) improves the model’s reasoning quality over time.
- Step 4: The order of examples matters — start with simple examples, then move to more complex ones.

### The Role of Example Quality and Quantity
In CoT and few-shot prompting, the quality and quantity of examples each serve distinct roles in guiding the model:

1. Example Quality
- Defined by relevance, diversity, and clarity: examples must accurately represent the task and provide a clear pattern.
- Used to ensure the model learns the correct task pattern, not spurious correlations (e.g., irrelevant details in the examples).
- Intuition: “Show the model the right way to do the task, so it doesn’t learn bad habits.”

2. Example Quantity
- Typically 2-5 examples for most tasks (k=2-5); more examples may help for complex tasks, but diminishing returns apply.
- Why? Because LLMs have limited context windows — too many examples will crowd the prompt and reduce performance.
- Intuition: “Give the model enough examples to learn the pattern, but not so many that it gets overwhelmed.”

3. Why Both Matter
- High-quality examples ensure the model learns the correct task; the right quantity ensures the model generalizes without being overwhelmed.
- Poor-quality examples (e.g., irrelevant or incorrect) will lead the model to produce bad outputs, even with many examples.
- Too few examples may not be enough for the model to learn the pattern; too many will waste context space.

### Analogy to Fine-Tuning
Fine-tuning also uses examples to update the model, but with key differences:
- Fine-tuning: Updates model parameters based on examples, making changes permanent.
- Prompt Engineering: Uses examples in the prompt to guide the model’s output, making changes temporary and flexible.

CoT/few-shot prompting follows the same goal (guiding the model with examples), but the “in-context” nature makes it more flexible than fine-tuning.

### TL;DR:
- Example Quality: Ensures the model learns the correct task pattern → avoids spurious correlations.
- Example Quantity: Balances generalization and context constraints → 2-5 examples are optimal for most tasks.

### A Toy Example: Math Problem Solving
The lecture walked through a simple toy environment: prompts are math word problems, and the task is to generate the correct answer with a reasoning chain.

#### Prompt Examples (CoT)
Example 1:
- Input: “John has 5 apples. He gives 2 to Mary, then buys 3 more. How many apples does John have now?”
- Reasoning: “John starts with 5 apples. He gives 2 away, so 5 - 2 = 3 apples left. Then he buys 3 more, so 3 + 3 = 6. ”
- Output: 6

Example 2:
- Input: “A store has 10 books. They sell 4 books in the morning and 2 in the afternoon. How many books are left?”
- Reasoning: “The store starts with 10 books. They sell 4 in the morning, so 10 - 4 = 6 books left. Then they sell 2 in the afternoon, so 6 - 2 = 4. ”
- Output: 4

Target Input: “Lisa has 7 pencils. She lends 3 to Tom, then finds 2 more. How many pencils does Lisa have now?”

#### Prompt Design
```python
def cot_prompt(examples, task, target_input):
prompt = ""
for x, r, y in examples:
prompt += f"Input: {x}\nReasoning: {r}\nOutput: {y}\n\n"
prompt += f"Task: {task}\nInput: {target_input}\nReasoning:"
return prompt
```

#### Key Prompting Functions
Generate few-shot prompt without reasoning:
```python
def few_shot_prompt(examples, task, target_input):
prompt = ""
for x, y in examples:
prompt += f"Input: {x}\nOutput: {y}\n\n"
prompt += f"Task: {task}\nInput: {target_input}\nOutput:"
return prompt
```

Evaluate prompt effectiveness (simplified):
```python
def evaluate_prompt(llm_output, ground_truth):
# Check if the final output matches the ground truth
output = llm_output.split("Output:")[-1].strip()
return output == str(ground_truth)

def evaluate_cot_prompt(llm_output, ground_truth):
# Check both reasoning coherence and output correctness
reasoning = llm_output.split("Reasoning:")[-1].split("Output:")[0].strip()
output = llm_output.split("Output:")[-1].strip()
coherence = "so" in reasoning or "because" in reasoning # Simple coherence check
correct = output == str(ground_truth)
return coherence and correct
```

### Conclusions
- Prompt quality vs performance: A well-designed prompt (few-shot/CoT) can outperform a poorly designed prompt even with a weaker model.
- Reasoning chains: Help improve accuracy for complex tasks but require more context — balance reasoning detail with context constraints.
- Flexibility: Prompt engineering is faster and cheaper than fine-tuning, making it ideal for rapid prototyping and niche tasks.
- Scaling: The toy code works for simple tasks, but production prompt engineering involves testing multiple prompt variants, optimizing for context window usage, and adapting to different LLM architectures.

Fine-tuning can only adapt the model to specific tasks permanently. Prompt engineering unlocks flexible, on-the-fly adaptation to new tasks with minimal effort. While toy demos like “math problem solving” are simple, the mechanics mirror how prompt engineering is used to optimize LLM performance in real-world applications.

## References
1. Brown, T., et al. “Language Models are Few-Shot Learners.” NeurIPS (2020).
2. Wei, J., et al. “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.” NeurIPS (2022).
3. Liu, P., et al. “Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing.” TACL (2021).
4. Zhang, Y., et al. “Automatic Prompt Engineering for Large Language Models.” EMNLP (2023).
5. OpenAI. “GPT-4 Prompt Engineering Guide.” OpenAI Blog (2023).
6. Anthropic. “Claude Prompt Engineering Best Practices.” Anthropic Documentation (2024).
7. DeepSeek Team. “Prompt Engineering for Mathematical Reasoning in LLMs.” arXiv preprint arXiv:2502.01905 (2025).
8. Qwen Team. “Qwen3 Prompt Engineering Guide.” arXiv preprint arXiv:2506.00898 (2025).