Prompt Engineering Guide: Techniques That Actually Work
By Modern AI Engineering · · 8 min read
Prompt engineering is the practice of writing the input to a language model so that it gives the output you need. The techniques that work are simple. Say exactly what you want. Give the model the context it lacks. Show examples. Fix the output format. Split hard tasks into steps. Then test the prompt on real inputs instead of trusting one good result.
This guide explains each technique, when to use it, and what to do when a prompt still fails.
What is prompt engineering?
A model only knows two things when it answers: what it learned in training and what is in the prompt. You cannot change the first in a normal request. Prompt engineering is the work of getting the second one right.
A helpful way to think about it is that you are briefing a capable new colleague who knows nothing about your project. If you say only that you want a summary of a report, you will get a generic summary. If you say who it is for, how long it should be and what to leave out, you will get something you can use.
The prompt is also a cost. Every token in it is paid for on every request, and a longer prompt takes longer to process. Good prompts are complete, not long.
The parts of a good prompt
Most strong prompts contain the same parts. Not every prompt needs all of them, but when a result disappoints, one of these is usually missing.
| Part | What it tells the model |
|---|---|
| Role and goal | What it is helping with, and for whom. |
| Context | The facts, documents or data it needs for this request. |
| Task | Exactly what to do, in plain words. |
| Constraints | Length, tone, what to include and what to leave out. |
| Examples | One or more samples of a good answer. |
| Output format | The structure of the reply, such as a list, a table or JSON. |
Zero-shot and few-shot prompting
Zero-shot prompting means you ask for the task with no examples. It works for common tasks the model has seen many times, such as translating or summarising. Start here, because it is the cheapest option.
Few-shot prompting means you add a few examples of input and output before the real input. The model picks up the pattern and follows it. This is one of the most reliable techniques there is. Use it when you need a particular style, a custom set of categories or a strict format that is hard to describe in words.
Choose the examples with care. The model copies them closely. If every example is short, the answers will be short. If every example carries the same label, the model will lean towards that label. Use a small, varied set that includes a hard case.
Chain-of-thought prompting
Chain-of-thought prompting asks the model to work through a problem in steps before it gives the final answer. A model writes one token at a time, and each token is based on what came before. If it writes out the steps first, the final answer can build on them. If it must answer at once, it has nothing to build on.
This helps on tasks with several steps: word problems, comparisons, decisions that depend on more than one rule. It costs extra tokens and extra time, so do not use it for simple lookups.
Reasoning models do a version of this by themselves before they reply. With those models a plain instruction about the goal often works better than telling them how to think.
Prompt chaining: split a hard task into steps
One long prompt that asks for five things at once often fails at one of them. Prompt chaining breaks the job into separate calls. The output of one call becomes the input of the next.
Suppose you want a reply to a customer email. The first call pulls out the question and the order number. The second drafts a reply from the right help article. The third checks the draft against your rules. Each prompt is short and has one job.
The gain is control. You can test each step alone, see which one failed, and fix only that one. The price is more calls, so use chaining when a single prompt is not reliable enough.
Prompt engineering best practices
These habits improve almost any prompt.
- Put long documents first and the question after them, and mark where each part begins and ends.
- Say what to do, not only what to avoid. A positive instruction is easier to follow.
- Give the reason for a rule. A model that knows why can apply the rule to cases you did not list.
- Ask for a format you can check by code, such as JSON with named fields.
- Tell the model what to do when it lacks the information. Allow it to say that it does not know.
- Keep the parts of the prompt that never change at the start. Many providers can cache that part, which lowers cost and delay.
- Change one thing at a time, so you know what caused the difference.
How do you test a prompt?
A prompt that works once has not been tested. Models give different outputs on different runs, and real inputs vary more than the one you tried.
- Collect twenty to fifty real inputs, including awkward ones.
- Write down what a good output looks like for each.
- Run the prompt on all of them and read the results.
- Group the failures by cause. Fix the most common cause first.
- Run the whole set again after each change, so a fix for one case does not break another.
When is prompt engineering not enough?
Prompting has limits, and it helps to know them early. If the model lacks facts, such as your company documents, no wording will supply them. You need to retrieve the right text and add it to the context. That is RAG. If the model must take actions or look things up, it needs tools. If you need the same behaviour across a very large number of requests and the prompt cannot hold it, fine-tuning may be the answer.
This is why the wider term context engineering is now common. It covers everything the model sees: instructions, retrieved documents, tool results and conversation history. In the AI Engineering Bootcamp, Module 9, The Art of Prompting, covers chain of thought, prompt chaining, prompt caching, context engineering and context compaction.
Frequently asked questions
What is prompt engineering in simple terms?
It is writing the input to an AI model clearly enough that it produces the output you need. It includes instructions, context, examples and the format of the answer.
What is the difference between zero-shot and few-shot prompting?
Zero-shot gives the task with no examples. Few-shot adds a few examples of input and output first, so the model can follow the pattern.
Is prompt engineering still worth learning?
Yes. Models are better at understanding loose requests than they used to be, but clear instructions, good context and testing still decide the quality of the result in a product.
Does chain-of-thought prompting always help?
No. It helps on tasks with several steps. On simple tasks it adds cost and delay for little gain, and reasoning models already work through steps without being told.