Introduction
Most people write a prompt, read the output, decide it did not work, and move on. They try something completely different next time - or stop using the tool altogether.
That approach treats the first output as a verdict. It is not. It is data.
The draft - test - refine loop gives you a structured way to improve any prompt. Instead of rewriting everything when a result misses the mark, you identify which element caused the gap and adjust that element. One change at a time. One comparison at a time.
This article walks through the loop: what each stage entails, how to identify the element responsible for an output gap, and how to apply targeted adjustments that yield measurably better results.
Start with the three stages.
What the Loop Is - and What It Is Not
The draft - test - refine loop is a three-stage process:
| Stage | Your Action | What to Look For |
|---|---|---|
| Draft | Build a prompt using the seven-element framework. Include the elements the task requires. Start with Task - that is always your anchor. | Nothing to evaluate yet. The draft is a hypothesis about which elements will produce the desired output. |
| Test | Submit the prompt. Read the full output. Evaluate it against a clear success criterion - not just "does it feel right," but a specific question: Is it the right format? The right length? The right tone? The right scope? | Identify the specific gap. Not "this is wrong" - but which element is missing, unclear, or absent. Name the gap before revising. |
| Refine | Adjust one element. Only one. Add a constraint, clarify the format, include an example, adjust the role, or rework the task instruction. Submit the revised prompt. Compare the output to the previous version. | A targeted change produces a readable comparison. Changing multiple elements at once makes it difficult to identify which adjustment produced the difference. |
The loop is not a troubleshooting guide for broken tools. It is a deliberate practice for improving a working process toward a more precise result.
It also does not guarantee a specific output. AI model outputs can vary depending on the model, the task, and the prompt's construction. What the loop typically does is reduce the gap between your intention and the output, version by version.
Stage 1 - Draft: Build a Starting Prompt
The draft stage is where you apply the seven-element framework. Your goal at this stage is not a perfect prompt - it is a structured one.
Start with the task. Every prompt needs one clear task instruction using a specific action verb: summarize, compare, list, draft, or explain. Vague verbs - help, write, do - leave too much open.
Then apply the remaining elements based on what the output needs to accomplish:
Task: Always required. State the job clearly using a specific verb.
Role: Apply when a professional identity meaningfully shapes the output.
Context: Add when the audience, purpose, or background changes the result.
Format: Specify when the output shape matters - a table, a list, a paragraph.
Tone: Name the voice when the default register is not right for the task.
Constraints: Set limits on length, scope, or content when the task is bounded.
Examples: Show a sample when the style is easier to demonstrate than to describe.
Do not try to include every element in the first draft. A prompt that includes every element, regardless of need, is often longer and less clear. Start with what the task clearly requires and test from there.
Stage 2 - Test: Evaluate the Output by Element
Testing is not reading the output and deciding whether you like it. It is a structured evaluation against a specific criterion.
Before you test, define your success criterion. Answer this question: "What would a successful output look like for this task?" Name the format, the length, the tone, or the scope - whatever matters most.
Then evaluate the output against that criterion and identify the gap by element. The table below maps common output gaps to the element most likely responsible for them.
| Output Gap | Likely Missing Element | What to Add or Adjust |
|---|---|---|
| Output is in the wrong format (prose when you needed a table, bullets when you needed a paragraph) | Format element missing or unclear | Add a format instruction: "Present this as a two-column table" or "Write as three short paragraphs." |
| Output is too long or includes content you did not need | Constraints element absent | Add a constraint: "Keep each point to one sentence" or "Do not include examples - definitions only." |
| Output uses the wrong register - too formal, too casual, or off-brand | Tone element missing or not specific enough | Name the tone precisely: "Conversational and direct" or "Confident but not promotional." |
| Output lacks audience relevance - too technical, too basic, or not targeted | Context element incomplete | Add audience and purpose: "This is for a non-technical team with no engineering background." |
| Output uses the wrong professional framing or vocabulary level | Role element absent | Add a role cue when professional identity changes the output: "You are a communications coordinator writing for an internal audience." |
| Output style is inconsistent with what you need - structure, sentence length, voice | Examples element absent | Add a short example: "Here is an example of the style I want: [one sentence in the target voice]." |
| Output misses the specific job - too broad, too narrow, or off-target | Task instruction is too vague | Rewrite the task with a more specific verb and scope: "List" instead of "Write about"; "Summarize in three points" instead of "Cover the main ideas." |
One gap at a time. If the output has two problems, pick the one that matters most and fix that element first. Running multiple changes in one revision makes the comparison harder to read.
Stage 3 - Refine: Adjust One Element
Refinement is a targeted edit, not a rewrite. You are changing one element based on what you identified in the test stage.
The rule: identify the element responsible for the gap, then adjust only that element. Submit the revised prompt. Compare the new output to the previous version.
Three adjustments tend to produce the biggest improvement in early prompts:
Add a Constraint
When output is too long, too broad, or includes content you did not want, a constraint is often the fastest fix.
| Version 1 - No Constraint: | Summarize the following project update for an internal team. Use plain language. |
| Version 2 - Constraint Added: | Summarize the following project update for an internal team. Use plain language. Keep each point to one sentence. Do not include technical terms - if a concept must be mentioned, describe it in plain language instead. |
What tends to happen: The constrained version typically produces a tighter, more audience-appropriate output. The underlying task and context did not change - one targeted element did.
Add a Format Instruction
When the output arrives in the wrong shape, a format instruction is usually the fix - not a rewrite of the entire prompt.
| Version 1 - No Format: | Explain the difference between a base AI model and a tool-connected AI system for a beginner audience. |
| Version 2 - Format Specified: | Explain the difference between a base AI model and a tool-connected AI system for a beginner audience. Format the response as a two-column comparison table. Use one sentence per row. Avoid technical jargon. |
What tends to happen: A format instruction often produces an output tailored to the intended use - a table for quick comparison, a list for scannable steps - without changing the task's content.
Add an Example
When the style is right but the execution is off - wrong sentence length, wrong pattern, inconsistent voice - an example tends to narrow the gap more efficiently than additional description.
| Version 1 - No Example: | Write three bullet points summarizing a product launch update. Use plain language. One sentence per bullet. |
| Version 2 - Example Added: | Write three bullet points summarizing a product launch update. Use plain language. One sentence per bullet. Here is an example of the style I want: "Our new tool is now available and ready to use." |
What tends to happen: A single-sentence example in the right voice often produces outputs that match the pattern more closely than a longer tonal description would. Showing tends to be more efficient than describing when style is the issue.
A Worked Iteration: Three Versions of One Prompt
Here is a complete example of the loop applied across three versions of the same prompt. The task: summarize a project update for a non-technical internal team.
Version 1 - Draft
| Prompt: | Summarize the following project update in three bullet points. Write for a non-technical internal team. Use plain language. |
Output gap identified: Two of the three bullets ran to two or three sentences. Both included technical terms - "API integration" and "testing pipeline" - that the audience would not recognize. Element responsible: Constraints absent.
Version 2 - First Refinement
| Prompt: | Summarize the following project update in three bullet points. Write for a non-technical internal team. Use plain language. Keep each bullet to one sentence. If a technical term must be mentioned, describe it briefly in plain language instead. |
Output gap identified: Bullets were now one sentence and avoided jargon. The tone was still slightly formal - "we are currently in the process of" rather than direct, plain-language phrasing the team would use. Element responsible: Tone not specified; role cue absent.
Version 3 - Second Refinement
| Prompt: | You are a communications coordinator writing for an internal non-technical team. Summarize the following project update in three bullet points. Keep each bullet to one sentence. If a technical term must be mentioned, describe it briefly in plain language. Here is an example of the style I want: "The team completed testing last week, and the feature is ready to launch." |
Output result: Output matched the target. One-sentence bullets. Plain language throughout. The role cue and example together shifted the register to match the audience better than the constraint alone did.
What This Iteration Showed
Three versions. Three targeted adjustments - Constraints, Role, Example. Each version identified one gap and fixed one element.
The first output was not a failure. It showed exactly which element to add next.
Changing one element at a time made it straightforward to identify which adjustment produced the improvement.
Common Mistakes When Iterating
Three patterns tend to slow down iteration for beginners:
Mistake 1: Changing Multiple Elements at Once
Adding a role, a constraint, and an example in the same revision makes it difficult to identify which adjustment produced the change. If the output improves, you do not know which element was responsible. If it does not improve, you have no clear signal about what to try next.
The fix: Change one element per revision. If the gap has multiple causes, prioritize the one that matters most and address the others in sequence.
Mistake 2: Treating the First Output as a Verdict
One output is not enough to evaluate a prompt. It shows you what the model produces with your current element combination - not what is possible with one or two adjustments.
The fix: The first output is a starting point. Use it to identify the gap, then apply one targeted adjustment and compare.
Mistake 3: Rewriting the Entire Prompt When Only One Element Is Off
When output is mostly right, but one thing is off - the format, the length, the tone - rewriting the full prompt often changes elements that were already working. This makes the next comparison harder to read.
The fix: Identify the specific element causing the gap. Edit that element only and keep the rest of the prompt unchanged.
Avoiding these three mistakes will not prevent all output gaps. But in many cases, addressing even one of them makes the next iteration more readable and the next adjustment easier to identify.
Key Takeaways
- The draft - test - refine loop treats the first prompt output as a starting point, not a verdict. Each output shows you what to adjust next.
- The test stage requires a specific success criterion before you evaluate. Name the format, length, tone, or scope you need - then compare the output against that standard.
- Identify the element responsible for the output gap before revising. Changing multiple elements at once makes the comparison harder to read and the cause of any improvement harder to identify.
- Three adjustments tend to produce the biggest early gains: adding a constraint, specifying a format, or including an example. Each targets a different common gap.
- The loop does not guarantee a specific output. What it typically does is reduce the gap between your intention and the result, version by version, one element at a time.
What to Try Next
Pick a prompt you have used recently - a summary, a draft, a list - and run it through one cycle of the loop.
Start with a success criterion. Define what a good output looks like for that task before you submit anything. Then evaluate the output against that standard, identify one element to adjust, and compare the results.
If the output from your second version is noticeably closer to the target, note which element made the difference. That is the beginning of a prompt habit that compounds over time.
Chapter 4 of the Learning Prompt Engineering eBook covers the full iteration process, including a worked example with all seven elements and guidance on building a personal prompt library.




0 Comments