← All chaptersChapter 2 of 8

Designing and Testing Robust Prompts

You create prompts that not only give one nice answer but remain reliable across different inputs.

After this chapterYou can build a professional prompt architecture, use examples purposefully, and compare versions with a small test set.
Your progress0 of 40 lessons
2.1

The anatomy of a professional prompt

Combine result, context, output, limits, and control; only use parts that change the risk or usability.

A professional prompt is a transfer of work. Start with the end result and then add the relevant context, desired output format, and necessary boundaries. Indicate which source information is leading and what should be done in case of missing data.

A role can guide perspective or jargon, but it does not prove expertise. Therefore, it is better to use concrete assessment criteria rather than an impressive job title. Conclude important tasks with a control instruction: report missing information, unsupported claims, and deviations from the assignment separately.

  • Result
  • Relevant context
  • Usable format
  • Real boundaries
  • Post-check
How you can use this

An analysis prompt defines decision, sources, criteria, and uncertainties; it does not only ask for 'deep thinking'.

Try this prompt
Result: [result]. Context and sources: [context]. Output: [form]. Limits: [limits]. Check before completion: [criteria]. Mark missing information and do not invent source facts.
Knowledge check

You are comparing two quotations. One omits the delivery lead time. Which prompt addition prevents an assumption from appearing as a quotation detail?

Your practical exercise

Write a complete prompt for a fictional task. Remove demonstrably unnecessary text; keep essential boundaries and exception rules, even if your first tests do not yet show why they matter.

Source for this lesson

OpenAI – Prompting
Task, context, output and boundaries; the choice of working method depends on the task and availability.
Checked: 2026-09-07

2.2

Using examples without blindly copying

Examples make the pattern and quality level concrete, but must not turn incorrect or incidental features into rules.

With example-driven instruction, you show input and desired output. This is especially helpful for classification, fixed formats, tone of voice, and exceptions. Choose examples that together cover the normal cases and important boundaries.

Explain what the model should adopt from an example: structure, level of detail, criterion, or tone. Do not automatically reuse names, facts, or incidental wording. Include at least one unusual or incomplete case and define the correct behavior, for example 'insufficient information' instead of guessing.

  • Representative examples
  • Name desired feature
  • Include edge case
  • No unwanted fact copying
How you can use this

An email classification shows examples of a general question, a complaint, an ambiguous question, and an out-of-scope message.

Try this prompt
For this task, follow the examples; use only these [features] from them. Do not copy names or facts. Classify the new input into [categories]. If no category fits, answer BUITEN_SCOPE and briefly explain why. Examples: [examples].
Knowledge check

Examples teach ChatGPT to put meeting actions into a table. All example actions have Noor as their owner; the new fictional notes mention Sam. Which instruction preserves what is reusable?

Your practical exercise

Create three examples and one edge case for your own classification or formatting task.

2.3

Requesting verifiable intermediate steps

Ask for relevant pieces of evidence and reasons for decisions, not for a hidden internal thought process.

For professional work, you want to be able to trace what an outcome is based on. Therefore, ask for source extractions, assumptions, calculations, decision criteria, and a concise justification. You can check such artifacts without pretending that a generated explanation fully reveals the actual internal reasoning process.

Choose intermediate steps that locate errors. In a comparison, these are, for example, criteria, source values, missing data, and scores. In a text check, these are claims, evidence, and suggested corrections. A long line of thought is not evidence; verifiable data and repeatable tests are.

  • Source extraction
  • Assumptions separately
  • Calculation or criterion
  • Concise justification
  • Remaining uncertainty
How you can use this

When choosing a supplier, do not ask for 'all thoughts', but for source values, weights, calculation, and sensitivity analysis.

Try this prompt
Analyze [question]. Show: 1) used source facts, 2) necessary assumptions, 3) applied criteria or calculations, 4) concise reason for the conclusion, and 5) remaining uncertainties. Do not invent missing values.
Knowledge check

ChatGPT recommends option B in a weighted comparison. What additional material helps you check whether B follows from the agreed data and weights?

Your practical exercise

Rewrite one prompt so that it produces five verifiable intermediate outputs.

2.4

Building a small but strong test set

Test normal cases, boundaries, and failures before reusing a prompt.

A prompt that handles one example correctly is not yet a reliable workflow. Create a test set with representative input, a difficult case, missing information, conflicting sources, and possibly an inappropriate instruction in the source material. Define in advance for each test what minimally correct behavior is.

Use only fictional data for these exercises. Assess both content and safety: does the output follow the schema, avoid inventing facts, keep source content subordinate to the task, and stop the workflow when required? Keep failed examples as regression tests.

  • Normal case
  • Incomplete case
  • Contradiction
  • Out of scope
  • Unreliable source instruction

Terms in plain language

Test set
A collection of input examples with correct and incorrect behaviour defined in advance.
Prompt injection
An instruction in retrieved content that tries to divert the AI from the actual task. Treat such content as data to be assessed.
Regression test
Running an earlier test again to see whether a change breaks something that used to work.
How you can use this

A summarisation prompt is tested with an ordinary report, missing decisions, conflicting data and a document containing instructions for the AI.

Try this prompt
Design six tests for [workflow]. For each test, provide input type, expected behavior, forbidden behavior, and evidence criterion. Use fictitious data. Include one prompt-injection-like source fragment.
Knowledge check

A summarising prompt has been tested four times on complete reports without conflicts. You now want to test how it handles conflicting agreements. Which addition has a suitable expectation that can be checked in advance?

Your practical exercise

Create six fictional tests for the study club in the chapter example, including an unclear message. Record the expected label and reason in advance. If possible, ask a colleague to challenge one expectation. Working alone: write the strongest alternative for one ambiguous case yourself and check both labels against the agreed definitions. Done: six expectations have source support, a genuinely ambiguous case is not forced into a category, and you state who assessed the work.

2.5

Fairly comparing prompt versions

Change one important element per round and compare based on pre-established criteria.

Do not compare prompt versions based on feeling or on a single random answer. Use the same test set and the same evaluation criteria. Preferably change only one factor, such as source instruction, output schema, or error handling. That way you will know which adjustment likely made the difference.

Record version number, change, reason, tests, and known limitation. Measure not only quality but also correction work, consistency, and required time. A version with a higher average score can still be unsuitable if it handles one critical edge case incorrectly.

A test set used to improve the prompt has become practice material for its designer. Also reserve some new cases for a later check. Repeat important tests where possible: a model can give different answers to the same input. A small set reveals specific defects, not a proven general reliability percentage.

  • Same test set
  • One main change
  • Criteria before test
  • Critical errors separately
  • Version log

Terms in plain language

Disqualifying error
A critical error defined in advance that makes a version unsuitable, even when other scores are good.
How you can use this

Compare v1 without source labels with v1.1 with source labels; do not simultaneously change tone, length, and task sequence.

Try this prompt
Compare prompt v1 and v2 on [testset]. Score per test: adherence to instructions, use of sources, completeness, safety, and recovery work. List critical errors separately; do not use an average to hide a knockout error.
Knowledge check

V2 improves five ordinary tests but invents a deadline in a sixth test. That error was defined in advance as critical; v1 correctly left the deadline open. Which change decision fits?

Your practical exercise

Test two prompt versions on the same five cases and write a short change decision.

Source for this lesson

OpenAI – Evaluation best practices
Task-specific test sets and human assessment. The API software is not needed for this course. The source announces the phase-out of the Evals platform: read-only from 31 October 2026 and closure on 30 November 2026. This does not end the general testing method; check the product timeline before choosing software.
Checked: 2026-09-08

Worked example

A small test set makes a prompt change visible

Fictional practice material; incorrect answers have been created deliberately for this exercise.

All messages, labels and test results are fictional. You are sorting messages for a study club. Labels: INSCHRIJVING for an explicit request to join, ANNULERING for an explicit cancellation, and ONDUIDELIJK when neither clearly applies.

Input

Prompt v1: choose INSCHRIJVING or ANNULERING for each message. Tests defined in advance: T1 “I would like to sign up.” → INSCHRIJVING; T2 “I am cancelling my participation.” → ANNULERING; T3 “When does it start?” → ONDUIDELIJK; T4 “Perhaps I will sign up, perhaps not.” → ONDUIDELIJK.

First practice answer

Fictional v1 output: T1 INSCHRIJVING; T2 ANNULERING; T3 INSCHRIJVING; T4 INSCHRIJVING. T1 and T2 fit; T3 and T4 force a decision. The test set and expected labels are not changed after this observation.

Check

One main change: add an explicit route for uncertainty. Prompt v1.1: use INSCHRIJVING or ANNULERING only for an unambiguous request; otherwise use ONDUIDELIJK. Assess the message content; do not infer intentions that are not stated. Retest all four messages, including the ones that passed before.

Improved version

Fictional v1.1 output: T1 INSCHRIJVING; T2 ANNULERING; T3 ONDUIDELIJK; T4 ONDUIDELIJK. For these four cases, the number of correct labels rises from 2 to 4. Decision: keep v1.1 for further testing. Four examples do not yet prove reliability with other wording.

Try it yourself

Before running the tests, add two cases: one polite cancellation message and one message that mentions both signing up and cancelling. Record the expected label.

View the model answer

“Thank you for the invitation, but I am cancelling.” → ANNULERING. “Sign me up, or no, I am not sure yet whether I am cancelling.” → ONDUIDELIJK. Keep actual output separate from these expected answers, then rerun the full set of six.

Chapter assignment

Bring everything together

Build a v1 prompt and a set of six test cases. Then create v1.1 with one justified change: fix an observed weakness or clarify a boundary that has not been tested sufficiently. Run the same cases again and document the result, regression check and remaining limitation. Refining a boundary without an observed error does not demonstrate a quality gain.

Maximum 10,000 characters per note.

Progress and notes are stored only in this browser on this device. Do not enter sensitive data. Download your notes regularly. This course sets no automatic expiry date. You can delete the data through your browser’s site-data settings; export anything you wish to keep first. Browser settings or cleanup may erase it earlier. These local notes are not sent to Finaudax.