Write the acceptance criteria first
Describe the output the task actually needs. For a meeting summary, accurate actions and explicit unknown owners may matter more than elegant prose. For data analysis, reproducible calculations and clear assumptions are essential. A generic “looks good” judgment gives little guidance when something fails.
Create representative test inputs
Include a normal example, a missing-information example, a conflicting-source example and an out-of-scope request. Use synthetic, public or explicitly approved material. Record what a knowledgeable reviewer expects and which aspects allow reasonable variation; not every writing task has one exact answer.
Separate material errors from style preferences
A fabricated deadline is different from a sentence that could be shorter. Mark unsupported claims, numerical errors, lost exceptions and unauthorized actions separately. Define which errors require the output to be rejected even if the rest of the response is useful.
Check against the source
Open cited documents and reproduce important calculations. A citation can be real while failing to support the associated statement. Ask whether the answer preserves relevant qualifications and identifies missing information. Evaluate the final result, not only whether it follows the requested format.
Record the conditions
Capture the tool or product, relevant configuration, prompt version, input, output, reviewer and test date. Compare alternatives under comparable conditions. A small local test provides evidence about those examples; it does not establish universal model quality or future performance.
Use findings to decide the next step
Group recurring errors, then change the task boundary, prompt, context or review process. Re-test important cases after a change. Track the time spent checking and correcting as part of the workflow. Stop or narrow the pilot when its outputs cannot be reviewed with confidence.
