Guides & case studies

Evaluation protocol

Show the input, the output and what failed.

Our proposed protocol for fair ChatGPT and Claude prompt comparisons: fixed inputs, task-specific criteria, saved outputs and transparent limitations.

Updated

Current status: protocol, not completed results

Our published templates are editorial starting points. We have not systematically evaluated the library on live models. This page describes the protocol we intend to use when actual outputs are available; it is not a report of a completed benchmark.

Source checking, a successful website build and a model-output test are different activities. We do not describe one as proof of another. Any future comparison should identify the exact tested task and keep conclusions within that evidence.

Fix the task and inputs before comparing

Use the same source material, brief and output requirements in each tool. Record the displayed model, relevant settings, date and tool availability. If one run can browse and the other cannot, disclose that difference or separate the tests.

Include an ordinary case and a difficult case: conflicting dates in a report, a mandatory qualification in copy or an inaccessible detail in an image. Repeat runs when making a stronger reliability claim; one output per tool is only a small case comparison.

Try this prompt

Before testing this prompt, create a record containing task, fixed inputs, expected output, important failure cases, model/settings/date and permitted tools. Define what would count as a consequential error. Do not generate predicted results.

Use task-specific review criteria

A report should preserve facts and uncertainty. Copy should preserve the offer while matching voice. A poster concept should communicate a distinct, supportable idea. Judge those separately from whether an answer sounds polished. Automated review can assist, but human checking remains necessary.

TaskEvidence to inspectDo not confuse with success
ReportCorrect figures, references, caveats and decision fitA fluent summary
Brand writingPreserved facts, observable voice rules and clear actionExtra adjectives
Poster conceptDifferent ideas, message clarity and brand constraintsVisual novelty alone
CodeExpected behaviour and relevant regression checksA passing build alone

Save failures and revision work

Keep the raw output, review notes and any correction prompt. Distinguish factual defects from preferences. An answer needing less substantive correction may be more useful even if another has a stronger opening. Describe actual editing time only when it was measured.

For publication, include the test materials that can be shared, relationship disclosures and the limits of the sample. A provider change can affect later results, so preserve the historical model/date record.

Try this prompt

Compare outputs A and B against this fixed rubric. Cite the passage supporting each finding. Separate factual error, missing requirement and style preference. Summarise required corrections and unresolved judgement calls. Do not infer a universal model ranking from these examples.

Common questions

Are all prompts on this site tested?

No. They are editorially reviewed templates. Systematic live-model evaluation is pending; completed tests will need their own inputs, outputs and dated record.

Can one comparison tell me which model is best?

It can inform that task under those conditions. It cannot establish a universal ranking or a precise reliability rate.

Keep exploring

Sources and editorial note

Guidance checked 9 October 2026. These AI-assisted guides and original templates are reviewed against the linked sources. Results depend on your inputs and available tools. Model recommendations are editorial starting points, not comparative test results.