Evaluation protocol
Show the input, the output and what failed.
Our proposed protocol for fair ChatGPT and Claude prompt comparisons: fixed inputs, task-specific criteria, saved outputs and transparent limitations.
Updated
Current status: protocol, not completed results
Our published templates are editorial starting points. We have not systematically evaluated the library on live models. This page describes the protocol we intend to use when actual outputs are available; it is not a report of a completed benchmark.
Source checking, a successful website build and a model-output test are different activities. We do not describe one as proof of another. Any future comparison should identify the exact tested task and keep conclusions within that evidence.
Fix the task and inputs before comparing
Use the same source material, brief and output requirements in each tool. Record the displayed model, relevant settings, date and tool availability. If one run can browse and the other cannot, disclose that difference or separate the tests.
Include an ordinary case and a difficult case: conflicting dates in a report, a mandatory qualification in copy or an inaccessible detail in an image. Repeat runs when making a stronger reliability claim; one output per tool is only a small case comparison.
Before testing this prompt, create a record containing task, fixed inputs, expected output, important failure cases, model/settings/date and permitted tools. Define what would count as a consequential error. Do not generate predicted results.
Use task-specific review criteria
A report should preserve facts and uncertainty. Copy should preserve the offer while matching voice. A poster concept should communicate a distinct, supportable idea. Judge those separately from whether an answer sounds polished. Automated review can assist, but human checking remains necessary.
| Task | Evidence to inspect | Do not confuse with success |
|---|---|---|
| Report | Correct figures, references, caveats and decision fit | A fluent summary |
| Brand writing | Preserved facts, observable voice rules and clear action | Extra adjectives |
| Poster concept | Different ideas, message clarity and brand constraints | Visual novelty alone |
| Code | Expected behaviour and relevant regression checks | A passing build alone |
Save failures and revision work
Keep the raw output, review notes and any correction prompt. Distinguish factual defects from preferences. An answer needing less substantive correction may be more useful even if another has a stronger opening. Describe actual editing time only when it was measured.
For publication, include the test materials that can be shared, relationship disclosures and the limits of the sample. A provider change can affect later results, so preserve the historical model/date record.
Compare outputs A and B against this fixed rubric. Cite the passage supporting each finding. Separate factual error, missing requirement and style preference. Summarise required corrections and unresolved judgement calls. Do not infer a universal model ranking from these examples.
Common questions
Are all prompts on this site tested?
No. They are editorially reviewed templates. Systematic live-model evaluation is pending; completed tests will need their own inputs, outputs and dated record.
Can one comparison tell me which model is best?
It can inform that task under those conditions. It cannot establish a universal ranking or a precise reliability rate.
Keep exploring
Sources and editorial note
Guidance checked 9 October 2026. These AI-assisted guides and original templates are reviewed against the linked sources. Results depend on your inputs and available tools. Model recommendations are editorial starting points, not comparative test results.