Compare models without hiding the test
Promptle Benchmark v1 is a reproducible protocol for testing the same prompt across ChatGPT, Claude, Codex, and other models. We publish the rubric and controls first; scored results will appear only when the raw output, model identifier, date, and settings can be preserved.
The test contract
- Use a fresh conversation or clean agent session for every run.
- Paste the prompt unchanged and supply the same source packet.
- Record provider, exact model label, date, enabled tools, and system-level settings visible to the tester.
- Save the complete output before scoring; do not silently repair it.
- Use two reviewers and resolve score differences with a written note.
Twenty-point scoring rubric
| Dimension | What reviewers inspect | Score |
|---|---|---|
| Instruction coverage | Every explicit deliverable is present | 0–4 |
| Constraint adherence | Format, length, exclusions, and boundaries are respected | 0–4 |
| Evidence discipline | Unsupported claims are avoided or clearly labeled | 0–4 |
| Actionability | The result can be used without inventing the next step | 0–4 |
| Calibration | Uncertainty, missing inputs, and assumptions are visible | 0–4 |
A zero means the requirement is absent or contradicted. Four means it is complete, specific, and needs no material repair. Report each dimension separately; a total score must never hide an evidence-discipline failure.
Model-by-model run sheet
Run the same task in each model family available to you. For chat models, disable memory and browsing unless those capabilities are part of the test. For coding agents, use the same repository snapshot, permissions, and validation commands. Report tool use separately so an agentic run is not presented as a plain-model comparison.
- ChatGPT: record the displayed model name, tools, memory state, and whether a custom GPT or project supplied extra instructions.
- Claude: record the displayed model name, project knowledge, enabled tools, and artifact or analysis modes used.
- Codex or another coding agent: record model, reasoning setting, sandbox, repository commit, commands run, tests, and final diff.
- Other models: record equivalent context-window, tool, retrieval, and sampling controls exposed by the provider.
Starter benchmark task
Task: Given a supplied product brief and five customer quotations, produce a one-page launch recommendation with three evidence-linked audience insights, two competing positioning options, a risk register, and five questions that must be answered before launch. Do not introduce statistics or customer claims absent from the packet.
Why this test: it measures synthesis, instruction coverage, restraint, and decision usefulness without rewarding trivia recall. Replace the source packet only between benchmark versions, never between models in the same run.
Publication policy
This page is the benchmark specification, not a claim that every listed model has already been tested. Promptle will not publish a leaderboard from undocumented runs or compare changing model labels as if they were stable products. Version 1 was reviewed on .
