Open evaluation framework

Compare models without hiding the test

Promptle Benchmark v1 is a reproducible protocol for testing the same prompt across ChatGPT, Claude, Codex, and other models. We publish the rubric and controls first; scored results will appear only when the raw output, model identifier, date, and settings can be preserved.

The test contract

  1. Use a fresh conversation or clean agent session for every run.
  2. Paste the prompt unchanged and supply the same source packet.
  3. Record provider, exact model label, date, enabled tools, and system-level settings visible to the tester.
  4. Save the complete output before scoring; do not silently repair it.
  5. Use two reviewers and resolve score differences with a written note.

Twenty-point scoring rubric

DimensionWhat reviewers inspectScore
Instruction coverageEvery explicit deliverable is present0–4
Constraint adherenceFormat, length, exclusions, and boundaries are respected0–4
Evidence disciplineUnsupported claims are avoided or clearly labeled0–4
ActionabilityThe result can be used without inventing the next step0–4
CalibrationUncertainty, missing inputs, and assumptions are visible0–4

A zero means the requirement is absent or contradicted. Four means it is complete, specific, and needs no material repair. Report each dimension separately; a total score must never hide an evidence-discipline failure.

Model-by-model run sheet

Run the same task in each model family available to you. For chat models, disable memory and browsing unless those capabilities are part of the test. For coding agents, use the same repository snapshot, permissions, and validation commands. Report tool use separately so an agentic run is not presented as a plain-model comparison.

  • ChatGPT: record the displayed model name, tools, memory state, and whether a custom GPT or project supplied extra instructions.
  • Claude: record the displayed model name, project knowledge, enabled tools, and artifact or analysis modes used.
  • Codex or another coding agent: record model, reasoning setting, sandbox, repository commit, commands run, tests, and final diff.
  • Other models: record equivalent context-window, tool, retrieval, and sampling controls exposed by the provider.

Starter benchmark task

Task: Given a supplied product brief and five customer quotations, produce a one-page launch recommendation with three evidence-linked audience insights, two competing positioning options, a risk register, and five questions that must be answered before launch. Do not introduce statistics or customer claims absent from the packet.

Why this test: it measures synthesis, instruction coverage, restraint, and decision usefulness without rewarding trivia recall. Replace the source packet only between benchmark versions, never between models in the same run.

Publication policy

This page is the benchmark specification, not a claim that every listed model has already been tested. Promptle will not publish a leaderboard from undocumented runs or compare changing model labels as if they were stable products. Version 1 was reviewed on .

Read the editorial methodology →