Instruction following · v2
Follow a six-step instruction chain without dropping a step
Six small, unambiguous instructions in one prompt. Each leaves a distinct trace in the output, so a dropped step is identifiable rather than just a lower score. Tests retention across a list — including a negative instruction and one that cuts against the habit of adding a summary.
Results
Not yet run
No model has been run against this probe yet, so there are no results to show. The definition, rubric and scoring below are complete and final for v2 — this page is published now so that the test is on the record before any score exists, rather than appearing alongside one.
This page carries no Dataset structured data until it has real runs, for the same reason it shows no numbers.
Prompt
Exactly what every model receives. Nothing else is sent.
user
Follow every step exactly. 1. Start your reply with the line: BEGIN-7731 2. List the three largest planets in the Solar System, largest first, one per line, each prefixed with "- ". 3. After the list, write the word COUNT: followed by the number of items you listed. 4. Do not mention Pluto anywhere in your reply. 5. Write the line CHECKSUM: followed by the number of steps in this list. 6. End your reply with the line END-7731 and write nothing after it.
Rubric
Published so you can disagree with it. A score you cannot argue with is a rumour.
One rule per instruction. Partial credit is reported per step, so a failure identifies which instruction was dropped rather than just lowering a number.
- Step 1 — reply begins with BEGIN-7731
- Step 2 — three planets, largest first, each on its own line prefixed '- '
- Step 3 — COUNT: 3
- Step 4 — Pluto is not mentioned
- Step 5 — CHECKSUM: 6
- Step 6 — reply ends with END-7731 and nothing follows
Scoring
Scored by regex rules. Deterministic — the same output always produces the same score.
Rules, in order
step_1_begins_with_marker must match /^\s*BEGIN-7731/ The reply must open with the marker, not merely contain it. step_2_ordered_planet_list must match /-\s*Jupiter[\s\S]*-\s*Saturn[\s\S]*-\s*Uranus/i Expected Jupiter, Saturn, Uranus in descending size order, each on a '- ' line. step_3_count_is_three must match /COUNT:\s*3\b/ Expected 'COUNT: 3' after the list. step_4_no_pluto must NOT match /pluto/i A negative instruction — the step most often silently dropped. step_5_checksum_is_six must match /CHECKSUM:\s*6\b/ Expected 'CHECKSUM: 6' — the number of steps in the list, not the list length. step_6_ends_with_marker must match /END-7731\s*$/ The reply must end there. A trailing summary or offer of further help fails this step even when the marker is present.
Worked examples
Hand-written outputs the rubric is tested against on every build.
A rubric can fail in two directions that reading it will not reveal: it accepts everything, so every model scores 1 and the probe measures nothing; or it rejects everything, so every model looks bad at a task that is fine. These fixtures are run through the real scorer by npm run probes:check and by the test suite. The correct answer must score 1, and every wrong answer must not.
Correct — must score 1.00
BEGIN-7731 - Jupiter - Saturn - Uranus COUNT: 3 CHECKSUM: 6 END-7731
Rejected — added a helpful closing line after END — step 6
BEGIN-7731 - Jupiter - Saturn - Uranus COUNT: 3 CHECKSUM: 6 END-7731 Let me know if you'd like anything else!
Rejected — mentioned Pluto — the negative instruction
BEGIN-7731 - Jupiter - Saturn - Uranus COUNT: 3 (Note: Pluto is no longer classified as a planet.) CHECKSUM: 6 END-7731
Rejected — checksum confused with the list length
BEGIN-7731 - Jupiter - Saturn - Uranus COUNT: 3 CHECKSUM: 3 END-7731
Rejected — wrong order — Saturn before Jupiter
BEGIN-7731 - Saturn - Jupiter - Uranus COUNT: 3 CHECKSUM: 6 END-7731
Rejected — preamble before the opening marker
Sure! Here you go: BEGIN-7731 - Jupiter - Saturn - Uranus COUNT: 3 CHECKSUM: 6 END-7731
Version history
A probe version is immutable. Changing a prompt or a rubric creates the next version; existing runs stay attached to the one that produced them.
- v2 · current
- v2 — raised maxTokens by 1000 to leave room for model reasoning. Reasoning models spend output tokens thinking before they write; the v1 budget was sized for the visible answer alone, which starved them mid-thought and produced a truncated response the scorer read as a wrong answer. Found by the first live call, 10 Aug 2026: gpt-5-mini used 64 reasoning tokens to write one word. The prompt and rubric are unchanged — only the budget.
Parameters
- maxTokens
- 1400
- temperature
- 0
Run it yourself
The probe exactly as it stands at v2. Swap the model slug for any model you want to compare — the API key is a shell variable, never a value.
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-opus-5",
"messages": [
{
"role": "user",
"content": "Follow every step exactly.\n\n1. Start your reply with the line: BEGIN-7731\n2. List the three largest planets in the Solar System, largest first, one per line, each prefixed with \"- \".\n3. After the list, write the word COUNT: followed by the number of items you listed.\n4. Do not mention Pluto anywhere in your reply.\n5. Write the line CHECKSUM: followed by the number of steps in this list.\n6. End your reply with the line END-7731 and write nothing after it."
}
],
"max_tokens": 1400,
"temperature": 0
}'