LLM Ground

Instruction following · v2

Follow a six-step instruction chain without dropping a step

Six small, unambiguous instructions in one prompt. Each leaves a distinct trace in the output, so a dropped step is identifiable rather than just a lower score. Tests retention across a list — including a negative instruction and one that cuts against the habit of adding a summary.

Results

Not yet run

No model has been run against this probe yet, so there are no results to show. The definition, rubric and scoring below are complete and final for v2 — this page is published now so that the test is on the record before any score exists, rather than appearing alongside one.

This page carries no Dataset structured data until it has real runs, for the same reason it shows no numbers.

Prompt

Exactly what every model receives. Nothing else is sent.

user

Follow every step exactly.

1. Start your reply with the line: BEGIN-7731
2. List the three largest planets in the Solar System, largest first, one per line, each prefixed with "- ".
3. After the list, write the word COUNT: followed by the number of items you listed.
4. Do not mention Pluto anywhere in your reply.
5. Write the line CHECKSUM: followed by the number of steps in this list.
6. End your reply with the line END-7731 and write nothing after it.

Rubric

Published so you can disagree with it. A score you cannot argue with is a rumour.

One rule per instruction. Partial credit is reported per step, so a failure identifies which instruction was dropped rather than just lowering a number.

  • Step 1 — reply begins with BEGIN-7731
  • Step 2 — three planets, largest first, each on its own line prefixed '- '
  • Step 3 — COUNT: 3
  • Step 4 — Pluto is not mentioned
  • Step 5 — CHECKSUM: 6
  • Step 6 — reply ends with END-7731 and nothing follows

Scoring

Scored by regex rules. Deterministic — the same output always produces the same score.

Rules, in order

step_1_begins_with_marker
  must match  /^\s*BEGIN-7731/
  The reply must open with the marker, not merely contain it.

step_2_ordered_planet_list
  must match  /-\s*Jupiter[\s\S]*-\s*Saturn[\s\S]*-\s*Uranus/i
  Expected Jupiter, Saturn, Uranus in descending size order, each on a '- ' line.

step_3_count_is_three
  must match  /COUNT:\s*3\b/
  Expected 'COUNT: 3' after the list.

step_4_no_pluto
  must NOT match  /pluto/i
  A negative instruction — the step most often silently dropped.

step_5_checksum_is_six
  must match  /CHECKSUM:\s*6\b/
  Expected 'CHECKSUM: 6' — the number of steps in the list, not the list length.

step_6_ends_with_marker
  must match  /END-7731\s*$/
  The reply must end there. A trailing summary or offer of further help fails this step even when the marker is present.

Worked examples

Hand-written outputs the rubric is tested against on every build.

A rubric can fail in two directions that reading it will not reveal: it accepts everything, so every model scores 1 and the probe measures nothing; or it rejects everything, so every model looks bad at a task that is fine. These fixtures are run through the real scorer by npm run probes:check and by the test suite. The correct answer must score 1, and every wrong answer must not.

Correct — must score 1.00

BEGIN-7731
- Jupiter
- Saturn
- Uranus
COUNT: 3
CHECKSUM: 6
END-7731

Rejected — added a helpful closing line after END — step 6

BEGIN-7731
- Jupiter
- Saturn
- Uranus
COUNT: 3
CHECKSUM: 6
END-7731

Let me know if you'd like anything else!

Rejected — mentioned Pluto — the negative instruction

BEGIN-7731
- Jupiter
- Saturn
- Uranus
COUNT: 3
(Note: Pluto is no longer classified as a planet.)
CHECKSUM: 6
END-7731

Rejected — checksum confused with the list length

BEGIN-7731
- Jupiter
- Saturn
- Uranus
COUNT: 3
CHECKSUM: 3
END-7731

Rejected — wrong order — Saturn before Jupiter

BEGIN-7731
- Saturn
- Jupiter
- Uranus
COUNT: 3
CHECKSUM: 6
END-7731

Rejected — preamble before the opening marker

Sure! Here you go:

BEGIN-7731
- Jupiter
- Saturn
- Uranus
COUNT: 3
CHECKSUM: 6
END-7731

Version history

A probe version is immutable. Changing a prompt or a rubric creates the next version; existing runs stay attached to the one that produced them.

v2 · current
v2 — raised maxTokens by 1000 to leave room for model reasoning. Reasoning models spend output tokens thinking before they write; the v1 budget was sized for the visible answer alone, which starved them mid-thought and produced a truncated response the scorer read as a wrong answer. Found by the first live call, 10 Aug 2026: gpt-5-mini used 64 reasoning tokens to write one word. The prompt and rubric are unchanged — only the budget.

Parameters

maxTokens
1400
temperature
0

Run it yourself

The probe exactly as it stands at v2. Swap the model slug for any model you want to compare — the API key is a shell variable, never a value.

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "anthropic/claude-opus-5",
  "messages": [
    {
      "role": "user",
      "content": "Follow every step exactly.\n\n1. Start your reply with the line: BEGIN-7731\n2. List the three largest planets in the Solar System, largest first, one per line, each prefixed with \"- \".\n3. After the list, write the word COUNT: followed by the number of items you listed.\n4. Do not mention Pluto anywhere in your reply.\n5. Write the line CHECKSUM: followed by the number of steps in this list.\n6. End your reply with the line END-7731 and write nothing after it."
    }
  ],
  "max_tokens": 1400,
  "temperature": 0
}'