Instruction following · v1
Emit a terminator and write nothing after it
Summarise a paragraph, then close with a fixed sentinel and stop. Tests the stop condition that machine-consumed output depends on. The failure that matters is not a missing sentinel — it is a correct one followed by one more helpful sentence.
Results
Not yet run
No model has been run against this probe yet, so there are no results to show. The definition, rubric and scoring below are complete and final for v1 — this page is published now so that the test is on the record before any score exists, rather than appearing alongside one.
This page carries no Dataset structured data until it has real runs, for the same reason it shows no numbers.
Prompt
Exactly what every model receives. Nothing else is sent.
system
You produce output for a program, not for a person. When given an output contract, follow it exactly. Nothing may follow the terminator.
user
Summarise the paragraph below in a single sentence. Then, on a new line, output exactly this terminator: <<END>> Write nothing after the terminator. Do not repeat it. Do not add a closing remark, an offer to help further, or any explanation of what you did. PARAGRAPH The migration ran for six hours against the replica before anyone noticed that the replica had been promoted the previous week, which meant the whole rehearsal had been executed against production data with the safety checks disabled. No rows were lost, because the script only ever wrote to a shadow table, but the incident review changed the team's rule from "rehearse on the replica" to "rehearse on a snapshot restored to a named host", since a replica's role can change without anyone touching the runbook.
Rubric
Published so you can disagree with it. A score you cannot argue with is a rumour.
Four properties, scored independently. The terminator must be present, must be the last thing in the output, must appear exactly once, and must be preceded by an actual summary. Trailing whitespace after the terminator is tolerated, being a transport artifact rather than content.
- The terminator <<END>> is present
- Nothing follows the terminator except whitespace
- The terminator appears exactly once
- No conversational preamble before the summary
- A summary is actually present before the terminator
Scoring
Scored by regex rules. Deterministic — the same output always produces the same score.
Rules, in order
emitted_the_terminator must match /<<END>>/ Without the sentinel a consumer has no way to know the response is complete. nothing_after_terminator must match /<<END>>\s*$/ The rule this probe exists for. Content after the sentinel is either silently discarded by a consumer reading up to it, or breaks one reading past it. terminator_appears_once must NOT match /<<END>>[\s\S]*<<END>>/ A repeated sentinel makes the first occurrence a false end-of-message. no_preamble must NOT match /^\s*(sure|certainly|of course|okay|here'?s|here is)\b/i A preamble shifts every byte offset a caller may have computed, and is content the contract did not ask for. summary_present must match /^[\s\S]*\S[\s\S]*<<END>>/ Emitting the terminator alone satisfies every other rule while answering nothing.
Worked examples
Hand-written outputs the rubric is tested against on every build.
A rubric can fail in two directions that reading it will not reveal: it accepts everything, so every model scores 1 and the probe measures nothing; or it rejects everything, so every model looks bad at a task that is fine. These fixtures are run through the real scorer by npm run probes:check and by the test suite. The correct answer must score 1, and every wrong answer must not.
Correct — must score 1.00
A six-hour migration rehearsal was unknowingly run against production data after the replica was promoted, causing no data loss but prompting a rule change to rehearse on a snapshot restored to a named host. <<END>>
Rejected — correct terminator, then one more helpful sentence — the failure this probe is for
A migration rehearsal ran against production data because the replica had been promoted, causing no data loss but changing the team's rehearsal rule. <<END>> Let me know if you'd like a shorter version.
Rejected — repeated the terminator
<<END>> A migration rehearsal ran against promoted-replica production data with no rows lost. <<END>>
Rejected — conversational preamble before the summary
Sure! Here's your summary: A migration rehearsal was accidentally run against production data after a replica promotion. <<END>>
Rejected — omitted the terminator entirely
A migration rehearsal ran against production data because the replica had been promoted the previous week; no rows were lost, but the rehearsal rule was changed.
Rejected — terminator only, no summary
<<END>>
Version history
A probe version is immutable. Changing a prompt or a rubric creates the next version; existing runs stay attached to the one that produced them.
- v1 · current
- Initial version.
Parameters
- maxTokens
- 1400
- temperature
- 0
Run it yourself
The probe exactly as it stands at v1. Swap the model slug for any model you want to compare — the API key is a shell variable, never a value.
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-opus-5",
"messages": [
{
"role": "system",
"content": "You produce output for a program, not for a person. When given an output contract, follow it exactly. Nothing may follow the terminator."
},
{
"role": "user",
"content": "Summarise the paragraph below in a single sentence.\n\nThen, on a new line, output exactly this terminator:\n\n<<END>>\n\nWrite nothing after the terminator. Do not repeat it. Do not add a closing remark,\nan offer to help further, or any explanation of what you did.\n\nPARAGRAPH\nThe migration ran for six hours against the replica before anyone noticed that the\nreplica had been promoted the previous week, which meant the whole rehearsal had been\nexecuted against production data with the safety checks disabled. No rows were lost,\nbecause the script only ever wrote to a shadow table, but the incident review changed\nthe team'\''s rule from \"rehearse on the replica\" to \"rehearse on a snapshot restored to\na named host\", since a replica'\''s role can change without anyone touching the runbook."
}
],
"max_tokens": 1400,
"temperature": 0
}'