tootsies

a discord bot for the tootsies server. ask, recap, discuss, ship features by typing.


Project maintained by mejasonmejason Hosted on GitHub Pages — Theme by mattgraham

Optimizing prompts for Opus: cut and reword, don’t add

The thesis of Epic #849. Read this before you touch any text Opus sees.

Every piece of text the model reads is both a cost and a lever: the constitution, the persona, each per-surface prompt, every tool description, every injected context-block label, even the eval judges. The default instinct when a model misbehaves is to add a rule to correct it. For Opus that instinct is almost always wrong. The right move is to find the existing text that is causing the behavior and cut or reword it.

Why “add a rule” fails on Opus specifically

The shared prompt was tuned, fix by fix, against Sonnet’s failure modes (sycophancy, hedging, stiffness, fence-sitting). Opus does not share those failure modes, and it differs in two ways that make accretion actively harmful:

  1. Opus follows instructions literally. A reductive verb it’s handed (“Compact”, “Synthesize”, “Drop the noise”) it obeys — it compacts. Sonnet largely ignores those words. So a phrase that is inert on Sonnet is a live instruction on Opus.
  2. Opus over-commits. Hand it a maximal instruction (“be exhaustive”, “the complete record”, “breadth so any question has a hit”) and it maximizes — it pads. A counter-rule stacked on top doesn’t cancel the first; now it’s following both, harder.

So when Opus drifts, adding a counter-instruction compounds the problem. It also bloats the prompt, and a longer prompt is itself a cost (tokens, latency, and more surface for cross-rule contradiction).

The method (what to do instead)

  1. Find the cause in existing text; remove or reword it. Don’t write a new rule. Ask: what am I already saying that produces this?
  2. Show before/after verbatim, every edit. No silent rewrites. The diff is the unit of review.
  3. Dry-run on REAL data, not synthetic. Pull production logs/notes (Railway, /debug/query) and run the actual pipeline. A fixture can’t show you what the live distribution does.
  4. Measure the axis you actually care about, with the real machinery. For recall that means max(prose_vector, tags_vector) cosine over the real query set — not a vibe read of the text. For length, the production answer_length telemetry. Name the metric before you edit.
  5. Ablate: cut vs add vs reword, and measure each. (See #858’s ablation table — cutting the regulars block made the symptom worse; an explicit clause fixed it; the real lever was a layer up.) The first thing you blame is usually not the cause.
  6. Length is not the enemy — bloat-without-payoff is. Don’t trim for its own sake. Measure whether the extra text buys the thing (searchability, depth, accuracy). If it doesn’t, it’s pure context cost and should go.
  7. Key everything per model; share only the safety floor. A rule earned against Sonnet’s failures mis-fires on Opus. The constitution HARD rules, the memory fence, and protected-path/preflight guarantees are the one universal base and are never loosened per-model. Everything else is per-model.

Worked examples (this epic)

Evaluating model effectiveness (how to measure, not just what to change)

Half of this epic’s lessons were about evaluation method, because a prompt change is only as trustworthy as the measurement that justified it. What we learned:

  1. Pick the metric and the axis before you edit. “Better” is meaningless. Name the axis — searchability, length, depth-of-record, fabrication rate, adherence — and define how you’ll measure each before touching text, so you can’t rationalize a change after the fact.
  2. Measure with the REAL pipeline, not a proxy. Recall means the production path: max(prose_vector, tags_vector) cosine over the real query set against the real embedding model — because that’s what actually retrieves the note. Eyeballing the prose for “searchability” measures nothing.
  3. n=2 is noise; trust n≈5. A two-sample dry run told a clean story (softened was “~30% shorter and deeper”); n=5 erased it (length a wash, depth inconclusive). Model output is high-variance — a clean small-n result is more likely sampling luck than signal. Run enough samples to see the spread, and report the spread (min–max + sd), not just the mean. A mean without a range hides whether the difference is real.
  4. Validate the metric itself — a broken metric is worse than none. The n=5 depth judge returned coverage counts above the fixed checklist size (depth

    1.0), i.e. it was nonsense, and a careless read would have “concluded” from it. Sanity-bound every derived metric (covered ≤ total) and discard it loudly when it violates the bound, rather than quoting a number that can’t be real.

  5. Isolate the parameter you’re testing. Hold everything else fixed (same source data, same queries, same fixed checklist as the depth denominator) so the only moving part is the one prompt change. A floating denominator (the judge re-deriving the moment count per call) injects variance that swamps the signal.
  6. Ablate cut vs add vs reword separately. Don’t test a bundle. #858’s table (baseline / cut-the-block / add-a-clause) is the model: it showed the obvious suspect (cut the regulars block) made it worse and the real lever was elsewhere. You cannot tell which token did the work unless you move one at a time.
  7. A wash is a result. If the metric says no difference, say so and decide on other grounds (prose hygiene, owner taste) — don’t manufacture a metric win to justify a change you already made. Honest “no measurable effect” is more useful to the next session than an oversold number.
  8. Dry-run on real production data. Pull live notes/logs (/debug/query, Railway EVENT logs); synthetic fixtures don’t carry the live distribution that produces the failure.
  9. A judge is itself a model — validate a FAIL before acting on it. The eval judge can be wrong, including failing a correct answer. Worked example (#893): the hardened recap eval failed an Opus recap for saying “SGA won the 2025-26 MVP” — and I first read that as an Opus weakness (“over-asserts an outside fact”). The statement is true; the recap had correctly grounded a real recent fact about the exact topic the room was debating, which is what recap is supposed to do. The judge erred — it couldn’t web-verify a recent result and treated “couldn’t verify” as fabrication, against its own rubric. Two compounding traps: (a) acting on a single judge verdict as ground truth, and (b) it nearly produced scar tissue — a “lean recap core” rebuild for a weakness that does not exist (the #861 wash STANDS). The bar a fabrication judge must enforce is “contradicted by search → FAIL,” never “couldn’t verify → FAIL.” Before you act on a judge FAIL, check the judge was right.

Eval fidelity: model production, or you measure nothing

An eval is only worth its verdict if it feeds the model what production feeds it and exercises the path production runs. Three failure modes this epic closed (#886, the ask/memory eval lift):

  1. Feed the production-shaped context, not a barer frame. Every real /ask supplies channel_descriptor (“where am I”) + asker_name (“who’s asking”) + the rich message buffer (reactions / reply-graph / author-tags). The synthetic cases passed none of it, so the model saw a thinner frame than any live call and the verdict measured a different distribution. Fix: feed the same context the cog feeds (cogs/ask.py), verbatim shape.
  2. Wire the REAL tools/pipeline — never a reimplementation. eval_ask_tools now builds the real SportsDataHub (ApiSportsProvider + SgoProvider) and calls the real format_scoreboard; the thin handlers just call hub.live_games()/upcoming_games()/recent_finals() exactly as the cog does. A handler that reimplements cog logic in the eval drifts from prod silently — then it grades a fiction. Reuse the production function; the eval’s job is routing + answer quality, not a parallel implementation.
  3. Grade the model actually SHIPPED per surface — read it from the DB (cost discipline). Don’t grade every model every run. The scheduled suite reads each surface’s enabled model from the per-guild /menu Models override (settings KV) via resolve_eval_models/merge_enabled: flagless → the single enabled model (the surface default UNION any guild override), so it pays 2x ONLY when a guild genuinely split the surface across models. Wiring detail worth remembering: GitHub’s runners cannot reach Railway’s internal DATABASE_URL host, so the read uses DATABASE_PUBLIC_URL (the public proxy) and is fully fail-open — no URL / unreachable → fall back to the surface default, never block the suite on a DB hiccup.
  4. Pin the production model in an ad-hoc dry run too, not only in the suite. Item 3 makes the scheduled suite grade the shipped model. A one-off A/B you run by hand needs the same discipline, and it is easy to skip because the default client is right there. Worked example (#2148): the music desk’s watch lane claimed a chart RE-ENTRY it could not prove, and the fix cut the phrase that offered the claim. The first A/B ran on the surface’s DEFAULT model and scored 0/6 on the BEFORE arm — the bad wording produced no bad output, so the result justified nothing and the edit looked unnecessary. Re-run pinned to claude-opus-4-8, the model that composed the live post, the BEFORE arm scored 3/5 and the AFTER arm 0/5. This is the same Opus-vs-Sonnet split the top of this file describes: a phrase that is inert on one model is a live instruction on another. So read the surface’s enabled model first, pin it, and only then believe the numbers. An unpinned A/B measures a model nobody ships.

Validate the gap is real before you build the fix (proportionality)

The fix is usually smaller than the issue makes it sound — measure the gap before building machinery for it. Worked example (#892): the issue asked for a live-embedding behavioral eval of chunked recall, and I started building it. On inspection the gap had mostly already been closed: the chunk mechanism (the two-stage rank_by_similarity rerank, base64-f32 encode/decode) was already deterministically unit-tested in tests/test_embeddings.py and gating CI, and a live-embedding eval would have been dormant anyway (evals.yml has no OPENAI_API_KEY, so embedding cases self-skip). The one genuinely untested seam was the cog wiring — that cogs/ask.py:_vector_recall passes chunk_rerank_top through. The proportional fix was a ~10-line deterministic wiring test, not a live eval harness. Pattern: enumerate what already covers the risk (unit tests, an existing eval, a dormant-in-CI gate) before adding more — sunk-cost on a half-built heavier fix is not a reason to ship it. (Pairs with the CLAUDE.md “proportional to the problem” rule and the #567 resume-spine → 10-line cog_unload example.)

The standing directive

Apply this to every piece of text Opus sees — constitution, persona, per-surface prompts, tool descriptions, context-block labels, eval judges. Each is a lever to be optimized by cutting and rewording for Opus, not a place to accrete rules. When you’re tempted to add: stop, find the cause, cut it instead.

See also

Worked example: the awards write-up (2026-08-30)

The owner asked for two things at once — “cut wording down, it’s too long and repetitive, we repeat the all time context too much”. Both were caused by existing text and existing DATA, not by a missing rule, so both fixes were cuts.

Worked example: the RIAA board takes + scar removal (2026-09-06)

The owner’s report was a style one – “we’ve been writing like AI lately instead of just reporting plainly” – against one shipped card: “2026 has been kind to Future’s catalog: 59 Platinum certifications so far this year, 9 clear of Kanye West’s 50 among today’s top acts, with SZA’s 27 not far behind.” Every part of that sentence traces to a line in utils/riaa_boards.py:compose_inputs. No rule was added. Four lessons, three of them about the SAME mechanism.

Method note: the axis was named before the edit (does the take carry an AI tell, and does it carry a fence the board does not have), it was counted over the REAL shipped corpus first, and the before/after ran on the real boards through the real wired compose path (compose_market_drop with the desk’s own _CONTEXT), scored by the real self-gate.

And one axis was NOT named before the edit, which is why the first rewrite shipped a regression into a PR. The n=5 pass measured the tells and the false fence – the things the owner reported – and nothing else. The invented gap only appeared when the second pass widened to every board and counted every guard. A prompt edit can move an axis you were not watching, so measure the guards the surface already runs, not only the axis you set out to fix.

Count the failure directly, not through the guard that happens to catch it. The invented gap was found through ungrounded_numbers, which flags a figure the source never states. That guard undercounts this failure: on the 2026 board the take wrote “28 ahead of Rod Wave (31)”, a computed gap whose value happens to equal another act’s row count, so it reads grounded and ships. The honest metric was a direct one – any “N clear of / ahead of” the block does not itself state.

The scar in the same file

The owner asked for the scar to come out in the same pass. Four kinds were there, and naming the kinds matters more than the four cuts, because each recurs:

The head also said one thing twice, twice: “do NOT list the rows” beside “do not recite the leaderboard”, and the card’s credit stated both before and after the rule that depends on it.

Worked example: ablating the RIAA framing line by line (2026-09-06)

Owner steer: “deleting prompt lines is the best strategy, inspect more and ship”. So every line of the RIAA board framing was deleted in turn and measured against its own baseline – five candidates, three boards, n=10 per arm, real model, real wired path, real self-gate.

The steer is right about ONE line and wrong about the other four, and the shape of the result is the useful part.

deletion platinum invented gap year identity claim acts named
(baseline) 1/10 0/10 3.0 / 2.6
the runner-up line 0/10 0/10 1.1 / 2.0
“Never invent a figure” 4/10 3/10 3.0
the credit-tag ban 5/10 1/10 3.0
“in a line or two, then stop” 7/10 1/10 2.9
the re-COUNT clause 4/10 1/10 3.0

Four unrelated cuts all degrade the SAME board, so the lever is not any one line. A shorter framing leaves the take room to invent an angle, and the angles it reaches for are the ones the guards exist to catch: a gap the block never stated, and an identity claim. Cutting “Never invent a figure” is the clearest case – it is one short sentence, it looks like boilerplate next to a longer grounding rule that already says the same thing, and removing it quadrupled invented numbers and took identity claims from 0/10 to 3/10.

Only the runner-up line was free of that, and only because deleting it removes the comparison itself. No runner-up in the take means no gap to work out. It is the one deletion shipped, and its cost is stated rather than hidden: act coverage falls from 3.0 names to 1.1 on the Platinum board. The card still draws every row, so the board is not lost – the take just stops narrating it.

One cut was board-dependent, which is its own warning. Deleting “in a line or two, then stop” is the best result on the YEAR board (123 chars against 142, one sentence, no new failures) and the worst on the Platinum board (7/10 invented). A cut measured on one board and shipped everywhere would have looked like a clean win. It appears in seven files across the repo; none of them were touched.

Method note. output_checks.count_sentences reads a decimal point as a sentence end, so any take carrying “336.5M” measures as four sentences. Every length reading here uses characters instead. A metric that is wrong in the same direction on every arm still ranks the arms correctly, but it cannot be reported as a sentence count, and the production answer_length telemetry has the same artifact.

Worked example: the age of the record (2026-09-06)

Symptom. Two X posts off kworb chart rows closed with an age: “pulling 14.2M streams a track over a decade old” and “a catalog cut over four decades old still finding new spikes”. Axiom showed 20 shipped watch-chart takes in 30 days with the same tail (“thirteen years deep”, “sixteen years after its release”, “twenty years on”). The owner asked to cut them.

Cause, found in existing text. The blobs for those lanes carry a position, a move and a play count, and no age. The shared _CONTEXT block (cogs/music_desk.py) carries an AGE / STAGE paragraph written for the Luminate tracking head, whose line really does say “catalog, out 4.2 years”. It opened as a fact about every post:

AGE / STAGE: the head names the release’s milestone – ‘debut week’, ‘first month’, … Put that milestone in your own words, fresh each post … a catalog album’s weekly number is a this-week read (‘46M this week, eleven years deep’) …

Twenty-odd lanes hand the model that paragraph beside a row with no age in it. The model did what it was told and supplied the age from its own knowledge. The _drop_ungrounded digit check already dropped “14 years old”, which is why the shipped tails are all in words.

Fix: reword the cause into a condition, no counter-rule stacked elsewhere.

AGE / STAGE: a Luminate head names the release’s milestone – … WHEN the head states one, put that milestone in your own words … When the material states NO age or milestone – a chart position, a rank move, a play count – say nothing about how old the title is: a chart row carries no release date, so its age is not in the material.

Measured (real model, n=6 per arm, the exact production blob + framing). Bruno Mars “Locked out of Heaven” #50 -> #33 on the Global Weekly: before 3/6 takes carried an age (“pushing 14 years old”, “years past its debut”, “aging like wine”), after 2/6 (“14 years deep”, “the 2012 cut”). The Billie Jean Spotify-jump blob went 0/6 -> 0/6 on this model. So the reword lowers the rate and does not zero it, and n=6 cannot separate 2/6 from 3/6. A prompt rule the model breaks intermittently gets a deterministic backstop (output_checks.ungrounded_age_claim, reason=ungrounded_age), the same call as has_career_claim and the #2081 digit check. Measured against 30 days of shipped takes (n=5774): the pattern matches 67; the lanes whose blob states an age (tracking, lookup_board) stand down, and the rest are true fabrications.

Lesson. A paragraph that is true for ONE lane’s material reads as an instruction on every lane that shares the block. When a shared block describes the material (“the head names X”), write it as a condition on the material, not a fact about it.

Worked example: the called-shot benchmark fence (2026-09-07)

#3014 handed the called-shot compose ONE new fact — what the artist’s own last album opened at. The existing fence justified itself with a sentence that the new fact made false:

before: ... a projection row carries a rank, a unit count, a sales split and a label, and nothing else. Anything more would be you guessing.

Per the method, the stale clause was reworded, not counter-ruled — no second rule stacked on top saying “but you may also use the benchmark”. The first attempt then over-corrected, and the live dry run is what caught it:

attempt 1: ... Two points are not a trend: say what the gap IS, never that it is a rise, a fall, a record or a best-since

Three live composes on the real building chart, all gating 0.92, and one of them wrote “up big from the 18k his last album opened with” — a direct violation that the judge passed without blinking. The rule was wrong, not the model. Handing it two numbers and forbidding it to say which is larger is incoherent: comparing is saying which is bigger, and the comparison is the entire reason the figure was added.

after: ... and the ONE prior opening above. That figure is this artist's OWN last album, and comparing the two numbers is allowed -- it is the reason you have it. What you have is TWO POINTS, not a career: never 'his biggest ever', never 'his best since', and never a claim about where his career is heading

Five further live composes, all 0.92, all landing on arithmetic over the two given numbers (“more than 5x his last album’s 18K opening”) with no career or trajectory claim.

Two lessons, both already in this doc and both re-earned the hard way:

  1. Ban the thing you actually mean. “Never say it rose” was a proxy for “never claim a career trajectory”, and the proxy caught the legitimate case. Opus follows literally, so an over-broad ban is either obeyed (losing the feature) or ignored (as here) — and an ignored rule is still prompt bloat.
  2. A passing self-gate is not a passing fence. The judge scored the violating post 0.92 because the post was accurate and well-formed. Only reading the output against the rule showed the rule and the output disagreeing.

Worked example: the gap with nobody in it (2026-09-09)

The projected-units board (chart_boards.projected_top_sales) shipped this to X on 2026-09-08, self-gate 0.88:

Projected Billboard 200 units for Sep. 19 span ~69K to ~34.8K, with HITS Daily Double forecasting a ~4K gap at the top.

Ella Langley’s “Dandelion” was the #1 at ~69K and the post never named it (owner: “this should’ve narrated dandelion”). The cause was in the existing text, twice over. The board’s headline (THE FACT the model reads) opened with the span and named the leader second, and its steer said:

before: Then lead with the SPREAD rather than who is #1: the ranking board already said that, and a units order IS the chart order, so without its own angle this is the same post twice. The fact here is the DISTANCE between these numbers

That clause was written in #2098 to stop four projection boards opening with the same leader. It over-reached: “rather than who is #1” reads, to a literal model, as “do not name who is #1”, and the take obeyed. The reword keeps the angle and restores the subject (_PROJ_STEERS["proj_top_sales"]):

after: Then NAME THE ALBUM ON TOP and its units figure: the leader is the subject of this take, and a gap with nobody named in it is not a post. YOUR ANGLE IS THE DISTANCE -- how far clear of #2 that album is, or how tightly the rest of the board is packed. The ranking board owns positions and moves; you own the leader's units and the gaps.

The headline now opens with the leader too ("Dandelion" by Ella Langley is tracking for ~69K units, the most on the projected Billboard 200 for Sep 19, per HITS, ~4K clear of the #2 album. The top 10 runs from ~69K down to ~34.8K.). The class sweep found the same clause on the platform play-count steer (_PLATFORM_STEERS["platform_streams"]) and reworded it the same way.

Live dry run on the same HITS document (scripts/dryrun_chart_boards.py --compose):

Hits Daily Double has Ella Langley’s “Dandelion” leading the projected Billboard 200 for Sep 19 at ~69K units, only about 4K clear of Olivia Rodrigo’s ~65K, tightest race atop the chart in a minute. – gate 0.88, SHIP

The rank board’s take from the same run named Dandelion too, and the two passed the wording-only dedup the chart cards ride (SequenceMatcher 0.43 against a 0.6 threshold; 0.49 against the rank take that actually shipped on 2026-09-08). The lesson is the “ban the thing you actually mean” one again: the thing meant was “do not build the take around the leader’s rank”, and the proxy “rather than who is #1” banned the subject.

Worked example: a clause the prompt demanded and the lane could not supply (2026-09-14)

Owner steer, on three shipped @tootsiesbar posts: “cut second sentence catalogue talk”.

The takes:

Billie Eilish cleared 8 billion streams on the year, a catalog run that refuses to let go. Kanye West cleared 8.3 billion streams on the year, a massive catalog-and-current run. Tyler, The Creator cleared 3.2 billion streams on the year, another massive catalog-and-current run.

The lane is the music desk’s ANNUAL year-to-date milestone (cogs/music_desk.py, _CONTEXT MID-YEAR MILESTONE + the streams-milestone note). Four of four annual takes that night ended in a clause that restates the figure in words. The self-gate scored them 0.90 and called the filler “one clause of real context”, so nothing downstream was going to catch it.

The cause was a mandate, not a missing ban. The paragraph read '<artist> cleared <X> streams on the year' + ONE short trailing clause, then STOP, and it named three sources for that clause: the momentum trend, the year-end projection, the rank. All three are withheld on this lane – momentum is gated to the non-annual streams facts, the projection is banned by the streams-milestone note in the same context, and a rank only arrives when a chart or qualifier block lands. So on most slots the model was told to write a clause and handed no fact to write it from. It wrote one anyway.

The paragraph already carried the counter-rule. “never a closing flourish” sat two clauses after the mandate and lost to it – the canonical shape this file warns about: a ban stacked beside the instruction that causes the behaviour does not cancel it.

The fix is one reword: the clause is CONDITIONAL on a fact. The format line now stops at the figure, and a trailing clause is allowed only to carry a hard fact a context block states. Given none, the post ends at the figure. The streams-milestone note appended for a measured report lost its own half of the sanction (“one sharp human beat about how big it is”), but only on the annual lane (streams_milestone_note(annual=...)). A NON-annual measured report heads its blob with a grounded release stage – “debut week”, “first month, 3 weeks in”, “catalog, out 4.2 years” – and the shared AGE / STAGE paragraph asks the take to put that milestone in its own words. The stage sits in the BLOB, not in a context block, so the annual wording would have stripped the stage framing off every weekly report whose fail-open supplemental blocks came back empty. The first version of this PR missed that and the Codex reviewer caught it. The lesson generalizes past this fix: a rule keyed to “a fact from a context block” silently excludes the reading’s own head, and a shared block is appended on more lanes than the one you are looking at – so check which lanes it reaches before you narrow what it allows. The annual head carries “year-to-date” and no age at all, so it loses nothing. Dry-run both halves: the non-annual takes still read “four years deep and barely slowing down” / “six weeks in and still holding a real number”.

Measured on the real wired path (compose_market_drop with the desk’s own _CONTEXT and the real blob; Sonnet, the lane’s default; the three reported artists, n=5 each, run with and without a Billboard-standing block):

arm takes ending in an invented clause
before 18 18
after 30 0

The chart fact still rides when the block names that artist (“cleared 8.3 billion streams on the year, and ‘Graduation’ is still charting at No. 148 on the Billboard 200 this week”), and the two artists the block did NOT name correctly ignored it.

Two guards were measured, not assumed, because a shorter take could have gone dark two ways. The self-gate scores the bare figure 0.90-0.92 against the real blob (the same score the flourish got), so nothing new is dropped. The arbitrated dedup still ships four same-shape annual takes in a row: the mechanical gate matches the shape in both arms, and the judge overturns it on the artist and the number. A first dedup run said the opposite and was wrong – the harness called the judge with the wrong signature, so every judge call raised and the fail-closed direction left the mechanical verdict standing. Check that your harness’s judge actually answers before you read its verdict.

A shared block is read by more surfaces than the one you are looking at (2026-09-14)

The second Codex finding on #3252, and the more general of the two. The clause rule was written as “a HARD FACT a context block below states”. That phrasing assumes the surface HAS context blocks. cogs/music_desk.py:_CONTEXT is shared with cogs/music_alert.py, whose _fire passes it with no supplemental blocks at all – so on the alert path the rule named a source that does not exist, and the only material the take had (the market blob, carrying the line’s own move) fell outside what the rule allowed. Fixed by keying the rule to the MATERIAL rather than to a block that may not be there.

The predicted symptom did not reproduce, and saying so matters more than the fix. The review expected the alert to discard the move and post the already-known year-to-date count. Measured on the real classifier + the real compose (an annual line_move, Sonnet, n=8 per arm), the figure rode in 8/8 of every arm, the flagged wording included – the model used the blob’s pace line regardless of what the rule said it could use. The fix is right because the wording was wrong, not because it rescued a post.

What the measurement DID find is a weakness that predates the PR. An annual line_move alert fires because the market-implied number moved, and the take almost never says so:

arm names the MOVE names the figure
main (pre-PR control) 0/8 8/8
the flagged wording 0/8 8/8
the fix 2/8 8/8

main scoring 0/8 is the whole point of running a control: without it this reads as a regression the PR caused, and it is not one. The cause is upstream of any clause rule – the MID-YEAR FORMAT makes the year-to-date figure the mandatory lead and demotes the move to a trailing clause, on a lane whose entire reason for firing is the move. That is a separate change on a separate lane, so it was FILED, not folded into this PR.

The scar in the same block (2026-09-14)

Owner steer in the same session: “cut scars”. Two of the four kinds this file names were in cogs/music_desk.py:_CONTEXT, and both are the FIRST kind – a rule with two definitions.

Ablated on the real wired path before cutting (compose_market_drop, the desk’s own _CONTEXT + the real blob, Sonnet, n=5) across three cases, so a cut that fixes one axis could not quietly break another:

case axis baseline scar cut
annual, no chart block ends at the figure 5/5 5/5
annual, chart block names the artist quotes the chart fact 5/5 5/5
weekly catalog album states the release stage 5/5 5/5

Mean length moved 9.0 → 9.0, 20.8 → 22.4 and 22.6 → 19.8 words: noise in both directions at n=5, no trend. So this cut is justified on prose hygiene alone, not on a measured win – the same honesty the #855 entry insists on. 287 characters out of a 3917-character block, and one contradiction the model could have resolved either way is gone.

Two candidates were examined and NOT cut, because each still earns its place: the AGE / STAGE paragraph’s closing “a chart row carries no release date, so its age is not in the material” (it reads like engineering detail, but it is the reason the rule holds, and that lane’s whole failure mode was the model supplying an age the row never stated), and the streams-milestone note’s “no exact digit-string” (it fences a FABRICATION the model can produce unprompted, not an echo of the material, so the stale-fence argument does not apply).

Worked example: the weekly streams beat (2026-09-29)

Owner steer on weekly streams posts that ended “a steady start”: “any commentary should come from the stats, like the qualifiers. If it’s just an empty human beat, remove it.”

The #3252 fix above made the clause CONDITIONAL on a fact, but only on the annual lane. The weekly lane kept three sources of an empty beat, and all three were cut, not countered:

A rank from the background block and a verified qualifier still ride as the one clause, because they are stats.

Worked example: the credit moves to the card (2026-09-29)

Owner steer, on an X post that ended “HITS Daily Double has it there.”: “drop them from the text post and make them on the card. The pill is enough.”

The cause was a mandate. attribution.credit_rule said “CREDIT HITS, AND DO IT IN THE SENTENCE”, and gave in-voice examples that named the source (“hits has her tracking ~74k”). A ban on the press-release tag sat beside it. The model obeyed the mandate and moved the name around the sentence. The rule now says only that the card names the data source, so the line names none.

A name in the instruction is a name in the take. The first rewrite said “THE CARD NAMES HITS DAILY DOUBLE AS THE SOURCE”. On the reconciliation lane, Sonnet still wrote “HITS Daily Double projected #3” on 3 of 8 takes. Three more places carried the name to the model: called_shot.describe (the take’s topic line), the block’s per-publisher labels, and the rule itself. Each was reworded by ROLE (“the industry projected #3”), and the rule stopped naming any source. The next run was 0 of 10.

Measured on the real compose inputs of six lanes (compose_market_drop, the desk’s own _CONTEXT + each lane’s framing):

arm model takes naming a source
before (main) terra 23/30
after terra 0/48
after Sonnet 1/48

The one Sonnet miss was on the reconciliation lane. The widened has_attribution_tag drops it after one _detag_credit re-ask, which now runs on every desk lane.

Platform boards are the exception, and it is deliberate. On a Spotify or Apple Music board the platform IS the chart (“#1 on spotify’s us chart”), so naming it is the fact, not a credit. Those boards keep their own chart phrasing (platform_chart_phrasing).