a discord bot for the tootsies server. ask, recap, discuss, ship features by typing.
The thesis of Epic #849. Read this before you touch any text Opus sees.
Every piece of text the model reads is both a cost and a lever: the constitution, the persona, each per-surface prompt, every tool description, every injected context-block label, even the eval judges. The default instinct when a model misbehaves is to add a rule to correct it. For Opus that instinct is almost always wrong. The right move is to find the existing text that is causing the behavior and cut or reword it.
The shared prompt was tuned, fix by fix, against Sonnet’s failure modes (sycophancy, hedging, stiffness, fence-sitting). Opus does not share those failure modes, and it differs in two ways that make accretion actively harmful:
So when Opus drifts, adding a counter-instruction compounds the problem. It also bloats the prompt, and a longer prompt is itself a cost (tokens, latency, and more surface for cross-rule contradiction).
/debug/query) and run the actual pipeline. A fixture can’t show you what the
live distribution does.max(prose_vector, tags_vector) cosine over the real query
set — not a vibe read of the text. For length, the production answer_length
telemetry. Name the metric before you edit.has_career_block / career_fact), not to delete it: with no fact in hand the ban is
still right.
And the ban was not even doing its job. Measured live on the projection lane, the
flat-ban arm wrote “a solid new entry with no chart history to lean on yet” about Dolly
Parton – an act with fifty-one Billboard 200 entries. A ban on career claims does not
produce silence about a career; it produces an UNGROUNDED GUESS in the other direction.
Supplying the fact is what fixed it: “it’d be her 52nd entry on that chart.”Measure the FIXTURE too, not only the change (same epic). The dry-run room blob said “this shift is dragging”, and the after-arm kept tacking on “, perfect for a slow shift”. That read as the fix failing. With a neutral room the same prompt dropped it from 5/5 lines to 2/5, and both survivors added information rather than grading. The cue was in the harness. A before/after where only one shape resists is a reason to check what the fixture is feeding the model.
has_prior_position) takes the take’s own rank and refuses any other position.forecast hint gained “is at ~61.5K units” as its model line. The rule
was NOT written as a list of banned verbs (the #1140 lesson: an enumerated list summons
what it names); the deterministic backstop (has_market_decline) carries the list
instead, where naming the words is the point. Then the dry run on the real blob +
history (n=5 per model) showed the reword was not the fix. With the earlier figure
still in the history block, Sonnet wrote the decline 3/5 (“is down to ~61.5K”, “down
from the ~65.2K mark”) and the guild’s GPT tier 2/5; on the Emmy board GPT went 5/5 ->
0/5 once the blob stopped carrying “down 8 pts 24h”, but Sonnet wrote “down from 62%”
off the history line. Every remaining decline was written off a figure the prompt
handed over. So the figures left the history block (_withhold_figures), and the
leak went to 0/6 on both models – and both models then returned EMPTY 6/6, on both
shapes, because the block’s other job was “return EMPTY if the numbers haven’t moved”
and it could no longer see the numbers. A block that must carry a fact for one job
and must not carry it for another is two blocks: the “did it move” judgment is now its
own Haiku call (market_history_news, 15/15 on the live cases) that reads the figures,
and the compose reads the lines with every figure withheld and a block that no longer
asks it to judge anything. When a cut breaks a second job the same text was doing,
split the readers; do not put the fact back.One example taught a wrong FACT (newsroom certification lane, 2026-09-15). The
Jennie RUBY card heroed “1x” over a take that said RIAA Gold. Nothing in the prompt told
the classifier to hero a bare multiplier. Its one certification example did: "figure":
"7x" for a 7x Platinum wire. The model copied the shape onto every rung, and the rungs
that carry no multiplier took a wrong award — “1x” is 1,000,000 units on the RIAA ladder
and Gold is 500,000. Measured over the 30 days before the report, 217 of 283
certification classifications answered a bare “Nx”, and 54 of those sat on Gold, Diamond
or Silver. The fix added no rule. The example now reads "figure": "7x Platinum", and a
second example shows the failing shape (a Gold, "figure": "Gold"). Measured on the real
classifier over seven live wires, n=4: 24 of 24 bare multipliers before, 24 of 24 award
names after. An example is a fact claim, not only a format. Check each one against
the domain, not only against the schema.
docs/opus_length_drift.md). Opus ran long on thread-continuation roasts via a
defend-then-soften reflex. Two things make this the textbook case:
_OPUS_REGULARS (the don’t-punch-down block) and cut it.
The ablation said otherwise: cutting it made the symptom worse (bf_hater
3.0 → 3.3) and wasn’t free (it’s the lane/villain guard). The real driver was
a length-sanctioning roast rule a layer up (“a full takedown … is the fun”).constitution.py CALIBRATION, Sonnet-era,
never Opus-optimized); the length rule lives in Layer 3 (_OPUS_LENGTH). Adding
a length clause to Layer 3 would just contradict Layer 1 one layer down —
two rules fighting. The fix is to reword the Layer 1 line so a roast lands
sharp in a line or two and drops the defend-then-soften reflex. (The same
investigation surfaced Layer-1↔Layer-3 duplication — STAY IN YOUR LANE in
both, DATA INTEGRITY’s 9 bullets vs the lean _OPUS_GROUNDING — i.e. more text
to consolidate, not add.)docs/opus_discourse_eval.md). The first measurement compared
Opus on the Sonnet-tuned 43k prompt against Sonnet, found Opus worse (a same-topic
head-to-head ~6-2), and nearly concluded “keep Sonnet, don’t build.” That measured
Opus on the wrong prompt — the exact confound this epic exists to fix. Given a
lean discourse core, Opus flips to the strongest of the three (open-question
1/8 → 8/8, flag-planted 10/16 → 15/16, open-loop Opus 3/4 vs Sonnet 0/4 on the real
surface). Four sub-traps it hit, each generalizable:
_OPUS_CITE / _OPUS_TOOLDISC into discourse leaked meta-commentary into a room
post — their “the user” / “your answer” 1:1 framing reads wrong in a broadcast post
(A/B META-leak 2/8 → 0/8 once swapped for a discourse-framed line). Reuse only the
surface-NEUTRAL component (_OPUS_REGULARS); reframe the rest for the surface._call) looked clean, but the live discourse() path has a forced-search retry
(thinking OFF) the offline harness never exercised — it degraded the lean output
(~4.0 sentences + selection-narration meta-leak vs the primary’s ~2.3). Skipping
that retry on the lean path was the fix. Grade the actual pipeline, not the prompt
in isolation.utils.x_reply_draft.tic_hits, calibrated on both corpora –
it trips 39% of a live week’s drafts vs 6% of the owner’s own posted replies) plus ONE
re-ask that hands the model the exact words to drop took tics 9/36 → 2/36 (25% →
6%, p=0.046, register a wash) – i.e. down to the owner’s own rate. The pair is the
lesson: “vary it, here’s everything you said” PRIMES the habit; “you wrote ‘the
real X’, say it without that” removes it. Same axis, same surface, opposite results,
and the difference is specificity. The retry is taken ONLY when it comes back clean,
so the guard can never degrade a good line – a fix that can only help is worth
shipping at a lower evidence bar than one that trades.Half of this epic’s lessons were about evaluation method, because a prompt change is only as trustworthy as the measurement that justified it. What we learned:
max(prose_vector, tags_vector) cosine over the real query set against the
real embedding model — because that’s what actually retrieves the note. Eyeballing
the prose for “searchability” measures nothing.1.0), i.e. it was nonsense, and a careless read would have “concluded” from it. Sanity-bound every derived metric (covered ≤ total) and discard it loudly when it violates the bound, rather than quoting a number that can’t be real.
/debug/query,
Railway EVENT logs); synthetic fixtures don’t carry the live distribution that
produces the failure.An eval is only worth its verdict if it feeds the model what production feeds it and exercises the path production runs. Three failure modes this epic closed (#886, the ask/memory eval lift):
/ask
supplies channel_descriptor (“where am I”) + asker_name (“who’s asking”) +
the rich message buffer (reactions / reply-graph / author-tags). The synthetic
cases passed none of it, so the model saw a thinner frame than any live call and
the verdict measured a different distribution. Fix: feed the same context the
cog feeds (cogs/ask.py), verbatim shape.eval_ask_tools
now builds the real SportsDataHub (ApiSportsProvider + SgoProvider) and
calls the real format_scoreboard; the thin handlers just call
hub.live_games()/upcoming_games()/recent_finals() exactly as the cog does.
A handler that reimplements cog logic in the eval drifts from prod silently —
then it grades a fiction. Reuse the production function; the eval’s job is
routing + answer quality, not a parallel implementation./menu Models override
(settings KV) via resolve_eval_models/merge_enabled: flagless → the single
enabled model (the surface default UNION any guild override), so it pays 2x ONLY
when a guild genuinely split the surface across models. Wiring detail worth
remembering: GitHub’s runners cannot reach Railway’s internal DATABASE_URL
host, so the read uses DATABASE_PUBLIC_URL (the public proxy) and is fully
fail-open — no URL / unreachable → fall back to the surface default, never block
the suite on a DB hiccup.claude-opus-4-8, the model that composed the live post, the BEFORE arm scored
3/5 and the AFTER arm 0/5. This is the same Opus-vs-Sonnet split the top of
this file describes: a phrase that is inert on one model is a live instruction
on another. So read the surface’s enabled model first, pin it, and only then
believe the numbers. An unpinned A/B measures a model nobody ships.The fix is usually smaller than the issue makes it sound — measure the gap before
building machinery for it. Worked example (#892): the issue asked for a
live-embedding behavioral eval of chunked recall, and I started building it. On
inspection the gap had mostly already been closed: the chunk mechanism (the
two-stage rank_by_similarity rerank, base64-f32 encode/decode) was already
deterministically unit-tested in tests/test_embeddings.py and gating CI, and a
live-embedding eval would have been dormant anyway (evals.yml has no
OPENAI_API_KEY, so embedding cases self-skip). The one genuinely untested seam
was the cog wiring — that cogs/ask.py:_vector_recall passes
chunk_rerank_top through. The proportional fix was a ~10-line deterministic
wiring test, not a live eval harness. Pattern: enumerate what already covers the
risk (unit tests, an existing eval, a dormant-in-CI gate) before adding more —
sunk-cost on a half-built heavier fix is not a reason to ship it. (Pairs with the
CLAUDE.md “proportional to the problem” rule and the #567 resume-spine → 10-line
cog_unload example.)
Apply this to every piece of text Opus sees — constitution, persona, per-surface prompts, tool descriptions, context-block labels, eval judges. Each is a lever to be optimized by cutting and rewording for Opus, not a place to accrete rules. When you’re tempted to add: stop, find the cause, cut it instead.
docs/opus_length_drift.md (#856/#858) — the worked investigation behind the
ask-roast example: the full ablation table, the 4-layer prompt assembly, the
cross-layer tension, and verbatim before/after rewords of the Layer-1 roast rule.The owner asked for two things at once — “cut wording down, it’s too long and repetitive, we repeat the all time context too much”. Both were caused by existing text and existing DATA, not by a missing rule, so both fixes were cuts.
_winner_receipts fed
the winner’s top_post_text — their all-time most-reacted message — into a prompt that
says “THIS period”. The model was doing as told with what it was given: 5/5 hype-ups
quoted a months-old graduation post as the week’s moment. No rule would have fixed that.
Diffing the seed’s highlight reel against the window-start snapshot gives the real
this-period post, and the leak went to 0/5. Pattern: when a model keeps reaching for the
wrong period, check what period the DATA is from before you write a rule about it.The owner’s report was a style one – “we’ve been writing like AI lately instead of just
reporting plainly” – against one shipped card: “2026 has been kind to Future’s catalog:
59 Platinum certifications so far this year, 9 clear of Kanye West’s 50 among today’s top
acts, with SZA’s 27 not far behind.” Every part of that sentence traces to a line in
utils/riaa_boards.py:compose_inputs. No rule was added. Four lessons, three of them
about the SAME mechanism.
ungrounded_numbers drops that before the self-gate, so the cost
is a spent compose. Invented-number rate, n=8 per wording on the same board: the original
2/8, the first rewrite 6/8.output_checks.has_row_arithmetic reads as a sum of the drawn rows and drops before the
self-gate (live 2026-09-05 17:06:34Z, one compose spent; the 18:03Z retry shipped only
because it happened to write “spanning”). Rewording the verb to “spans” did NOT fix it:
n=5, 5 of 5 still said “combined”. Putting the word SPANNING in the instruction did:
5 of 5 said “spanning”, 0 of 5 said “combined”. Seeding the word you want beats
removing the word you don’t, and it beats banning it.Method note: the axis was named before the edit (does the take carry an AI tell, and does
it carry a fence the board does not have), it was counted over the REAL shipped corpus
first, and the before/after ran on the real boards through the real wired compose path
(compose_market_drop with the desk’s own _CONTEXT), scored by the real self-gate.
And one axis was NOT named before the edit, which is why the first rewrite shipped a regression into a PR. The n=5 pass measured the tells and the false fence – the things the owner reported – and nothing else. The invented gap only appeared when the second pass widened to every board and counted every guard. A prompt edit can move an axis you were not watching, so measure the guards the surface already runs, not only the axis you set out to fix.
Count the failure directly, not through the guard that happens to catch it. The
invented gap was found through ungrounded_numbers, which flags a figure the source never
states. That guard undercounts this failure: on the 2026 board the take wrote “28 ahead of
Rod Wave (31)”, a computed gap whose value happens to equal another act’s row count, so it
reads grounded and ships. The honest metric was a direct one – any “N clear of / ahead of”
the block does not itself state.
The owner asked for the scar to come out in the same pass. Four kinds were there, and naming the kinds matters more than the four cuts, because each recurs:
_number_grounding reads the framing, so a singles take quoting 10 would have
passed the guard. The claim survives; the proof is gone.The head also said one thing twice, twice: “do NOT list the rows” beside “do not recite the leaderboard”, and the card’s credit stated both before and after the rule that depends on it.
Owner steer: “deleting prompt lines is the best strategy, inspect more and ship”. So every line of the RIAA board framing was deleted in turn and measured against its own baseline – five candidates, three boards, n=10 per arm, real model, real wired path, real self-gate.
The steer is right about ONE line and wrong about the other four, and the shape of the result is the useful part.
| deletion | platinum invented gap | year identity claim | acts named |
|---|---|---|---|
| (baseline) | 1/10 | 0/10 | 3.0 / 2.6 |
| the runner-up line | 0/10 | 0/10 | 1.1 / 2.0 |
| “Never invent a figure” | 4/10 | 3/10 | 3.0 |
| the credit-tag ban | 5/10 | 1/10 | 3.0 |
| “in a line or two, then stop” | 7/10 | 1/10 | 2.9 |
| the re-COUNT clause | 4/10 | 1/10 | 3.0 |
Four unrelated cuts all degrade the SAME board, so the lever is not any one line. A shorter framing leaves the take room to invent an angle, and the angles it reaches for are the ones the guards exist to catch: a gap the block never stated, and an identity claim. Cutting “Never invent a figure” is the clearest case – it is one short sentence, it looks like boilerplate next to a longer grounding rule that already says the same thing, and removing it quadrupled invented numbers and took identity claims from 0/10 to 3/10.
Only the runner-up line was free of that, and only because deleting it removes the comparison itself. No runner-up in the take means no gap to work out. It is the one deletion shipped, and its cost is stated rather than hidden: act coverage falls from 3.0 names to 1.1 on the Platinum board. The card still draws every row, so the board is not lost – the take just stops narrating it.
One cut was board-dependent, which is its own warning. Deleting “in a line or two, then stop” is the best result on the YEAR board (123 chars against 142, one sentence, no new failures) and the worst on the Platinum board (7/10 invented). A cut measured on one board and shipped everywhere would have looked like a clean win. It appears in seven files across the repo; none of them were touched.
Method note. output_checks.count_sentences reads a decimal point as a sentence end,
so any take carrying “336.5M” measures as four sentences. Every length reading here uses
characters instead. A metric that is wrong in the same direction on every arm still ranks
the arms correctly, but it cannot be reported as a sentence count, and the production
answer_length telemetry has the same artifact.
Symptom. Two X posts off kworb chart rows closed with an age: “pulling 14.2M streams a track over a decade old” and “a catalog cut over four decades old still finding new spikes”. Axiom showed 20 shipped watch-chart takes in 30 days with the same tail (“thirteen years deep”, “sixteen years after its release”, “twenty years on”). The owner asked to cut them.
Cause, found in existing text. The blobs for those lanes carry a position, a move and a
play count, and no age. The shared _CONTEXT block (cogs/music_desk.py) carries an AGE /
STAGE paragraph written for the Luminate tracking head, whose line really does say
“catalog, out 4.2 years”. It opened as a fact about every post:
AGE / STAGE: the head names the release’s milestone – ‘debut week’, ‘first month’, … Put that milestone in your own words, fresh each post … a catalog album’s weekly number is a this-week read (‘46M this week, eleven years deep’) …
Twenty-odd lanes hand the model that paragraph beside a row with no age in it. The model
did what it was told and supplied the age from its own knowledge. The _drop_ungrounded
digit check already dropped “14 years old”, which is why the shipped tails are all in words.
Fix: reword the cause into a condition, no counter-rule stacked elsewhere.
AGE / STAGE: a Luminate head names the release’s milestone – … WHEN the head states one, put that milestone in your own words … When the material states NO age or milestone – a chart position, a rank move, a play count – say nothing about how old the title is: a chart row carries no release date, so its age is not in the material.
Measured (real model, n=6 per arm, the exact production blob + framing). Bruno Mars
“Locked out of Heaven” #50 -> #33 on the Global Weekly: before 3/6 takes carried an age
(“pushing 14 years old”, “years past its debut”, “aging like wine”), after 2/6 (“14 years
deep”, “the 2012 cut”). The Billie Jean Spotify-jump blob went 0/6 -> 0/6 on this model.
So the reword lowers the rate and does not zero it, and n=6 cannot separate 2/6 from 3/6.
A prompt rule the model breaks intermittently gets a deterministic backstop
(output_checks.ungrounded_age_claim, reason=ungrounded_age), the same call as
has_career_claim and the #2081 digit check. Measured against 30 days of shipped takes
(n=5774): the pattern matches 67; the lanes whose blob states an age (tracking,
lookup_board) stand down, and the rest are true fabrications.
Lesson. A paragraph that is true for ONE lane’s material reads as an instruction on every lane that shares the block. When a shared block describes the material (“the head names X”), write it as a condition on the material, not a fact about it.
#3014 handed the called-shot compose ONE new fact — what the artist’s own last album opened at. The existing fence justified itself with a sentence that the new fact made false:
before:
... a projection row carries a rank, a unit count, a sales split and a label, and nothing else. Anything more would be you guessing.
Per the method, the stale clause was reworded, not counter-ruled — no second rule stacked on top saying “but you may also use the benchmark”. The first attempt then over-corrected, and the live dry run is what caught it:
attempt 1:
... Two points are not a trend: say what the gap IS, never that it is a rise, a fall, a record or a best-since
Three live composes on the real building chart, all gating 0.92, and one of them wrote “up big from the 18k his last album opened with” — a direct violation that the judge passed without blinking. The rule was wrong, not the model. Handing it two numbers and forbidding it to say which is larger is incoherent: comparing is saying which is bigger, and the comparison is the entire reason the figure was added.
after:
... and the ONE prior opening above. That figure is this artist's OWN last album, and comparing the two numbers is allowed -- it is the reason you have it. What you have is TWO POINTS, not a career: never 'his biggest ever', never 'his best since', and never a claim about where his career is heading
Five further live composes, all 0.92, all landing on arithmetic over the two given numbers (“more than 5x his last album’s 18K opening”) with no career or trajectory claim.
Two lessons, both already in this doc and both re-earned the hard way:
The projected-units board (chart_boards.projected_top_sales) shipped this to X on
2026-09-08, self-gate 0.88:
Projected Billboard 200 units for Sep. 19 span ~69K to ~34.8K, with HITS Daily Double forecasting a ~4K gap at the top.
Ella Langley’s “Dandelion” was the #1 at ~69K and the post never named it (owner: “this should’ve narrated dandelion”). The cause was in the existing text, twice over. The board’s headline (THE FACT the model reads) opened with the span and named the leader second, and its steer said:
before:
Then lead with the SPREAD rather than who is #1: the ranking board already said that, and a units order IS the chart order, so without its own angle this is the same post twice. The fact here is the DISTANCE between these numbers
That clause was written in #2098 to stop four projection boards opening with the same
leader. It over-reached: “rather than who is #1” reads, to a literal model, as “do not
name who is #1”, and the take obeyed. The reword keeps the angle and restores the
subject (_PROJ_STEERS["proj_top_sales"]):
after:
Then NAME THE ALBUM ON TOP and its units figure: the leader is the subject of this take, and a gap with nobody named in it is not a post. YOUR ANGLE IS THE DISTANCE -- how far clear of #2 that album is, or how tightly the rest of the board is packed. The ranking board owns positions and moves; you own the leader's units and the gaps.
The headline now opens with the leader too ("Dandelion" by Ella Langley is tracking for
~69K units, the most on the projected Billboard 200 for Sep 19, per HITS, ~4K clear of
the #2 album. The top 10 runs from ~69K down to ~34.8K.). The class sweep found the
same clause on the platform play-count steer (_PLATFORM_STEERS["platform_streams"])
and reworded it the same way.
Live dry run on the same HITS document (scripts/dryrun_chart_boards.py --compose):
Hits Daily Double has Ella Langley’s “Dandelion” leading the projected Billboard 200 for Sep 19 at ~69K units, only about 4K clear of Olivia Rodrigo’s ~65K, tightest race atop the chart in a minute. – gate 0.88, SHIP
The rank board’s take from the same run named Dandelion too, and the two passed the
wording-only dedup the chart cards ride (SequenceMatcher 0.43 against a 0.6 threshold;
0.49 against the rank take that actually shipped on 2026-09-08). The lesson is the
“ban the thing you actually mean” one again: the thing meant was “do not build the take
around the leader’s rank”, and the proxy “rather than who is #1” banned the subject.
Owner steer, on three shipped @tootsiesbar posts: “cut second sentence catalogue talk”.
The takes:
Billie Eilish cleared 8 billion streams on the year, a catalog run that refuses to let go. Kanye West cleared 8.3 billion streams on the year, a massive catalog-and-current run. Tyler, The Creator cleared 3.2 billion streams on the year, another massive catalog-and-current run.
The lane is the music desk’s ANNUAL year-to-date milestone (cogs/music_desk.py,
_CONTEXT MID-YEAR MILESTONE + the streams-milestone note). Four of four annual takes
that night ended in a clause that restates the figure in words. The self-gate scored
them 0.90 and called the filler “one clause of real context”, so nothing downstream was
going to catch it.
The cause was a mandate, not a missing ban. The paragraph read '<artist> cleared <X>
streams on the year' + ONE short trailing clause, then STOP, and it named three sources
for that clause: the momentum trend, the year-end projection, the rank. All three are
withheld on this lane – momentum is gated to the non-annual streams facts, the
projection is banned by the streams-milestone note in the same context, and a rank only
arrives when a chart or qualifier block lands. So on most slots the model was told to
write a clause and handed no fact to write it from. It wrote one anyway.
The paragraph already carried the counter-rule. “never a closing flourish” sat two clauses after the mandate and lost to it – the canonical shape this file warns about: a ban stacked beside the instruction that causes the behaviour does not cancel it.
The fix is one reword: the clause is CONDITIONAL on a fact. The format line now stops
at the figure, and a trailing clause is allowed only to carry a hard fact a context block
states. Given none, the post ends at the figure. The streams-milestone note appended for a
measured report lost its own half of the sanction (“one sharp human beat about how big it
is”), but only on the annual lane (streams_milestone_note(annual=...)). A NON-annual
measured report heads its blob with a grounded release stage – “debut week”, “first month,
3 weeks in”, “catalog, out 4.2 years” – and the shared AGE / STAGE paragraph asks the take
to put that milestone in its own words. The stage sits in the BLOB, not in a context block,
so the annual wording would have stripped the stage framing off every weekly report whose
fail-open supplemental blocks came back empty. The first version of this PR missed that
and the Codex reviewer caught it. The lesson generalizes past this fix: a rule keyed to “a
fact from a context block” silently excludes the reading’s own head, and a shared block is
appended on more lanes than the one you are looking at – so check which lanes it reaches
before you narrow what it allows. The annual head carries “year-to-date” and no age at all,
so it loses nothing. Dry-run both halves: the non-annual takes still read “four years deep
and barely slowing down” / “six weeks in and still holding a real number”.
Measured on the real wired path (compose_market_drop with the desk’s own _CONTEXT and
the real blob; Sonnet, the lane’s default; the three reported artists, n=5 each, run with
and without a Billboard-standing block):
| arm | takes | ending in an invented clause |
|---|---|---|
| before | 18 | 18 |
| after | 30 | 0 |
The chart fact still rides when the block names that artist (“cleared 8.3 billion streams on the year, and ‘Graduation’ is still charting at No. 148 on the Billboard 200 this week”), and the two artists the block did NOT name correctly ignored it.
Two guards were measured, not assumed, because a shorter take could have gone dark two ways. The self-gate scores the bare figure 0.90-0.92 against the real blob (the same score the flourish got), so nothing new is dropped. The arbitrated dedup still ships four same-shape annual takes in a row: the mechanical gate matches the shape in both arms, and the judge overturns it on the artist and the number. A first dedup run said the opposite and was wrong – the harness called the judge with the wrong signature, so every judge call raised and the fail-closed direction left the mechanical verdict standing. Check that your harness’s judge actually answers before you read its verdict.
The second Codex finding on #3252, and the more general of the two. The clause rule was
written as “a HARD FACT a context block below states”. That phrasing assumes the
surface HAS context blocks. cogs/music_desk.py:_CONTEXT is shared with
cogs/music_alert.py, whose _fire passes it with no supplemental blocks at all –
so on the alert path the rule named a source that does not exist, and the only material
the take had (the market blob, carrying the line’s own move) fell outside what the rule
allowed. Fixed by keying the rule to the MATERIAL rather than to a block that may not be
there.
The predicted symptom did not reproduce, and saying so matters more than the fix. The
review expected the alert to discard the move and post the already-known year-to-date
count. Measured on the real classifier + the real compose (an annual line_move, Sonnet,
n=8 per arm), the figure rode in 8/8 of every arm, the flagged wording included – the
model used the blob’s pace line regardless of what the rule said it could use. The fix is
right because the wording was wrong, not because it rescued a post.
What the measurement DID find is a weakness that predates the PR. An annual
line_move alert fires because the market-implied number moved, and the take almost
never says so:
| arm | names the MOVE | names the figure |
|---|---|---|
main (pre-PR control) |
0/8 | 8/8 |
| the flagged wording | 0/8 | 8/8 |
| the fix | 2/8 | 8/8 |
main scoring 0/8 is the whole point of running a control: without it this reads as a
regression the PR caused, and it is not one. The cause is upstream of any clause rule –
the MID-YEAR FORMAT makes the year-to-date figure the mandatory lead and demotes the move
to a trailing clause, on a lane whose entire reason for firing is the move. That is a
separate change on a separate lane, so it was FILED, not folded into this PR.
Owner steer in the same session: “cut scars”. Two of the four kinds this file names were
in cogs/music_desk.py:_CONTEXT, and both are the FIRST kind – a rule with two
definitions.
claude_client._POST_LENGTH (“ONE
sharp sentence, ~20 words, then stop … Pick the SINGLE sharpest beat and let it
carry”), plus the market-drop tail (“never the figure AND a projection AND a trend AND a
flourish; if the figure alone lands, stop AT the figure”). _CONTEXT then restated both
– “HARD LENGTH LIMIT (the #1 rule): ONE sentence, ~25 words” and “Stacking the
figure AND a projection AND a trend AND a flourish is a banned paragraph”. The second
clause is near-verbatim the system tail; the word counts are simply different numbers for
the same rule. Cut to the one thing this block alone knows: the several context blocks
are OPTIONS for the single clause, not a checklist.format_reading has put humanize_units(value_so_far)
– “8B” – in the blob since the owner’s “no one needs the exact count” steer. A fence for
material that no longer exists is the “scope instruction outlives its scope” case again.Ablated on the real wired path before cutting (compose_market_drop, the desk’s own
_CONTEXT + the real blob, Sonnet, n=5) across three cases, so a cut that fixes one axis
could not quietly break another:
| case | axis | baseline | scar cut |
|---|---|---|---|
| annual, no chart block | ends at the figure | 5/5 | 5/5 |
| annual, chart block names the artist | quotes the chart fact | 5/5 | 5/5 |
| weekly catalog album | states the release stage | 5/5 | 5/5 |
Mean length moved 9.0 → 9.0, 20.8 → 22.4 and 22.6 → 19.8 words: noise in both directions at n=5, no trend. So this cut is justified on prose hygiene alone, not on a measured win – the same honesty the #855 entry insists on. 287 characters out of a 3917-character block, and one contradiction the model could have resolved either way is gone.
Two candidates were examined and NOT cut, because each still earns its place: the AGE / STAGE paragraph’s closing “a chart row carries no release date, so its age is not in the material” (it reads like engineering detail, but it is the reason the rule holds, and that lane’s whole failure mode was the model supplying an age the row never stated), and the streams-milestone note’s “no exact digit-string” (it fences a FABRICATION the model can produce unprompted, not an echo of the material, so the stale-fence argument does not apply).
Owner steer on weekly streams posts that ended “a steady start”: “any commentary should come from the stats, like the qualifiers. If it’s just an empty human beat, remove it.”
The #3252 fix above made the clause CONDITIONAL on a fact, but only on the annual lane. The weekly lane kept three sources of an empty beat, and all three were cut, not countered:
streams_milestone_note said “Then one sharp human beat about how big it is.” It now
has one rule for every measured report: a clause carries a hard fact the material
states, else the figure is the whole post.lookup_artist_momentum are gone.include_stage=False. Market reads keep the
stage, because the paragraph is shared with them.A rank from the background block and a verified qualifier still ride as the one clause, because they are stats.
Owner steer, on an X post that ended “HITS Daily Double has it there.”: “drop them from the text post and make them on the card. The pill is enough.”
The cause was a mandate. attribution.credit_rule said “CREDIT HITS, AND DO IT IN
THE SENTENCE”, and gave in-voice examples that named the source (“hits has her tracking
~74k”). A ban on the press-release tag sat beside it. The model obeyed the mandate and
moved the name around the sentence. The rule now says only that the card names the data
source, so the line names none.
A name in the instruction is a name in the take. The first rewrite said “THE CARD
NAMES HITS DAILY DOUBLE AS THE SOURCE”. On the reconciliation lane, Sonnet still wrote
“HITS Daily Double projected #3” on 3 of 8 takes. Three more places carried the name to
the model: called_shot.describe (the take’s topic line), the block’s per-publisher
labels, and the rule itself. Each was reworded by ROLE (“the industry projected #3”),
and the rule stopped naming any source. The next run was 0 of 10.
Measured on the real compose inputs of six lanes (compose_market_drop, the desk’s own
_CONTEXT + each lane’s framing):
| arm | model | takes naming a source |
|---|---|---|
before (main) |
terra | 23/30 |
| after | terra | 0/48 |
| after | Sonnet | 1/48 |
The one Sonnet miss was on the reconciliation lane. The widened has_attribution_tag
drops it after one _detag_credit re-ask, which now runs on every desk lane.
Platform boards are the exception, and it is deliberate. On a Spotify or Apple Music board
the platform IS the chart (“#1 on spotify’s us chart”), so naming it is the fact, not a
credit. Those boards keep their own chart phrasing (platform_chart_phrasing).