a discord bot for the tootsies server. ask, recap, discuss, ship features by typing.
What every external API hands us vs. what we surface to the model. This is the “are we dropping anything” ledger: each integration lists the fields we extract and the fields we deliberately omit (with the reason). The rule is either 100% of the answer-valuable fields, or an explicit entry here saying what’s missing and why — so an omission is a documented decision, never a silent gap.
Scope note: “surface to the model” means the field reaches Claude’s context (a rendered block or a tool result), not telemetry. The constitution’s data-minimization rule governs telemetry + her outbound replies, NOT the in-context input we feed the model, so enriching context is in-spec. A handful of fields are withheld from output by the HARD RULES (no personal-info disclosure) even when fetched — those are noted.
Every surfaced field is also self-describing in its rendered block or tool description (a header legend or inline label), so the model knows what each value means and what to do with it — not just that it exists.
utils/industry_feeds.py, #reporting-oversight)The trade press read directly via public RSS (no key), so a first-party number is in hand
rather than chased through a lagging web-confirm. Five sources in FEEDS: Pollstar news,
IQ Magazine, Music Business Worldwide, Billboard music-news, and Billboard Chart Beat
(added 2026-09-28: Billboard’s own chart announcements, which the music-news feed does not
carry). One namespace-driven parser. The cache holds a feed 90 seconds, because the
newsroom’s feed watcher reads every 2 minutes (see docs/ARCHITECTURE.md).
title, link, summary (the teaser, tags stripped + entities decoded),
published (RFC-2822 → ISO, so the desks’ own time parser reads it), the lead
image_url (from content:encoded / enclosure), categories, author. That is
100% of what a feed item carries for reporting purposes.content:encoded article body — the teaser plus the
headline carry the number and the who/what, and reposting a full article body verbatim
is the redistribution an RSS feed does not invite (a fact + attribution + link does). A
desk that needs the full body fetches the article link on demand.relay_number=True — a first-party trade number may be
stated and attributed to the source (“per Pollstar”), not routed through the aggregator
verify path (which exists to check a reposter’s claim against the first-party source —
and this IS that source).utils/riaa.py, #2839)RIAA’s own certification database, read directly (no key). Four views: the recent feed (every award newest first, 30 rows a page, paged by the site’s own “show more” JSON call), one award’s timeline (its full level history with dates), the top-tallies list (the 100 highest-certified albums, certified units in millions, 4 pages) and the awards-by-artist search (one row per credited act: certified units, Gold / Platinum / Multi-Platinum / Diamond counts, SOLO titles only per the site’s own note). The site holds no singles tally list (probed 2026-09-02: every sub-tab value answers the album list).
award_id, artist, title, label, level (certified units
in millions). Parsed (artist row): artist, units_millions, gold, platinum,
multi_platinum, diamond. 100% of both rows’ cells (the “type” cell is “All
types” on the solo row and is not surfaced).awards_for): every award whose artist OR title carries
the act’s name, the feed row’s cells plus the badge’s program code (ST/DI main,
LA Latin, MT the retired ringtone program). This is the only view that sees a
FEATURE credit (“LIFE IS GOOD (FEAT. DRAKE)” sits on Future’s row).contributors cast: 264 main-program rows, about one in twenty hides a
credited act (“23” hides Miley Cyrus, Wiz Khalifa and Juicy J; “UMBRELLA” hides JAY-Z;
“CLOSER” hides Halsey). riaa._OMITTED_CREDITS names those rows by hand and
scripts/riaa_credit_audit.py finds and re-checks them; the audit is a MANUAL tool, so
Deezer is not a runtime dependency of this feed and no card’s number rests on it. The
sweep now WRITES its accepted rows to utils/data/riaa_omitted_credits.json (--write,
#2944), which _OMITTED_CREDITS merges under the curated entries – so Deezer is a
BUILD-TIME source whose output is committed as a reviewable diff, never a run-time
dependency. The
sweep keeps SINGLE rows only: an album row’s credit is the album artist, so a track
search on an album title matches an unrelated track. Measured on the Diamond ALBUMS
population, 150 rows gave 22 candidates and all 22 were false (Drake’s album “Take Care”
read as the track “Take Care (feat. Rihanna)”, Adele’s “21” as “21 grammes”).artist_rankings, #2918): the SAME awards-by-artist
tab with an EMPTY name – RIAA’s own ranking of every act it has certified, most
certified units first, 30 rows a page (show-more action load_more_artist_default).
Measured live 2026-09-03: Drake 333.0M, Kanye West 271.0M, Taylor Swift 247.5M, and the
reason it exists – The Beatles 207.5M, Garth Brooks 200.5M and Elvis Presley 189.5M,
none of whom appear on any streaming list. This is the CANDIDATE population for the
certified-units and Platinum-singles boards, which cannot read a ready-made list the
way the Diamond and year boards do. The rows count SOLO titles, so they choose WHO to
read; each act’s feature-inclusive total still comes from the per-act award search.
Four pages (120 acts, down to 62.5M solo) leaves the headroom measured in
ARTIST_RANK_MAX_PAGES.utils/riaa_boards.py): the fastest roads to Diamond among today’s top acts’ singles (first certification date to Diamond date off the title timeline), the highest-certified singles and albums per genre (Pop, R&B/hip-hop, Country, Rock; the genre-filtered award search, Diamond and up, #2860), the top ten albums with
their level, and the top acts’ Diamond counts, singles and albums as two boards, with
features and collaborations included (diamond_tally per format: every main-program
Diamond of that format the act is credited on, attributed
through artist_watch.credit_matches, the chart boards’ one segment-boundary rule,
on the artist cell and the title’s feature clause; the act’s most-certified title
beside it; the Latin program counted apart and never ranked; the desk’s watch list,
every act walked before a “most” is claimed). Measured 2026-09-02: Drake 16 singles
and 1 album. The take is fenced to credited titles of the one format and to today’s
top acts.award_id (the row’s id, our seen-key), artist, title,
label, format (SINGLE | ALBUM), level (the badge code: 0 Gold, 1 Platinum, N = N×
Platinum, 10+ Diamond) and certified (the certification date). 100% of the row’s
cells. The timeline gives every (date, level) pair; the detail block’s genre and
release date are parsed out of the same payload but not yet surfaced.relay_number=True – this IS the certifying authority, so the figure is
stated on source trust and attributed (“per RIAA”), never routed back through the
Grok/Perplexity verify hop. Scope stays US + UK certs only (RIAA is US).utils/pollstar_charts.py, #reporting-oversight)Pollstar’s public chart JSON (data.pollstar.com, no key), the live-music numbers
Billboard does not publish. Anon free-tier only — no auth is ever sent, and only the
preview rows the endpoint returns are taken (top ~10 concert / ~50 Mediabase). One parser
for every chart in CHARTS.
CHARTS[*].chart_id
is only a SEED, and the client resolves the current issue before every read: each chart’s
config endpoint (/charts/{seed}/{type}/config) lists every issue ever published under
variants[].charts[], and latest_issue_id takes the newest by DATE. Cached 6h, keyed by
issue so a new week busts the body cache at once; an unreadable config falls back to the
seed, leaving the board’s freshness fence to refuse it rather than post a stale week as
current. Reading the seed directly is what froze the touring + radio lanes for 15 days
while integration health read 100%.rank, subject, week_of, the chart page url, image, and
the full values map (raw) + display map (formatted) of every column the endpoint
returns — Live 75 (avg tickets, gross, shows, agency), Global Concert Pulse (avg gross,
avg ticket, avg price, shows), New Tours (avgticketssold, avggross, genre,
datefrom/dateto tour run, tourduration, eventcount), Artist Power Index
(live/streaming/airplay/social composite), Mediabase (this/last week, song, label, total
spins). 100% of the free-tier columns land in ChartRow.utils/pollstar_boards.py): the ranked board CARD draws one
value per row — rank, the subject, and the headline metric (chart.metric) — plus, on a
radio chart, the SONG as the row label with the artist beside it (the chart ranks songs
by spins). The model’s TAKE (compose_inputs) is handed EVERY parsed column, each named
in-line — song + label on radio; avg tickets, price, shows on the tour charts — plus a
deterministic last-week rank move, so no column is dropped from what the bot can say. The
card stays single-value by design (a ranked standing is one number); the extra columns
reach the reader through the take, not a second card column.rank
column and the anonymous preview returns rows in ALPHABETICAL order, capped at ~10, so
the board can only ever show the FIRST slice of the alphabet — one week that was ten acts
all named A or B, with every tour by a later-alphabet artist absent. An earlier fix
(#2433, owner report 2026-08-19) made it a dated ROSTER (PollstarChart.ranked=False) to
kill a false “#1 biggest” claim — Adelitas Way (alphabetically first) had read as topping
the list at $6,113 while Backstreet Boys’ $2.8M sat at #7. But the roster fix removed only
the ranking claim, not the sampling bias: “New tours going on sale, soonest first” still
read as THE soonest new tours when it was only the soonest among the A–B sample. An
alphabetical fragment cannot honestly represent “new tours going on sale”, so the board is
no longer posted (prefer absent over a misleading partial). new_tours is dropped from
CONCERT_BOARD_KEYS; the chart stays parseable in CHARTS, and build_board still
renders the roster if called directly, so the fetch + roster code + tests are kept to
re-enable cleanly if a licensed full-chart read replaces the alphabetical preview. The
other three concert charts keep their true ranking and stay boarded.rank = 1..10 with the headline metric strictly
descending, and the five boarded Mediabase formats (Top 40, Country, Urban, Rhythmic,
Hot AC) return thisweek = 1..10 in rank order with spins descending across the top 10.
So each preview IS the true top 10, not an alphabetical slice. The difference is
structural: those charts carry a real rank column, New Tours (an on-sale list) does not,
so only New Tours falls back to alphabetical. (Mediabase Country’s spins dip below
monotonic past rank ~12 because Mediabase ranks on an audience formula, not raw spins;
the boarded top 10 stays clean and rank-ordered.)PollstarChart.attribution (“Pollstar
utils/billboard.py)The first-party chart read (server-rendered HTML, no key). One parser covers every
chart in CHARTS; past weeks are addressable by date.
_WEEKLY_EXTRA_CHARTS and no standings treatment, the same
terms as Artist 100 / Top Album Sales / Digital Song Sales. Slugs, row counts and labels
were read off each page’s own H1, not guessed off the slug.catalog-albums, current-albums,
hot-dance-electronic-songs, hot-christian-songs, hot-gospel-songs and
rhythmic-songs return rows but parse no week date and pick up “Also appears on”
artifacts in the title — a different page markup the shared parser does not cover, so
they need a parser check before any of them ships. tiktok-billboard-top-50,
vinyl-albums, soundtracks, world-albums, dance-electronic-albums and the airplay
family (pop-songs, adult-pop-songs, alternative-songs, adult-contemporary,
r-b-hip-hop-airplay) returned nothing at all, so their slugs are wrong and need
discovery first.rank, title, artist, last_week, peak, weeks_on, plus the
derived is_debut / is_reentry / is_new_peak / movement() and the chart’s
Week of date. That is 100% of what a Billboard chart row publishes — the page
carries no other per-row datum.movements) rather than read, because Billboard doesn’t publish them as
fields — they fall out of rank vs last_week vs peak. Week-over-week entries and
exits need two weeks (diff_weeks), since a song that fell OFF the chart is by
definition absent from the current page.charts-static.billboard.com thumbnails
— the entity-image path already resolves art, and a chart thumbnail is a worse source
than the catalog one); the per-row “chart detail” expander (an AJAX sub-request per
row — 100 extra fetches for award/imprint trivia no surface asks for, and a volume
we’re deliberately not spending against a host whose robots.txt names us).source_blocked), never a stand-in.artist_entries, #2154 + #2587): the per-artist
/artist/<slug>/chart-history/<code>/ page, server-rendered like the weekly charts.
Surfaced: each entry’s TITLE, DEBUT DATE, and all-time PEAK position ((title, date,
peak) triples). Two career facts are built from them, strongest first:
music_news.entry_ordinal_fact, #2154) ranks the reported
title by its debut date (“JUNGLE’s fourth Billboard 200 entry”), so the page’s
peak-ordered rows don’t mislead, two same-titled releases resolve to the newest, and
an album’s same-week co-debuts (one shared date) are suppressed;music_news.peak_history_fact, #2587) uses the peak column +
the CURRENT week’s position for “their first #1” (position 1, no prior entry peaked #1)
or “their highest-charting entry yet” (position beats every prior entry’s peak). Keyed
on the current position, so the claim is fresh news, never an all-time peak reached
months ago; a missing peak on any row, or a same-titled re-release, suppresses it.
Scoped to the two flagship charts (billboard200→tlp, hot100→hsi), the ones the
newsroom’s charts_named resolves. Omitted (intentional): weeks-on-chart and the
peak DATE — the page publishes them, but no surface consumes them yet. Fail-open + QUIET:
a blocked / unknown-slug / not-yet-listed read yields no fact
(billboard_artist_history reason=…), never a wrong one, and it rides its OWN circuit
breaker so a history outage cannot short-circuit the authoritative weekly-chart read.utils/chart_history.py, cogs/chart_history.py, 2026-09-20)Not a new upstream: the DURABLE copy of three integrations below plus Nora Music’s X posts, written to Postgres so the history is read from the store, never re-fetched.
CHART_HISTORY_CHARTS (default hot100 + billboard200); plus, for the
Billboard 200 top 10, the PRINTED UNITS off Billboard’s weekly chart story
(utils/billboard_stories.py, 2026-09-20): total equivalent units, the album-sales /
SEA / TEA split where the story breaks it down (the No. 1 and a debut), the
week-over-week change, an approx flag for a qualified figure, and the story URL.
Omitted: the row’s history_slug (a lookup key the artist-entries read resolves
live), the page’s image/link fields (never parsed), and everything else the story
says (stream counts, vinyl/CD breakdowns, career context: prose the desk reads live).source_ref. The structured chart_entries exist only on documents
since 2026-04; every document since 2025-01 carries the chart_data TSV twin, and
parse_projection falls back to it (parse_chart_data), so the walk reads the whole
dataset (measured 2026-09-21: 83 midweek + 89 building documents, of which only 17 +
19 had entries). Omitted: label marketshare (a per-document aggregate, not a
release fact; the live read still carries it) and the building chart’s FINAL /
building status line.KXALBUMEQUIV, KXPUREALBUMS,
KXALBUMSTREAMS, KXSONGSTREAMS, KXALBUMDEBUT plus every Luminate family the desk’s
discovery lists, KXALBUMEQUIVY and KXALBUMSTREAMSU included): the parsed release + artist, the
metric, the tracking window and its chart week, the latest implied figure and ladder
depth, the settled expiration_value (Luminate’s count) and the binary result, plus
a daily snapshot of the implied figure while live. Omitted: the per-rung prices (the
snapshot keeps the implied median only; the desk’s candle read stays live).chart_history tool + the X mention reply’s history lines, #3562 E):
every stored column above that answers a question reaches the model in its rendered
block (chart_history.render_history_lines): the weeks with rank / last week / peak /
weeks on / debut / re-entry and the printed units with their split, change and
approx flag; the projections with rank / units / split / change / last week per
source and kind; the markets with status, window, implied figure + rung count + its
time, settled count + its time, and the binary result; the daily rows with rank /
streams / peak / days on / day move / NEW; the track days with US and Global streams.
Omitted from the block: the join keys, the Sanity/tweet source_ref and story URL,
the market’s series ticker, the daily rows’ Spotify ids and release date (lookup
keys, not answers), and the per-day market snapshots (the desk’s own path read; a
later forecast surface reads them). One line in the block is DERIVED, not stored:
the next-week OUTLOOK (chart_priors, #3562 B) is this store’s own arithmetic on its
own rows, labelled as such in the section head, and it carries the cell and the range
it came from so the model can say where the number is from.source_ref, for their
own posts (early_look / release_radar / post) and third-party citations
(cited). Everything else on the post (engagement, media, the app links) is not a
projection and is not stored. Their numbers are theirs: stored for comparison, never
re-posted.utils/hits.py)The industry first-week projections + Mediabase radio feeds, read off HITS’ public
Sanity GROQ API (no key; access reasoning in docs/BILLBOARD_SOURCES.md §5b).
units (Activity),
change_pct, the full consumption split — album_sales (pure), tea, sea —
last_week (None = a projected DEBUT), label marketshare, the tracking-week end
and the derived chart_date (the Billboard week the projection is FOR). Columns are
read by label, so the midweek chart’s subset (Activity + Albums only) yields
None for the absent fields rather than shifted values.genre — load-bearing: the
docs concatenate ten format charts each ranked 1–50, so rank is only meaningful
per-format via by_genre), adds, plays_this_week/plays_last_week + the
derived spin_gain, rank, last week.albumReleases calendar — release date,
artist, album, label, and the row’s BENCHMARK columns: the artist’s PREVIOUS album’s
release date and that album’s first-week units. This is the only forward-looking
document any chart source here publishes, and the only one carrying a benchmark. It is
read from the chartData TSV, not the structured twin: chart_entries carries
only {album, artist, label, release_date} and drops both benchmark columns. The column
meaning was validated against a known outcome (Carly Rae Jepsen’s “Day and Night” reads
10/27/22 + 20,800, matching “The Loneliest Time”’s real first week) rather than inferred
from the format, since the document has no header row. -- in both trailing columns
means “no prior album”, which stays None — a 0 would read as “opened at nothing”.
Consumer: the release_calendar board on the projection lane (#3013). It ranks by
the benchmark rather than by date and drops rows without one — a dry run showed the
near rows are mostly acts with no prior album, so a chronological cut left nine of ten
value cells blank and named a 33.9K act the biggest while a 233.3K one sat past the cut.hits_benchmark,
180d) as the calendar is read, keyed on the normalized artist+album, and read back
weeks later when that album appears in a projection (benchmark_for). Fail-open at
both ends — a store miss costs one absent clause, never the read or the post.activity (year-to-date
album-equivalent units), album_sales (pure), song_sales, audio_streams and the
release date. TSV-only again — this document ships no chart_entries at all.
Consumer: the ytd_albums board on the projection lane (#3013).chart_type declares five (Activity, Albums, Songs, Audio, Released), so exactly one is
undeclared. Four are anchored — Activity by position and scale, Albums by checking acts
whose physical sales run opposite ways (BTS 989,698 against Drake 16,968), Songs, and
Audio as the only billions-scale column — leaving field [7] unidentified. It is not
read. A mislabeled number is worse than a missing one here because nothing downstream
could tell it was wrong; a test asserts no parsed field ever equals it.chart_data TSV twin of every PROJECTION doc (there the
same data as chart_entries, unlabeled — we read the structured shape; the two TSV
reads above exist because those documents have no usable structured twin); the
year-end chart types (annual, one doc each); the six DEAD chart types — song-revenue,
song-streams, ytd-song-streams, ytd-latin-song-streams, ytd-country-albums and
the noisemakers / rainmaker / track document families, all last updated between
2024 and Oct 2025 (surveyed 2026-09-06, #3013 — do not rebuild against them); editorial
post documents (we take their data, not their journalism).format_projection_block, not left to a prompt.utils/chart_boards.py: the four daily projection boards read the full
row set (rank, units, pure album_sales, last_week), and the weekly debut-sales
board joins the building chart’s units to the printed Billboard 200 debuts under
the chart_date == week_of alignment fence. All HITS-sourced cards carry the HITS
source pill; projected boards additionally date themselves “PROJECTED FOR …”.
proj_pure_sales (#2098) reads the same rows but ranks on album_sales alone, and
drops any row the document gives no pure-sales figure for rather than substituting
units — the two are different measures and a card calling streaming activity a sale
would be the one way that board could mislead.utils/kworb.py, #youtube-chart-art)The “Today’s Most Viewed Music Videos on YouTube” pages — the global top-500 and the anglophone top-300, read as one merged list. Their read is CARD ART, not model context: nothing on these rows reaches Claude, so the ledger question here is only whether we drop something the art path could use. (The other kworb pages — Spotify, Apple, iTunes, radio — have no entry yet; that is a documented gap, not a decision.)
maxresdefault (1280×720) then
sddefault (640×480), never hqdefault (480×360). Every video id serves hqdefault,
so taking it would turn a genuine miss into a soft, cover-cropped card.chart_boards.video_views_board,
the worldwide most-viewed card, from 2026-08-10 until the owner cut it (“lets do the us
only and cut the international one”). The pages are still read, and the MARKET-ART path
is now their only consumer, so no row here reaches a model or a card — only a thumbnail
does. The US read that replaced the board is the weekly /youtube/insights/us.html
chart below.utils/kworb.py, #3191)Every kworb page serves ETag and Last-Modified, and until 2026-09-12 the client read
neither. They are not chart content, so they reach no model and no card; they are read for
two jobs.
ETag + Last-Modified as the conditional-request
validators, so an unchanged page answers 304 with no body; and Last-Modified again as
the page’s own refresh time (refreshed_at), which is the only timestamp we hold that
says when a CHART changed rather than when we looked.Content-Length, Server, Accept-Ranges. None of them says
anything about the chart, and a byte count is not a change signal — kworb can republish
the same length with different rows.refreshed_at rides the chart_crown event for
diagnosis only. A page’s refresh minute is when kworb PUBLISHED a change, not the minute
a title reached #1, so printing it as the latter would be a fabrication.utils/kworb.py, 2026-09-19)The per-track chart-history page (spotify/track/<id>.html), read by track_streams.
Since 2026-09 it renders TWO sections in the same table shape: <div class="weekly">
(one row per chart week) and then <div class="daily"> (one row per day); each has a
Total and a Peak row and a per-country column set.
Total row per country (total, the
Global column; countries, the rest) and the last dated row’s Global streams
(daily). The Total is the streams kworb COUNTED while the song sat on each chart,
which equals a lifetime total only for a song charting daily since release; for a
catalog song it is a fraction (“Shape of You”: 1.33B here, 5.10B lifetime). So no
ladder reads it: the per-track streams ladder takes the act’s songs page
(track_lifetime, joined by track id, else title). The streams-hero backup
(release_streams) takes the LARGER of this Total and the songs-page row, since
each is a lower bound (the songs page lags a brand-new song: “Joseph” 279K there
against 7.59M here, 4 days out), and this page’s per-country split.Peak row (position + streams at peak,
which the daily chart rows already carry).TrackDay (TrackStreams.days, oldest first): the Global and US
columns, found by header name, as (pos, streams), -- as None. kworb keeps about
30 days of it (33 rows on a song 338 days on the chart), so a song’s first month is
recoverable and then gone; the chart-history store’s daily walk reads it once per
track into chart_history_track_days. Still omitted: the other country columns
(CA, AU, GB, …), which no consumer reads.utils/kworb.py, research round 4, 2026-09-21)The Spotify daily chart pages (spotify/country/us_daily.html, global_daily.html) name
the chart DATE in their title (Spotify Daily Chart - United States - 2026/09/19): the
day the streams are FOR. kworb publishes a day one or two days after it (read 2026-09-21,
the US page showed 2026/09/19), so the read’s date is not the streams’ date.
ChartRow.chart_date, parsed once per page by parse_chart_page_date
and stamped on every row; None when the page names none. The chart-history store keys
chart_history_daily and chart_history_track_days on it and stores nothing from a
page that names no date (absent over invented). No lane prints it yet.utils/kworb.py, owner ask 2026-08-10)The all-artists ranking (spotify/artists.html) and the monthly-listeners ranking
(spotify/listeners.html), consumed by /ask’s spotify_chart standing line, the
watch sweep, and now the artist leaderboard boards (utils/chart_boards.py).
Daily +/- day-over-day change (the
listener-movers board’s whole basis — parsed since 2026-08-10), and peak
listeners (listeners page).utils/kworb.py, owner ask 2026-08-31)The per-country YouTube chart for the US (youtube/insights/us.html), behind the
us_video_top + us_video_new boards. A DIFFERENT page and metric from the
most-viewed video charts above: those are worldwide, daily, and by VIEWS; this is US,
weekly, and by STREAMS. It is the only YouTube page we read that says whether a title
is new, and since 2026-09-14 it is the only YouTube page that feeds a BOARD at all.
The page size flips: kworb has served it at 100 rows and at 20, and the boards need
it whole, so kworb._US_VIDEO_CHART_SIZES holds both measured sizes and a read matching
neither darkens both boards with a chart_fetch shape_change.
P+); artist and title (split on
the first ` - , the radio chart's idiom); weeks on chart; peak position; the week's
streams; the signed week-over-week stream change; and the NEW / RE markers as
two SEPARATE booleans (is_new, is_reentry`) — both print no delta, so only the
marker distinguishes a first-ever week from a title returning after 49 weeks.(x?) multiplier column, which is a chart-mechanics
weighting no card explains or uses.utils/kworb.py, owner ask 2026-08-30)The all-time most-viewed YOUTUBE ARTISTS leaderboard (youtube/archive.html — kworb
files it under “archive”, but its own nav calls it Artists). ~1,732 acts. It backs the
market card’s ranking strip on a YouTube market, so the scale is read on the same
platform the card’s number is on. No consumer since 2026-09-29: the artist standing
strip it backed was removed from every card (owner steer). The client read and parser stay,
but nothing calls them.
youtube_artists fetches with a high min_rows floor (500) and a short read
is DISCARDED rather than ranked. _fetch used to return the partial parse below the
floor, which left the floor advisory — fixed with this page (Codex #2702), and a
no-op for every min_rows=1 caller, where under the floor already meant zero rows. The Spotify ranking pages need no such
floor: they print their own #.charts.youtube.com is a script app; probed
2026-08-30, no artist name in the markup). So the strip is an all-time career total and
says so in its own words.utils/kworb.py, album card, owner ask 2026-08-16)The per-artist songs page (spotify/artist/<id>_songs.html, parse_artist_songs):
every tracked song of one artist with its cumulative total + latest daily. It
backs the fresh-drop album-streams card (chart_boards.album_streams_board) —
match an album’s iTunes tracklist to these rows and each track’s daily comes free.
open.spotify.com/track/<id>
link, NOT kworb’s internal /track/<id>.html — a different link shape, its own
matcher); cumulative total; latest daily (None when the page shows no daily
figure — kworb populates that cell only while a song is charting, which for an
album is its release week; a deep cut past that reads None, never a false 0).<3 inside the link text. The four link matchers
(_ARTIST_LINK_RE / _TRACK_LINK_RE / _SPOTIFY_TRACK_LINK_RE /
_SPOTIFY_ALBUM_LINK_RE) share _LINK_TEXT, which reads a < as text unless a letter
or / follows it, and _strip_tags strips only a real tag. Before that, the row failed
the link match and the parser skipped it as a summary row: the page read 74 of 75 songs,
and the “purple” milestone card said 12th song from the album past 100M where the true
count was 13. Every kworb table parser rides the same two primitives, so the chart pages
hold the same fix.utils/kworb.py, owner ask 2026-08-28)The per-artist ALBUMS page (spotify/artist/<id>_albums.html, parse_artist_albums):
every tracked ALBUM of one artist with its cumulative total + latest daily – the same
shape as the songs page above, one level up. It is the PREFERRED source for an album’s
lifetime streams (utils/stream_totals.py), because it is Spotify’s own published
per-release figure: the older tracklist sum only approximates it (a track’s own total
carries the plays it took as a single, so summing over-counts a release that has one),
and iTunes does not return every act’s albums at all, so for some acts the sum could
never resolve (measured 2026-08-28: an iTunes album search for Bad Bunny returns his
singles and features and none of his albums).
open.spotify.com/album/<id> link); cumulative total; latest daily (None when
the page shows no daily figure, the same rule the songs page follows).StreamTotals.album_total, per-artist board rows). The music desk’s market card reads
BOTH the total and the daily (StreamTotals.album_page_streams), so a first-week units
card can state where the release stands on the day it posts. Before that, the desk’s
only stream figure for a release was its biggest TRACK’s, and the card led on the
opening-day number a chart account had published days earlier (#3613)._album_track_fold drops parenthetical
qualifiers, so a standard edition and its deluxe / international / edits siblings fold
together – four titles collide on the live Miley Cyrus page (2026-09-22), among them
“Something Beautiful” with its (Deluxe) and (Edits). album_page_total and
page_daily_streams each maximize independently, so calling both can pair one release’s
lifetime total with another’s daily figure and present it as one album’s own numbers.
chart_boards.album_page_row returns the row the published total belongs to, and a
caller that wants both figures reads them off it (Codex on #3615).(Deluxe) row’s 359,060,373 under the standard’s name. Live Kalshi subjects do carry
these qualifiers – Black Boy (Alternative) -- Dahi was an open album-equivalent market
on 2026-09-22. album_page_row therefore tries _norm first (it keeps the qualifier’s
words) and only then _album_track_fold, which still covers the two sources spelling one
release differently. album_page_total delegates to it, so the boards and the card can
never disagree about which row a title is.utils/kworb.py, #3013)Two feeds kworb publishes that the client did not read, added 2026-09-06 after a survey of every path the site exposes.
/spotify/country/us_weekly.html). We already read the US daily
and the global weekly, so the missing combination was the US market on the weekly clock
— the closest streaming read to the Luminate tracking week our units come from. Same
table shape, so parse_daily_chart covers it unchanged. Surfaced: the full
ChartRow (pos, artist, title, streams, weeks-on, peak, track id), as a fourth
ChartStanding leg. The two weeklies are named apart in the tag and blob (“#1 US wk” /
“#3 wk”): they rank the same song over the same days in different markets, so a bare
“this week” number is the shape that lets a US rank get published as a global one./charts/deezer/us.html, /charts/deezer/ww.html) on the same table shape as the
Apple pages, so it is technically free to add. It was added on 2026-09-06 and removed
the same day: the standing owner steer of 2026-09-05 is “never use Deezer”, which
the new-releases board section below records, and that steer is about the SOURCE, not
about one board. Deezer stays where it already is — the catalog facade, artist photos,
Versuz credits and the alert lanes’ date gates — and no new Deezer read gets added.
It went dark on 2026-09-14 (#deezer-403), and that is now a standing risk to plan
around. Deezer began answering HTTP 403 to Railway’s IP, on every endpoint, in about
27ms — an edge reject, not a quota. A descriptive User-Agent does not lift it: the
integration probe already sends one and is refused the same way, so this is an IP-level
block on the datacenter, the mirror image of the iTunes throttle Deezer was adopted to
hedge (#371). Two consequences worth holding: nothing in the repo can fix it from code,
and Deezer was the ONLY source of an artist PHOTO, so the block blanked music cards
until wikidata.artist_portrait was added under it as a second provider. Check it with
GET /debug/integrations?names=deezer before theorising about any missing artist photo
— it answers from Railway’s own IP, which a local probe cannot (Deezer answers a
session’s proxy normally while refusing production)./youtube/trending.html). A different signal from the most-viewed
pages and a different table, so it gets its own parse_trending_chart + TrendingRow
rather than a flag: most-viewed ranks by raw view count (a size chart the same few
videos hold for weeks), trending ranks by BREADTH. Surfaced: pos, video id, title,
countries (how many countries it is trending in — live range 2 to 85) and kworb’s own
per-country highlights summary kept VERBATIM. The fourth cell is a country count where
the most-viewed pages carry views, which is exactly why a shared parser would have
published “85 views” for a video trending in 85 countries. Consumer: the
yt_trending board on the platform lane (#3013); every row’s value is suffixed
“countries” so the figure can never be read as a view count.*_totals.html cumulative variants of each Spotify
chart, /spotify/toplists.html (all-time most-streamed), and the per-country Spotify
dailies/weeklies beyond US/GB/global — all live, none with a consumer yet. Shazam
(/charts/shazam/) is live and deliberately deferred, not dropped: it is a genuine
leading discovery indicator and is the next one worth adding (#3013 ranks it fourth).utils/genius.py)web_pages source links, song art image
URL, stats.pageviews, the song page URL.language tags (English-assumed); non-credited media
artists.utils/chart_credits.py, owner ask 2026-08-10): the
producer/writer boards read producer_artists + writer_artists per Hot 100
row, artist-verified against the chart credit and cached durably 45d
(kv_cache: chart_credits; verified misses 7d). Everything else on the song
object stays deliberately unread there — the reference lookup above is the
full-coverage consumer.utils/versuz_catalog._from_genius, #2885 item 2):
the third source for a registered song’s cast, asked only when Deezer and
iTunes both leave it blank (a posse cut both catalogs credit to the lead
alone). Reads the search hit’s title, primary_artist (the title/artist
matchers gate it) and featured_artists; the /songs/<id> read is the
fallback when the hit carries none. The catalog’s lead and link stay; Genius
adds the features. Producer/writer credits are deliberately unread here.utils/ktt2.py, #ktt2)The music FORUM, read as discourse (search_ktt2) and as a breaking signal
(ktt2_breaking). No key and no API: every section page server-renders its Apollo
GraphQL cache into __NEXT_DATA__, so one GET yields typed JSON entities. Access
reasoning + the three pre-ship checks are in the module docstring.
title, postCount (rendered as the traction signal),
createdAt + the last reply time (which drive both the breaking velocity and the
discourse recency rank), isPinned / isLocked, KTT2’s own resolved artists and
album entities, and the opening post’s prose. Plus the board’s own
trendingArtists list, which is a ranking we do not have to compute.id, slug and every URL. Their
terms bar commercial use and “public display” of their materials, so KTT2 is an input
we synthesize from, never a link we paste — the same rule Reddit follows for the
owner’s “not reddit-nerdy” steer. KttThread has no url/slug/id FIELD, so the value
never reaches a formatter, and post_text strips every URL out of post bodies via the SHARED url_guardrail.strip_urls (the
second door: a release thread’s opening post is mostly artwork and YouTube links).
tests/test_ktt2.py asserts both, against the real fixture.likeCount (we take the board’s
temperature, not who said what — the constitution’s data-minimization line); the
per-thread reply bodies (we read the opening post only, so one GET serves both tools);
the archive (there IS no server-side search — ?search= and /search?q= both return
the unfiltered section, so both reads work off the ~50-thread live front page, which
is the right scope for “what is the board on right now” anyway).ktt2_breaking joins each thread’s artists to the
desk’s WATCHED list (utils/watchlist_source.py, ~136 names off kworb) and marks those
threads [ours], ranked first — so “what are our artists blowing up about” is
answerable. The join is on KTT2’s OWN resolved artist entities, never on words in the
title, and it reuses artist_watch.Watchlist.contains, so the styling gaps that module
already handles come free: measured live, KTT2’s JAŸ-Z and A$AP Rocky both matched
the kworb spelling with no new alias. An empty watchlist (kworb down) marks nothing and
degrades to plain velocity order — never to silence.search_ktt2 and ktt2_breaking /ask tools.utils/apple_music.py)collectionArtistName, duration, release
year, track/disc number, genre, explicitness, copyright, artwork (100px +
derived 600px), preview URL, clean track/collection URLs, track count.wrapperType/kind (internal classification, used
only to filter); artistLinkUrl (iTunes itms:// deep link, not
web-browsable — web URL kept instead); purchase price/currency.cogs/music_desk.py:_genre_debut_groups, 2026-09-05): the album
search’s primaryGenreName is READ as the genre of a TRUE debut on the HITS Top 50 (a row on no
row of the printed Billboard 200), so a first-week album reaches the hip-hop / R&B genre board the
printed Billboard genre chart cannot carry yet. The tag is Apple’s top genre for most albums and a
SUB-genre for some (live: “R&B/Soul” for SZA’s SOS, plain “Rap” for Rod Wave), so
chart_boards.APPLE_GENRE_GROUPS maps both levels. A result counts only when it matches the row on
the album|artist key; a tag outside the kept groups leaves the album absent.utils/versuz_catalog.py, #2874): the fallback credit source. iTunes has
no contributors field; the cast is the artistName billing split into acts (“Future, Metro Boomin &
Kendrick Lamar”) plus the parenthetical credit in trackName (“(feat. Akon, T.I., …)”). An Apple
Music link is read by its track id (apple_music_id + Lookup) before the keyword search.top_songs_rss) — im:releaseDate is READ (#2724). The
feed carries a release date per entry and it is the ORIGINAL release, not a
reissue: the live R&B chart dates “September” to 1978 and “I Wanna Dance with
Somebody” to 1987. _normalize_rss_song parses it into year, so /guess can drop
an off-decade chart row in an era game with NO extra call. This matters because
the “hot now” chart is not all recent music — that same chart runs from the 1960s
to this week. Still omitted from the RSS row: im:price, im:collection,
im:image, category (the chart rows feed the game, which needs title / artist /
preview / year only).sequence_mismatch, #slime-language-3):
the fuzzy title match (substring containment, else a 0.6 difflib ratio) cannot see a
sequel. “Slime Language 3” vs “Slime Language 2” scores 0.938, “Slime Season 2” lands
exactly on the floor, and “Slime Language” is a clean substring — so a release card for
an album iTunes had not yet indexed shipped a sibling’s cover and link. The trailing
installment number (arabic, or roman for “Culture II”) is now compared on its own before
either fuzzy test, and a difference — including one side having no number — rejects the
row. Read only from the TAIL past a format word (“Slime Language 2 - EP”), so a one-token
title (“1999”), a leading number (“7 rings”), and a trailing year (“Woodstock 1999”) carry
no marker and the rule never fires on them. Composite roman numerals are PARSED, not
table-looked-up — a table stopping at XII left “Slime Language XIII” and “… XIV” both
markerless, switching the guard off for exactly the series it exists for, and Chicago
numbers its albums into the thirties. Three tests keep ordinary words out: two letters
minimum (bare “i”/”v”/”x” stay words), a canonical spelling, and a value at most 40 —
which is what rejects “mix”, canonical roman for 1009. Shared with Deezer’s own matcher.utils/deezer.py)contributors (featured on a track; the release’s BILLED
acts on an album detail, read by album_cover’s identity guard), album, album cover art,
duration, preview URL, rank, track id. album_cover reads the album-search
cover (cover_xl..cover) as the SECOND cover source behind Apple: a fresh
release often reaches Deezer before Apple’s search indexes it, so a release card
Apple can’t cover still gets its real art. It is trusted as verified
(deezer_cover, exempt from the crosspost media gate), so BOTH guards are STRICT
(stricter than the fuzzy matchers the release-date lookups use): the title must
be an EXACT normalized-title match (a longer same-artist title like “Love” vs
“Love Songs” is refused), and every requested-artist token must be in the
candidate credit (want <= cand), so a same-title album by another act, a
60%-overlap namesake (“Big Time Rush” vs “Big Time Machine”), or an unverifiable
non-Latin credit is refused. Two reads make the identity guard reach a COLLECTIVE
(#2683): the album search returns ONE top-level credit, and a collective bills its
record under the collective’s name — Deezer credits “Slime Language 3” to “Young Stoner
Life” alone, so every spelling the classifier produces (“YSL & Young Thug”, “Young
Thug”) failed want <= cand and the card shipped with no cover at all. When the
top-level credit misses, the album DETAIL’s contributors[] decides — the release’s own
billed acts, ["Young Stoner Life", "Young Thug"] — where the requested billing is split
into its individual ACTS and EVERY act must match ONE single contributor (its tokens a
subset of that contributor’s, or its whole name that contributor’s initialism — “ysl” is
the initialism of “Young Stoner Life” — and an initialism ALONE is refused unless another
act of the same billing matches a contributor directly, because three letters collide:
“BTS” matched the unrelated contributor “Behind The Scenes”), paid only on that fallback
and capped at
_MAX_CONTRIBUTOR_LOOKUPS detail reads across both search terms. Per-act, per-contributor
matching is what makes the widening safe, and two looser shapes were rejected in review
(#2684): matching any ONE contributor accepted “Drake & Young Thug” against a release
billed to Young Thug alone, and covering the tokens from the UNION of all contributors
lost their boundaries, so “Lil Baby” passed ["Lil Durk", "Baby Keem"]. Both put
unverified art on the timeline, which the media gate exempts. This widens WHICH credit is read, never
whether one is required: it runs after CORE-title equality already matched, so it asks
“is this act billed on THIS record”, and an act absent from the contributors is still
refused. The lead-credit rule (#2216) is untouched — it governs an ARTIST-LEVEL card that
names no release, and album_cover is never called without a title. The search also
runs TWO terms (title artist, then the bare title): appending the artist can rank
the real album out of the results entirely — “Slime Language 3 Young Stoner Life”
returned only “Slime Language 2”, while the bare title returned the right album first —
the same trap catalog_art.resolve_music_row already handles on the iTunes side. Both
terms run the same guards, so a second term can only turn a miss into a correct hit. The looser album_title_matches (60% token
overlap) that the release-DATE lookups use now applies the shared installment rule too
(apple_music.sequence_mismatch): “Slime Language 3” and “Slime Language” share 2 of 3
tokens, so album_release_date returned the FIRST album’s 2018-08-17 for a record hours
old, which would frame a brand-new drop as eight-year-old catalog.utils/versuz_catalog.py, #2874): the /track/{id} read’s contributors[] is the
ONE structured source of a song’s credited cast (the search rows do not carry it, verified live
2026-09-03): lead = the track’s artist.name, and each contributor carries a role (Main /
Featured, read since #2885): the board names the Featured acts first, then the title’s own
“(feat. …)” credit, then the co-leads last-billed first (“Future, Metro Boomin & Kendrick
Lamar” is Kendrick’s verse). A name can carry both roles (Young Thug on “pick up the phone”). Read for
every song a Versuz night registers; surfaced as ft <first feature> on the poll and the board.
Gap: a posse cut is sometimes credited to the lead alone (“We Takin’ Over” -> [DJ Khaled]), so an
empty cast falls through to iTunes.isrc (not answer-valuable). Known API limitation, not a drop./track/{id}.release_date is deliberately NOT used as a release YEAR (#2724).
The field exists on the single-track read, but it reports the date of the edition
Deezer serves, not the original release. Measured against the live R&B chart: it
dates Prince’s 1986 “Kiss” to 2007, Al Green’s 1971 “Let’s Stay Together” to 2015,
Diana Ross’ 1980 “Upside Down” to 2003, Otis Redding’s 1968 “Dock of the Bay” to
1992 and Quincy Jones’ 1989 “The Places You Find Love” to 2005. A “2000s+” era
game reading those dates would keep every one of them. So a yearless Deezer row
that needs a year is resolved through iTunes instead (music.resolve_clip),
which returns the original date for the same nine tracks. The ISRC’s year digits
have the same defect (a reissue gets a new ISRC: Nirvana’s 1991 “Smells Like Teen
Spirit” carries USGF1994…), so they are not a shortcut either.recent_releases (Deezer’s all-genres album chart + per-album date lookups) was the board’s
only source once and left it dark for its whole recorded life (#2947); paired with Apple’s
feed (#2950) its rows were the long-tail noise on the card (Sleep, Rubber Band Gun, Paris
Paloma) and its tracks chart was weeks stale (#2956). The reader, its album-date cache, the
deezer_fetch event, the ops-monitor health branch and the dashboard panel entries are gone.
The stream-stat’s song-date check moved to the iTunes Search API (apple_music.song_release_date).
Deezer stays in the catalog facade, artist photos, Versuz credits and the alert lanes’ date
gates – those are other surfaces’ calls.utils/apple_releases.py) — the releases board’s sourceGET rss.applemarketingtools.com/api/v2/us/music/most-played/100/albums.json — 100
albums ranked by current plays, each row carrying its own releaseDate, in ONE keyless
request. It joined the board because reading Deezer alone left the surface DARK: Deezer’s
all-genres chart is LONG-TAIL, so on Friday 2026-09-04 it held three albums released that
day and missed Beyoncé’s B’DAY (20th ANNIVERSARY DELUXE EDITION), the day’s biggest
drop. Apple held four of that day’s drops, Beyoncé’s among them. It was paired with Deezer’s
chart in #2950, then made the only source on the owner’s steer (never Deezer)./browse/new-releases answers 403 for this app (re-measured 2026-09-04: the
client-credentials token call returns 200, the browse call 403). Deezer’s editorial
/releases feed returns an empty list. Apple publishes most-played only — the
new-releases, most-recent and coming-soon feed names all 404, and the legacy iTunes
RSS newreleases path 400s.artistName (the billed act), name (the album), and
releaseDate — the three fields the board lists and windows by, 100% of what it consumes.artworkUrl100, url,
id, kind, contentAdvisoryRating, genres. The board is a scannable list of names,
not a per-album detail card; a hero cover would come off artworkUrl100 on this same row.
Omission by surface need, not an API limitation.apple_releases_fetch event +
the apple_releases integration probe. A DIFFERENT host from the iTunes Search API in
utils/apple_music.py, with its own egress and its own failure mode, so it is probed and
rated separately and does not ride that module’s circuit breaker.utils/itunes_charts.py, #3469) — the iTunes ALBUMS chartGET itunes.apple.com/us/rss/topalbums/limit=100/explicit=true/json — the iTunes STORE’s
own Top Albums chart (a SALES ranking of album purchases), up to 100 rows in one keyless
request. The explicit=true segment is required. Without it Apple serves the CLEAN
chart, which removes every explicit album. On 2026-09-28 Tinashe’s Popstar (Bonus
Edition) was the store #1 and was absent from the default feed (65 rows against 100),
so no lane saw the #1 that @chartdata reported.
It exists because kworb mirrors no iTunes albums page (its /charts/itunes/ is songs;
the albums URLs 404), so a “#1 on iTunes” ALBUM claim had no reader: on 2026-09-18 Miley
Cyrus’ Bass Persuades led this chart while the song of that name sat at #32 on iTunes
songs, and nothing posted. A THIRD Apple host beside the Search API (apple_music.py)
and the marketing feed (apple_releases.py): its own probe (itunes_charts), its own
breaker, its own event (itunes_chart_fetch).im:name (the album title), im:artist (the billed act),
id.attributes.im:id (Apple’s album id — the identity the day-over-day diff keys on,
so a deluxe edition is its own entry), im:releaseDate (as YYYY-MM-DD, carried on
the row for the debut lanes’ date gate), and the row’s POSITION (the feed’s order; it
prints no rank column). The feed’s updated stamp is kept as refreshed_at() for the
crown lanes’ lead-time telemetry.pos_change and is_new. The feed prints today’s ranks only,
so the reader stores the previous UTC day’s final snapshot in kv_cache and diffs
against it (see the module docstring for why day-over-day and not read-over-read, and
why the first day is quiet). Every consumer of kworb’s rank-only rows then reads these
rows unchanged (ItunesAlbumRow extends kworb.AppleSongRow).im:image (three cover sizes — the cards resolve art through
catalog_art off the same Apple id, so the feed’s copy is redundant), im:price,
im:itemCount (track count), im:contentType, category (one genre id; the genre
boards read Deezer + the iTunes Search API for that), rights, link, title (the
“name - artist” composite). Omission by surface need, not an API limitation._credit drops an exact repeated party, keeping the
first, so the card reads the act once. An edition listing (“(The Miley Edition)”) keeps
its marker in song_pool.song_key, so it is its own entry beside the standard album.chart_boards.PLATFORM_CHART_ROWS
leaves it_albums out on purpose); the top-10 board, the new-entries board, the watch
cards and the alert lanes need only the top of the chart. Under 10 rows is shape_change.it_albums watch cards and
platform boards, the newsroom’s chart-claim row, and (2026-09-20) the X mentions lane’s
per-chart units + reply facts (cogs/x_mentions.py _chart_reader), which had answered a
“#1 on iTunes albums” question with the Apple Music rank because it read kworb only.ETag / Last-Modified (a conditional
GET answers 200 with a full body), so the client holds a parse 5 minutes
(ITUNES_CHARTS_CACHE_TTL) instead of revalidating like kworb. Health on the
itunes_chart_fetch event + the itunes_charts integration probe.utils/spotify.py) — DORMANT (#2447)/browse/new-releases (client-credentials auth) was the board’s first intended source,
but that Browse endpoint is 403 for any app created after Spotify’s Nov-2024 Web API
deprecation, so the board reads Apple + Deezer (both above). Re-measured 2026-09-04
against the live production credentials: the token call returns 200 and the browse call
returns 403, so the client cannot serve this board and NOTHING imports it. utils/spotify.py
spotify_fetch event + the spotify probe are kept for a possible future Spotify
use; nothing emits spotify_fetch today, and the event has never fired. Historical: it surfaced artists[0].name, name,
release_date (100% of what the board consumed) behind optional SPOTIFY_CLIENT_ID +
SPOTIFY_CLIENT_SECRET.utils/songstats.py, utils/streaming_stats.py, utils/milestones.py)The RapidAPI aggregator — /tracks/stats, /artists/stats, /tracks/historic_stats.
Each payload’s stats array carries ~20 platform blocks (spotify, apple_music,
amazon, deezer, youtube, tiktok, instagram, shazam, soundcloud, tidal, itunes,
beatport, traxsource, tracklist, facebook, twitter, songkick, bandsintown, radio).
streams_total (track + artist career),
monthly_listeners_current, followers_total (each explicitly Spotify-attributed)
— the milestone/verify/desk headline numbers. PLUS the free cross-platform reach
clause (streaming_stats.format_reach_context, appended to every Songstats context
blob — verify / desk / /ask / milestone — and to the milestone take): YouTube
video_views_total + radio radio_plays_total (both real absolute counts) +
Apple Music PRESENCE (charts_current, charted_countries_total,
playlists_editorial_current, artist charted_tracks_current). format_stream_stats
also names the top non-Spotify streams_total platforms (only Spotify + SoundCloud
report one). Every stream total passes a _MIN_STREAM_TOTAL (1,000) FLOOR first
(streaming_stats.meaningful_totals, the one reader of platform_totals): Songstats
indexes a brand-new release BEFORE it carries that release’s stream counts, so the
payload arrives with the real platform at 0 and a marginal one holding a two-figure
number. On 2026-08-29 Alabama Shakes’ “I Must Be Dreaming” returned {spotify: 0,
soundcloud: 15} and the new-release card shipped to X reading “SOUNDCLOUD STREAMS /
15” beside a ~29K first-week units projection. Under the floor such a payload reports
NO totals, so stream_hero yields no figure and _streams_for_isrc reports ok=False
– the kworb backup gets its try and the new-release lane stands down (it fails CLOSED
on the streams by design). The floor is a VALUE test, not a platform test, so a
genuinely SoundCloud-first act keeps its headline. All of this rides in the SAME
response — every stats call returns all
~20 platform blocks — so the reach clause adds ZERO extra Songstats calls.views_total/followers_total/videos_total — virality),
Instagram, Shazam (shazams_total), SiriusXM royalties, Beatport/Traxsource
DJ charts, and followers_total on the social networks. Numeric and free off the
same payload; slot into format_reach_context when wanted. Captured here so the
richness isn’t forgotten.utils/musicbrainz.py)disambiguation.utils/chart_ages.py, owner ask 2026-08-10): the
oldest-songs board reads title + artist-credit + first-release-date
(year only) per chart row via the same release-group search, title+artist
fold-verified before a year is accepted, cached durably 120d (kv_cache:
chart_ages; verified misses 14d), paced ~1.1s per live lookup under MB’s
1 req/s etiquette.utils/wikidata.py)utils/chart_data.py)<sup> markup (stripped for clean cells);
trend/debut arrows (the peak number is authoritative).utils/reference.py)utils/steam_art.py, #2741)The catalog rung for a VIDEO GAME subject, the way TMDB is film’s and Deezer is
music’s. Public store endpoints, no key. It exists because the Wikipedia lead-image
rung structurally cannot serve a game: pageimages EXCLUDES non-free files, and a
game article leads with its copyrighted box art (measured 2026-08-31 — every game
tested returns no image, while free-licensed subjects return one).
storesearch, matched EXACTLY on a folded name, never
“first result”); the portrait library capsule library_600x900_2x (600x900) when
present; else the first store screenshot path_full (1920x1080).header_image (460x215) and path_thumbnail (600x338) — both under the caller’s
400px big_enough floor, so a full-bleed cover would render visibly soft.background_raw — REJECTED on evidence, not on size. It is the store PAGE
WALLPAPER (deliberately abstract atmosphere), and rendered full-bleed for
STAR WARS Zero Company it is a near-empty starfield that reads as a blank card:
worse than the branded floor it was meant to beat. It passes a naive size check
(1438x810), so test_background_raw_is_not_used_as_art guards the choice.metascore — this is an ART
rung only; the card’s numbers come from Kalshi, and pulling store metadata into a
market card would mix two sources of truth.video_game kind, not on a name match. Exact name
equality proves the strings match, never that the subject IS the game: measured over
27 adversarial non-game things, “Uno” (the card game) matches the Ubisoft video game
exactly, and UNO’s Wikipedia article has no free lead image either — so a name-only
gate would have shipped a video game’s art for a card game. video_game was split
out of thing in image_subjects to answer that question honestly._deterministic_art, 400px shortest edge):
the screenshot fallback is not always large (Portal’s first screenshot is 640x360),
and resolve_market_image’s deterministic rung does not size-check on its own.steam in utils/integration_probes.py (keyless; the search leg, which
is the half that can go silently dead — if it stops answering, every game card falls
back to the branded floor with no exception raised anywhere).utils/markets.py, utils/sportsdata/board.py)fairOdds (sharp
consensus) alongside the book line, event URL.parse_game_board): every ou/yn market in the
~1,900-key tree — game lines, game props (corners, first-to-score, cards,
BTTS), and player props across all periods (quarters/halves/OT), each with
the book’s implied % and fair where it diverges. Labels derive from SGO’s
marketName, so a market type with no bespoke parser still flows./v2/account/usage (get_usage, quota-EXEMPT so it
answers even while degraded) reports tier + per-period requests AND entities
(the rookie tier caps entities at 100k/month — the cap that 429’d /bet dark).
Polled every ~30min → quota events → ops-monitor quota_low. The SAME read’s
per-minute requests/entities metric self-calibrates the proactive rate
limiter (order #46, SportsGameOddsClient.note_usage) to the account’s real
reported cap instead of the static SGO_RATE_LIMIT_PER_MIN guess, so a crater
caused by that guess being wrong for the account’s actual tier self-corrects.utils/markets.py)createdAt; full CLOB order-book depth
beyond top-of-book (surfaced via the detail tool, not the always-on block).utils/markets.py)createdAt, market owner; order-book depth
beyond 5 levels (readability); >72h candle history.title + parent event_title are qualified with the
platform their metric is measured on, read off that market’s own
rules_primary (markets.qualify_platform_metric, applied in
_kalshi_market_to_snapshot): “Morgan Wallen: Highest daily view count” →
“…Highest daily YouTube view count”, since Kalshi’s titles name a metric but
not where it’s counted and every surface renders the title verbatim. Grounded, not
guessed — exactly one platform must be named in the rules (never inferred from the
series ticker), so an ambiguous or already-qualified title passes through
untouched. rules_primary itself stays detail-only (below), so this adds no field.KXMLBGAME-26SEP211835TORBAL = Sep 21 2026, 6:35pm ET; NFL / NBA / NHL / soccer carry
the date only, KXNFLGAME-26SEP20INDKC; verified live 2026-09-21 against the event’s
expiration), and the event subtitle names the same date (TOR vs BAL (Sep 21)), which
the FTS indexes as 'sep' '21'. Every game-line caller now passes the game’s
start_ts through fetch_prediction_snapshots -> kalshi_snapshots -> kalshi_search,
which adds the date tokens to the FTS query, drops candidates whose ticker is another
Eastern date (or, with a time, a start more than 4h away – the doubleheader), and
names the date in the picker’s query. Before this the picker saw every open game
between two teams and a series game could price off a neighbouring day’s market (the
dry run on #3551 caught three MLB games priced off the wrong day). A ticker with no
date segment stays eligible; the /ask tools pass no kickoff and are unchanged.utils/the_odds_api.py)The SGO-down resilience backstop (#725 epic). Lower tier — 20,000 credits/month,
cost = #markets × #regions per /odds or /event-odds call; /sports + /events
are FREE, /scores = 1 (2 with daysFrom). x-requests-remaining is captured on
every call (client.credits_remaining) so the budget is observable.
Endpoints covered (all live v4 endpoints):
/odds — matchup, commence time, moneyline + spread + total, best price per
side across all books as a $100 payout at the consensus line; commence_time_to
bounds to the bookie’s next-2-days window./scores (ScoreEvent) — live + completed scores (home/away score keyed back
from the team name, completed, last_update). The live-scores provider is
RETIRED from the hub chain (#3552). TheOddsApiScoreProvider filled SGO-only-sport
scores while sgo.degraded; with SGO quota-dead for weeks (#3229) that gate stayed
open, and its 24/7 /scores reads (2 credits each, four sport keys, every refresh)
spent the whole month’s credits by the 20th — the reason no feed could price the
NFL Sunday slate. The free ESPN provider covers live + finals for every team sport,
always on (#3569), so nothing is left for a metered scores read to fill. The class
stays in utils/sportsdata/providers.py (the dry-run harness); it is wired nowhere
in bot.py. /scores now serves ONE caller: the Bookie’s settlement leg, below.
americanfootball_nfl_preseason, distinct from the regular-season
americanfootball_nfl. In the preseason window the regular-season key returns
EMPTY odds AND scores (the season has not started), so the Bookie was dark on NFL
and a preseason bet could never settle. The Bookie now queries BOTH keys (the slate
fetch _supplement_with_odds_api + the settlement _odds_api_finals, both via
_odds_api_keys_for); both keys carry the internal americanfootball tag
(_ODDS_SPORT_TAG), so a game and its final reconcile to the same match_key. An
added preseason game is stamped with the preseason key as meta['league'] so the
per-game odds/value/board surfaces re-key it correctly (odds_api_sport_key)./scores leg runs only while ESPN cannot settle
(#3552). It was ALWAYS-ON (#nfl-preseason) because API-Sports’ baseball and NFL
hosts are on a FREE plan (reads only seasons 2022–2024), so a 2026 MLB/NFL final
never arrives there while the breaker stays CLOSED. The free ESPN leg now meets
that need for every team sport (#3569; verified on 2026-09-20, six real NFL bets
settled off ESPN), so the metered leg is gated on Bookie._espn_settling() — the
ESPN provider is wired and the ESPN client’s breaker is closed — and a healthy day
spends no /scores credit at all. _settle_game is idempotent so an overlap never
double-pays; the stranded-bet detector and the 96h void remain the visibility if
ESPN misses a final./odds h2h slate read is the LAST pricing layer, spent only on a game the
free sources could not price (#3552). _supplement_with_odds_api used to run
FIRST and, whenever sgo.degraded, fetch every bettable sport from scratch. Now
the ESPN pregame line (folded at conversion, donated across rows from the
provider’s cache) and the prediction markets price the slate first, and a sport is
fetched only when it still has an unpriced game the Odds API has not already been
asked about in the last hour (_odds_api_misses, keyed by match_key, recorded only
when the source answered). Bound per sport: one credit per 10 minutes for a game it
prices (_H2H_CACHE_TTL_SECONDS, the pre-match line does not move in play), one per
hour for a coverage hole. The fetch’s other in-window events still join the slate.OUT_OF_USAGE_CREDITS, or a bad key) puts the
client in quota_dead for 6 hours: enabled reads False, every caller’s own gate
skips the metered path with no round-trip, the skip is labeled error=quota_dead
(not no_key) in market_fetch. The API does not publish the account’s monthly
reset date, so the client re-probes with one real call per window (a 401 bills
nothing) and one 200 clears the flag./events (EventInfo) — id, teams, sport, commence time (no odds, FREE) for
cheap event discovery./sports (SportInfo) — key, group, title, description, active, has_outrights./event-odds (PlayerProp) — per-event player props / alternate / period
markets: player (the outcome description), market key, Over/Under side, line
(point), payout, book. Wired into the /ask lookup_player_props tool as
the SGO-down BACKSTOP (#725): when SGO props come back empty, the named game is
resolved to its Odds API event via the FREE /events lookup and that ONE event’s
props are pulled (one metered /event-odds call, never a league sweep),
entity-filtered to the game’s players and rendered by format_odds_api_player_props.
The SAME /event-odds endpoint also backs the deep-game-market board backstop
(#725): get_event_board requests the deeper game markets (totals, both-teams-to-
score, corners, team totals, 2nd-half totals — ODDS_BOARD_MARKET_KEYS), parsed by
the pure parse_odds_api_board into the same GameBoard the SGO board produces so
format_market_edges/format_game_board render unchanged. Powers the commentator’s
market-edge beat and the /ask break_down_board tool when SGO is degraded.
Budget-guarded: the deep-board fetch only fires when sgo.degraded AND
client.has_enhancement_budget (credits above the _ENHANCEMENT_RESERVE of 5,000),
so the credits the /bet SGO-down slate + settlement backstop depend on are never
spent on the enhancement surfaces. ~5 credits per game per cache window (8 min).includeLinks=true bettable links — the per-outcome bet-slip link (else the
per-bookmaker event link) captured onto EventOdds.links ({outcome: (url, book)})
by _best_link_per_outcome: the best-priced book that has a usable link, since
_clean_link drops {state}-templated US-regulated books (BetMGM/BetRivers) we
can’t fill. Powers the betting VALUE alert (cogs/betting_value.py) — the value
is at the LAGGING book, so the card’s title links the sportsbook where you place it
(sportsbook_link in utils/sportsdata/format.py matches the value side), falling
back to the prediction-market link when no clean book link exists. includeLinks is
a paid-plan feature that adds NO credit cost (billing is markets×regions), so it
rides the existing /odds slate call for free.regions=us_ex exchange links (Kalshi-first card link, #kalshi-links) — the US
betting EXCHANGES (Kalshi/Polymarket/Novig/ProphetX) come back as bookmakers in the
us_ex region, PRE-MATCHED to the game by full team name. _event_links_per_book
captures each book’s clean event-page link onto EventOdds.book_links ({book: url},
keeping venue identity unlike _best_link_per_outcome); exchange_link picks
Kalshi first (CFTC-regulated, web-bettable) else Polymarket. Powers the betting
board + line-move alert card link (us_bettable_link in cogs/betting_board.py):
a US reader clicks the card title and goes to a market they can actually bet — the
fix for our own Kalshi discovery being unable to resolve a US-team-sport game line
(0/15, the city↔mascot wall). us_ex is a separate 1-credit-per-market region
call (billing is markets×regions) — cached per (sport, markets, regions). Coverage:
Kalshi on MLB/NBA/WNBA, Polymarket broad; soccer is Polymarket-only, tennis absent.
Venue-coherent card (the #kalshi-links follow-up): the us_ex Kalshi link embeds the
exact Kalshi event ticker (kalshi_ticker_from_link), so coherent_card_snap fetches
THAT event by ticker (kalshi_snap_by_ticker → KalshiClient.get_event_markets, free) and
renders the card’s chart + odds from Kalshi too — chart, odds, and link all one venue, no
“odds via Polymarket” split. Budget guard: the us_ex call is SKIPPED entirely when our
own prediction discovery already matched a Kalshi market (it already carries a Kalshi
link+chart+odds); us_ex is only reached for games our own discovery can’t resolve. A ticker
fetch miss keeps the Polymarket chart under the Kalshi link (the footer self-labels it).
NBA vs WNBA routing: one internal basketball tag spans TWO Odds-API sports, so the
sport_key is resolved by the game’s meta['league'] (bookie.odds_api_sport_key,
reusing sports.WNBA_LEAGUES) — a WNBA game routes to basketball_wnba (which has Kalshi
coverage) instead of basketball_nba. Shared by both the board’s exchange link and the
value alert’s sportsbook link.PlayerProp, where which book offers a prop matters, and on
EventOdds.links now, where you need to know which book to bet at);
includeSids/includeBetLimits/includeRotationNumbers (bet-slip ids + limits —
not answer-valuable for our surfaces); the multi-region books (uk/eu/au — us only, to
hold the credit cost at 1× region); the historical endpoints (/historical/*
— 10× credit cost, no live-resilience value); /participants (1 credit, just
full_name+id per team — we name-match already) + /event-markets (1 credit,
the available-market catalog per game — a future caller could use it to request
only the prop markets a game actually offers instead of a hardcoded list).x-requests-remaining/-used ride
every response (credits_remaining/credits_used), refreshed via a FREE
/sports call (refresh_usage). 188/20,000 used (~99% headroom) — the healthy
one when SGO + API-Sports were near their caps. Polled → quota events; also
the real-time market_fetch credits_remaining finding (#734).get_outrights, #2268): the same
/sports/{key}/odds path with markets=outrights on one region, so 1 credit
per call. Eleven sports carry has_outrights (verified live 2026-08-10: NFL
Super Bowl, NBA / NCAAB / NCAAF championships, MLB World Series, NHL, the four
golf majors, and the US presidential market); the sports boards use the five
league ones. Surfaced: every entrant’s name and American price, from the
SINGLE book quoting the MOST outcomes — the fullest field, chosen
deterministically rather than by taking whichever book the API happened to list
first — plus that book’s name for the card’s source pill. Omitted
(intentional): the other books, and any de-vigging. The implied percentages a
title-odds board shows sum to over 100% because the book’s margin is in them; the
card credits the book by name rather than presenting a blended number no source
published. Cached 12h, far slower than a futures field moves. Health rides
market_fetch source=the_odds_api_outrights.utils/api_sports.py)x-apisports-key account — v3.football
(soccer/World Cup), v1.basketball (NBA), and, for the Bookie’s settlement feed
(#mlb-betting), v1.baseball (MLB, league 1), v1.hockey (NHL, league 57),
v1.american-football (NFL, league 1), and v1.mma (UFC/MMA /fights, season-less).
League ids env-overridable (API_SPORTS_{MLB,NHL,NFL}_LEAGUE); seasons derived
per-call (MLB = calendar year, NHL/NFL = start-year; MMA needs none). All verified
live 2026-07-03 (Pro plan, 7,500/day). No tennis host exists (v1.tennis doesn’t
resolve), so tennis has no API-Sports finals feed — the free ESPN backstop
(utils/espn.py, #1113) fills tennis scores + is the intended tennis settler._team_game_to_snapshot) absorbs two shape
quirks: NHL scores are flat ints (scores.home, not .total) and NFL nests
id/status/date under an inner game object; only the final-score + status
(settlement-critical) fields are consumed for MLB/NHL/NFL, not the depth blocks.
MMA (_fight_to_snapshot) is a different shape again — fighters.first/second
with an explicit winner boolean (no scores), an event slug + weight category
(no league) — so the winner is encoded as a 1-0 ‘score’ for the shared settlement
path; a no-winner draw/no-contest reads 0-0 → a 2-way PUSH. Only UFC settles here,
so the Odds-API MMA supplement is fold-only (never adds a non-UFC promotion’s fight)./status (get_status, a free account read) reports
the plan + daily request cap (requests.current/limit_day; the Pro tier
sat at 90% — 6,767/7,500 — the day SGO died). Polled every ~30min → quota
events → ops-monitor quota_low. (Per-minute limits ride the response headers.)utils/highlightly.py)x-ratelimit-requests-limit/
-remaining ride every response, captured PASSIVELY off the post-game highlight
calls (the free Basic tier is only 100 req/day, too small to poll). Surfaced via
quota events when a recent call populated the headers.utils/espn.py, #1113)The FREE, keyless last-resort score/finals backstop for the SGO-only INDIVIDUAL
sports — tennis (SGO is its only source; no API-Sports tennis host) and
MMA/UFC (API-Sports /fights is its only settler; The Odds API /scores is
0-completed for MMA). Public site.api.espn.com site API, no auth. Wired as
EspnScoreProvider, the actual LAST provider in the hub chain (after
API-Sports + SGO; the metered Odds API scores provider left the chain in #3552), so
first-provider-wins dedup makes it a pure fill-in. Since #3569 it is the free WORKHORSE (owner steer 2026-09-21: “the
free source is the workhorse; the metered feeds are accuracy references, not
dependencies”): the capability matrix’s _ESPN_SETTLES names every team sport
plus tennis, bot.py derives the always-on gate from it, so ESPN’s live, upcoming
and finals reads run for all of them whatever the primaries say. API-Sports and SGO
rows still win the dedup when they exist. Only MMA keeps the api_sports.degraded
gate (API-Sports /fights is its settler). Polite by construction (“a free
service, not a resource”): a day slate with no game in progress and no kickoff inside
10 minutes is cached 5 minutes (_ESPN_TEAM_TTL_QUIET), a slate with something moving
30 seconds (_ESPN_TEAM_TTL), and concurrent live/upcoming/finals reads share one
in-flight fetch per (league, day). Measured before the flip: ~800 ESPN calls an hour,
almost all from the bet slate refresh; the TTLs are what keep the always-on flip at or
below that.
Endpoints covered:
/sports/tennis/{atp,wta}/scoreboard — an event is a TOURNAMENT; matches
nest under event.groupings[].competitions[] (by draw). SINGLES kept (clean 1v1
names for the fold; SGO prices singles), doubles dropped./sports/mma/ufc/scoreboard — an event is a fight CARD; fights are
event.competitions[] directly.displayName), the
authoritative per-competitor winner boolean, status (pre/in/post + completed),
the numeric set/round score when present, start time. The winner is encoded 1-0
on a COMPLETED match (mirrors API-Sports _fight_to_snapshot) so the shared
settlement path pays the winner unchanged; a no-winner finish reads 0-0 → a 2-way
PUSH. match_key folds fighter names via canonical_fighter, so an ESPN row
reconciles with an SGO/API-Sports bet. recent_finals is bounded to a 48h window
(a tennis scoreboard returns a tournament’s whole ~1,158-match draw; the window
cuts it to the ~dozens that just finished).status.type.detail (“11:48 - 2nd Quarter”,
“Final”) rides as TeamGame.status_detail and, once live or final, as the
snapshot period, so an ESPN-sourced game renders its clock.EspnScoreProvider.upcoming_games returns the
not-yet-started TEAM games inside the caller’s window (pregame_scan_days +
starts_within, the API-Sports provider’s own scan), for a team sport whose
primary is degraded, under the same per-sport gate as the live/final reads. It
used to return [] (“the pregame nudge rides SGO/Odds”); on 2026-09-20 SGO, The
Odds API and the API-Sports NFL host were all dead at once, so no feed put a game
on the bettable slate before kickoff. ESPN settles NFL + MLB, so its pregame rows
keep the rule that a slate game must come from a feed that can also settle it; the
prediction markets then price the row. Only a row ESPN marks pre is kept (a
postponed game keeps its scheduled start and is neither live nor completed), and
the Bookie’s enabled-sports set (extra_sports) gates the new team sports the same
way it gates the API-Sports day slates, so an unbet MLB slate never crowds an
enabled sport out of the bounded pricing pass. Individual sports (tennis/MMA) are
not surfaced pre-start.
espn_team_game_to_snapshot
folds the scoreboard’s moneyline (home/away, plus the draw for soccer; book
names the sportsbook) onto a pre row at conversion, stamped odds_source=espn,
so ESPN’s upcoming slate arrives PRICED with zero extra requests; a live or
finished row carries none (ESPN drops the line at kickoff). The Bookie’s
_supplement_with_espn_odds used to re-read every league’s scoreboard through the
raw client for that line, one request per league per day per slate refresh; it
now reads the provider’s upcoming_games (its per-(league, day) slate cache, which
the slate fetch just filled) and copies the line onto an API-Sports / SGO row of the
same matchup (fold_line_from_snapshot), so the fold costs no ESPN request in the
normal case. On a tie in priced sides the slate keeps the sportsbook-fed row over
the ESPN-stamped one (_slate_rank)./sports/{path}/summary?event= play feed (#3404): scoringPlays[] — the
play text (the whole call: “Joshua Palmer 43 Yd pass from Josh Allen (Tyler Bass
Kick)”), scoringType/type.text (touchdown / field goal / safety), team,
period.number + clock.displayValue, and awayScore/homeScore after the play
— all surfaced verbatim in a GameEvent (espn_summary_events). drives.current.
plays[] (football; drives.previous[-1].plays between drives; a top-level
plays[] for the other clock sports) — the last 6 plays’ text, period/clock,
start.downDistanceText, scoringPlay (espn_summary_recent_plays). Omitted:
per-play statYardage, yardsAfterCatch, win probability, wallclock (the play
text already states the yards; the rest is not answer-valuable in chat), and the
drive summaries (description/result).quota/
usage_fetch, correctly absent from the usage poll). Health is the market_fetch
source=espn ok-rate (ops-monitor integration-health + the ops-only espn
health-cog Watch) + the circuit_breaker integration=espn breaker; reachability
from the datacenter IP is the /debug/integrations espn probe (confirmed green
from prod: 200/31ms). A failed read now carries http_status (#3584), so a
429 throttle is distinguishable from an outage or a bad path — every miss used to
read as a bare fetch_failed._fetch_url could raise instead of returning None. retry_http re-raises once
its attempts are spent and CircuitBreaker.call re-raises too, so a sustained
refusal escaped the client; the hub catches a raising provider and discards
EVERYTHING it returned, so one throttled (league, day) read sank every other league
and day in the same gather. Measured: the bet slate fell from 16 games to 4. It now
fails open to None, per the contract its docstring always stated._RETRY_AFTER_CEILING_SECS), and
honours ESPN’s own Retry-After when a 429 carries one — the same shape SGO and
Kalshi already had.The STATS reads (#2268). The same host publishes, free and keyless, the
team-league data the sports stats boards are built on. Three more endpoints, each
its own market_fetch source so one dead read is separable from a healthy one:
apis/v2/sports/{path}/standings (source=espn_standings) — the league
table. Surfaced: per team the group (conference/division), the published
seed or table rank, games played, W-L-T, table points, win percentage, games
behind, the ACTIVE STREAK string (“W6”), the last-ten record, point
differential, the league’s own record string, and the team badge. Also the
season the ROWS are from (seasonDisplayName), which is the only field naming
it — the payload’s TOP-LEVEL season names the season ESPN currently points at,
which in an offseason has not been played yet. Per-league quirks, all verified
live 2026-08-10 and all load-bearing: the NBA omits gamesPlayed (W+L is its
games played); the NBA and MLB publish a points stat that is NOT table points
(Detroit at 60-22 came back with 19), so only a points league may show it; MLB’s
playoffSeed is a playoff seed, not a record position; soccer publishes no
streak at all; entries do NOT arrive in table order. Omitted: per-team
splits beyond home/away, the division sub-tables, and clinch markers.apis/site/v3/sports/{path}/leaders (source=espn_leaders) — season stat
leaders, ten deep per category, with each player’s team. Surfaced: the
category name + label, the athlete, the team abbreviation, the numeric value and
ESPN’s own formatting of it, and requestedSeason (the season the numbers ARE).
Quirk: MLB puts the player’s WHOLE batting line in displayValue
(“137-426, 35 HR, 24 2B, …”) on every category, so a value carrying a comma is
re-rendered from the number. Omitted: the categories outside each league’s
curated set (ESPN returns 20 for MLB; nobody posts double plays).apis/site/v2/sports/{path}/scoreboard?dates=YYYYMMDD (source=espn_results)
— one calendar day of team games. Surfaced: both sides by homeAway (a team
game identifies its sides, unlike the 1v1 path), display names + abbreviations,
scores, the per-competitor winner flag, completion state, start time, badges.market_fetch
ok-rates above, sharing the one integration=espn breaker (every read on this
host funnels through _fetch_url, so a bad prefix trips the same breaker the
scoreboard path uses).utils/box_office.py, the cinema desk’s box-office source)A scrape, not an API, so “the fields” are the chart page’s COLUMNS. BOM prints three
chart layouts and they do not share a column ORDER, which is why parse_chart reads
the <th> row and maps each column by NAME (#2266). A header-less snippet falls back
to the original positional heuristics.
| page | columns |
|---|---|
weekend /weekend/<YYYY>W<NN>/ |
Rank, LW, Release, Gross, %± LW, Theaters, Change, Average, Total Gross, Weeks, Distributor |
daily /date/<YYYY-MM-DD>/ |
TD, YD, Release, Daily, %± YD, %± LW, Theaters, Avg, To Date, Days, Distributor |
year /year/<YYYY>/ |
Rank, Release, Genre, Budget, Running Time, Gross, Theaters, Total Gross, Release Date, Distributor |
gross and grossToDate, but a mid-2026 release shows a LARGER year figure than its
to-date figure). So the year-to-date board prints only the Gross column – the one
BOM ranks the page by, which makes our order BOM’s order – and never the second one.sort=maxNumTheaters
per its own header link), not a current one, so a per-screen read belongs only to the
weekend page.movie_detail is the budget source),
the “New This Week” / “Estimated” flag columns, and the daily page’s second %± column.weeks comes only from a “Weeks” header. The daily page’s equivalent column is
DAYS, and the older positional parser silently read it into weeks. Nothing consumes
daily rows today, so this is a correction, not a regression.utils/tmdb.py, the cinema desk’s release + enrichment source)now_playing / discover_streaming / new_tv /
trending_movies / releases_between): id, title, release date, popularity, vote
average + count, overview, poster path.trending_movies order IS the ranking, and it is not the popularity field
(#2266): TMDB ranks that endpoint by its own trending score, so row 2 can carry a
lower popularity than row 3. A caller keeps the returned order; re-sorting by
popularity would publish a different ranking under TMDB’s name.releases_between filters on release_date with region + with_release_type=3,
not on primary_release_date – the primary date is a foreign film’s HOME-market
date, so the calendar would print a day the film does not open in the US on. It is
also why /movie/upcoming is unused: that list mixes in RE-RELEASES, which arrive
carrying their original release date (a 2016 date on a coming-soon board).region + with_release_type filter above selects a film that
has ANY US theatrical date in the window – a re-release included – but the
release_date field in the response is still the film’s PRIMARY date. So the filter
did not keep re-releases out at all; it let them in and handed us the original date.
Measured on the live 2026-09-15..2026-09-28 window: Avengers: Endgame came back
FIRST on popularity dated 2019-04-26, with Ghost in the Shell (1996-03-29) and
The Transformers: The Movie (1986-08-08) behind it. All three drew on the card
under the hero “Opening Next”, and the take called a 2019 film the most anticipated
of the fortnight. The response carries no field that holds the real US date for those
rows, so releases_between DROPS every row dated outside the window it asked for –
absent beats invented. The cost is nothing in practice: page one held 17-20 in-window
rows across five consecutive 14-day windows, for a card that draws 10 and needs 5.
tmdb_calendar_filter reports each drop, because the drop is otherwise silent.
/movie/{id}/release_dates is the endpoint that DOES carry the per-region, per-type
dates, if a future caller needs to date a re-release rather than drop it.movie_detail, one round trip with append_to_response=credits):
runtime, budget, worldwide-lifetime revenue, genres, tagline, overview, DIRECTOR and
top CAST, poster path.cinema_numbers._add_credits now enriches the theatrical
titles the roundup can name (bounded to 4 per slot against TMDB’s 600/min budget,
fail-open per title) and _release_line renders them in [brackets] – into the BLOB,
which is what cinema_desk_score grounds on. Supplying a field the prompt asks for is
the fix; asking without supplying is the bug./{movie,tv}/{id}/videos – trailer_url, the watch-link resolver’s
deterministic rung): site, key, type, official flag. It ranks an official Trailer,
then any Trailer, then an official Teaser, and takes YouTube keys only (the point is a
link X and Discord unfurl into a player, and TMDB’s other sites don’t embed the same
way). The rest of each video row is deliberately unread – name, size, language,
region, published_at – because the endpoint’s job here is one url, not a listing.
What this endpoint does NOT tell you is its own staleness, and that is the field
gap that mattered: it is community-maintained, so it holds nothing for a trailer that
dropped an hour ago and looks identical to a title that genuinely has no trailer.
Nothing in the response separates those two cases, so the resolver does not try – it
treats an empty answer as unknown and falls to a YouTube search (#trailer-lane).utils/omdb.py)The cinema desk’s critical scores. One endpoint (omdbapi.com/?t=<title>), read by the
scores story lane, the critics scoreboard board lane, and utils/reference_movie.py.
ALWAYS SEND THE YEAR (y=) – now ENFORCED IN CODE, not just written here. A bare
t=<title> is ambiguous and OMDb answers it with the older famous match. t=Moana
returns the 2016 animated film (96% RT, 81 Metacritic), not the 2026 live-action release
(31% RT, no Metacritic). The critics board shipped without the year on 2026-08-09 and
printed the 2016 scores under the 2026 title; the take then repeated them and the card
went to X. y= filters EXACTLY, so a film that opened last December needs a retry
against the previous year – the shape CinemaDesk._critic_scores uses.
Writing the rule down was not enough. It lived in that ONE caller, and three others could
still ask the ambiguous question: cinema_numbers and reference_movie both build the
year as release_date[:4] if release_date else "", so an undated film silently degraded
to a bare lookup, and entity_image’s poster rung never had a year at all. So
OMDBClient.by_title and .poster now REFUSE a request without a usable 4-digit year
(is_usable_year, which also rejects the "None" and "" a caller’s str(year) really
produces). The refusal is emitted as omdb_fetch ok=false reason=no_year – visible on
purpose, because a caller that quietly stopped resolving scores would otherwise look
exactly like the coverage gap described below, and those two must never be confused.
Measured live on 2026-09-14: bare t=Moana still answers 96% RT, y=2026 answers 32%.
Ratings array,
IMDb rating, title, year, rated, released, awards, box office..poster()
remains on the client for a caller that knows the year, but entity_image’s film/tv rung
no longer calls it. That rung could only ask by bare title – the classifier hands back
{kind, name} with no year, and the only thing that could supply one is TMDB, which has
just missed by the time the rung is reached. A bare title there returns the RIGHT name
over the WRONG film’s poster (t=The Uprising serves the 2013 film while the 2026
release is the subject), and it is reached precisely when a same-title hit is most likely
to be a different film. A miss now hands the card the branded floor instead, which this
repo ranks above a confidently wrong image.utils/rotten_tomatoes.py)The critics scoreboard’s RT column. There is no API — this reads RT’s public pages,
which is an owner decision (2026-08-11) and not an agreement. It sits outside their terms
of service and can break with no warning, so every path fails open and the rt_fetch
event carries the fill rate.
t=, not under i=<imdbID>, and not with tomatoes=true, which returns
every tomato* field as "N/A". So the board’s dash claimed “unreviewed” and meant
“our source lacks it”./search?search=<title> for the /m/<slug> candidates, then
the film page for media-scorecard-json (criticsScore, audienceScore) and the
JSON-LD Movie block (name, dateCreated).Moana returns
moana_2026, moana_2016, moana_2 and moana as the first four hits. The client
opens candidates in order and takes the first whose own release year is within one of
the chart’s year; nothing matching returns None. One year of slack is required, not
cosmetic — RT dates the 2016 Moana as 2017, by wide release.criticsScore.score), its review count and Certified Fresh
flag, the Popcornmeter (audienceScore.score), the film’s RT title, release year and
URL.rt_scores experiment, default STAGING, and STAGING means
what it means everywhere else: OFF never uses RT; STAGING uses it only when the BOARD
is also staged, so the card carrying these scores lands in #bot-logs for a mod to
check against the real site; PRODUCTION always uses it. The experiment routes nothing
by itself — cinema_boards decides delivery and defaults to PRODUCTION — so reading
“not OFF” as “use it” would have put unaudited scores straight into the room.moana_2 among the first hits for “Moana”. Titles are
compared normalized, which includes stripping RT’s own disambiguating suffix — it
titles a reused name The Odyssey (2026), and a strict compare dropped 3 of 8 real
matches on live data.utils/perplexity.py)strip_industry_projection cuts the sentence before the compose sees it. The
day-one / day-two STREAM partials the lane actually wants survive, including when
HITS is credited for them, and so does a bare HITS RANK. No other surface strips
anything.purpose=music_qualifier, owner ask 2026-08-24):
the music desk looks up the RANK a milestone gives an artist (“first female rapper
to hit the mark this year”, “highest debut”) and DOUBLE-verifies it — a Perplexity
lookup of the web plus an INDEPENDENT Grok LOOKUP of X (qualifier_lookup_pulse,
owner ask 2026-09-10; it replaced the Grok claim_verify_pulse corroboration read,
which only re-asked X about the web’s angle), else a second Perplexity read, then
the claude.qualifier_verify judge (fail-closed). Measured 2026-09-10 two hours
after Pop Crave’s Doja Cat post: the web read returned an unrelated “#41 among
artists”, the X read returned the exact claim with the account and time.
utils/milestone_qualifier.py; gated on the milestone_qualifiers experiment
(graduated, master-only).utils/grok_search.py, #1390)Real-time X (Twitter) grounding via Grok’s server-side x_search Agent tool
(Responses API on api.x.ai), the sibling of Perplexity. Wired into: the on-demand
/ask search_x tool (alongside the ScrapeCreators social tools) AND a live X-pulse
grounding block on the scheduled trend/take surfaces — discourse (one source among
Perplexity/markets), music (a fresh-hit signal: “what tracks/artists are people
posting about on X”; she names the real track, never pastes an X link into the
links-only channel), recap (what X is saying about what the room’s on; same
not quiet gate as Perplexity, #880), the trending-clip reactions (clip-scoped:
what people are saying around the pool’s trending moments, so a “did you see this”
lands on WHY it’s buzzing – context for her take, never a “who said what” relay),
and the market surfaces market_drop + market_alert (the topic-scoped X CROWD
READ via the shared market_topic_pulse helper, so the take plays crowd-vs-reality
the bare %s can’t – an alert can even explain WHY a market just moved; sentiment
only, the compose still quotes numbers ONLY from the market blob). Two lanes are
carved OUT of that wiring, all structurally (withhold the input, do not only ban
the output). The rule: a card that REPORTS does not get the crowd read (owner
steer, 2026-08-31). The crowd read serves the crowd-vs-reality angle, and that angle
needs a live question to be about. Three kinds of card have none:
forecast_move plus its
closing range / ladder ending_soon branches, AND market_drop’s forecast /
forecast_rank projection angles. The post IS a projected number. This block is
what turned a clean figure into “…forecast at ~7.1M, well below the streaming
dominance the timeline keeps citing” (live 2026-08-31).settled_out report, and
market_drop’s >= 95% report lock. The question is answered.frontrunner. The card states who is out front; the crowd
does not get a vote on that.The gate flag (forecast_lane, which drives the self-gate’s opinion-clause hard
fail) stays scoped to (1) alone, because its wording is forecast-specific. Which
INPUT the compose gets and which GATE rubric applies are two separate questions. The withheld
input is the WHOLE fix, and that is measured, not assumed. An added “do not
editorialize” instruction shipped alongside it at first and was then cut: on the
real blob, n=8 per arm, crowd-read-out scored 0/8 opinion tails with the rule and
0/8 without, while crowd-read-in scored 3/8 without. Either lever alone fixes it,
so the instruction was pure prompt cost – and the enumerated-ban shape #1140
warns about. tests/test_market_alert.py ratchets it: a test now FAILS if the ban
is added back.
Every carve-out is known BEFORE the fetch, so the Grok call (and Perplexity, on the
market_drop path) is skipped outright, not paid for and discarded. SELECTIVE by design (pricier than Perplexity, so
only where live X sentiment is the point — not blanket-wired). Default model
grok-4.3 (GROK_SEARCH_MODEL-overridable): a live dry-run (2026-07) measured it
at ~3 x_search calls / ~8-11k input tokens / ~$0.03 per call with 7-9 real X
citations (~5-8× Perplexity), vs the frontier grok-4.5 which drives the agentic
loop ~14 tool calls / ~275k tokens / ~$0.3-1.2 for the SAME all-X citation
quality — a 10-20× cost delta with no quality gain, so the cheap model is the
default (max_tool_calls does NOT rein grok-4.5 in — model choice is the only cost
lever). grok-4.6 was measured the same way (2026-09, #model-upgrade) and stays
OFF the default for the same reason: on the same two grounding queries it ran
4-6× slower (48-69s vs 12s), read 1.5-4× the input tokens (12.0k/42.2k vs 8.0k/10.2k),
and returned the SAME or FEWER citations (5 vs 7 on one, 4 vs 4 on the other) — on top
of a rate 1.6×/2.4× above 4.3. Citation yield is the only thing a grounding block is
bought for, so the newer model is a straight loss here. Gated on GROK_API_KEY;
fail-open. Emits grok_search.
url_citation annotations, appended as a linkable SOURCES block so
the model can cite/link the actual posts). Optional allowed_x_handles scopes to
named accounts; from_date/to_date bound the window.web_search results (X-only by design).xAI API reference (verified live against api.x.ai, 2026-07-15). The single
place we keep the confirmed xAI facts, so a future session (image failover #1389,
text backend #1391, video #1392) doesn’t re-derive them:
/v1/chat/completions (OpenAI-compat, function tools only) ·
/v1/responses (the Agent Tools API — server-side tools live HERE) · /v1/messages
(Anthropic-compat, returns 200) · /v1/images/generations + /v1/images/edits ·
/v1/models, /v1/language-models (pricing) · /v1/api-key (key metadata).
Base URLs: https://api.x.ai/v1 (OpenAI SDK), https://api.x.ai (Anthropic SDK).grok-4.6 (the current top model,
$2/$6 per 1M in/out, 500k ctx) · grok-4.5 ($2/$6 per 1M in/out, <200k; $4/$12 ≥200k) ·
grok-4.3 ($1.25/$2.50; $2.50/$5, 1M ctx) · grok-4.20-{reasoning,non-reasoning,multi-agent}
($1.25/$2.50) · grok-build-0.1 ($1/$2, a code model) · grok-imagine-image /
-quality (image) · grok-imagine-video / -1.5 (Aurora video). Cached input
~10× cheaper. There is NO grok-4.1-fast/mini (a common web-doc error).types are x_search, web_search, code_interpreter,
collections_search. Request them in tools:[{"type":"x_search", ...}] on
/v1/responses. xAI features grok-4.5 as the tools model, but we default to
grok-4.3 — it drives the x_search loop far leaner (~3 vs ~14 tool calls) for the
same all-X citation quality at ~10-20× lower cost (dry-run-measured, see above). The
old chat-completions live_search tool + search_parameters are 410-retired —
Agent Tools is the only path.x_search/web_search/code_interpreter = $5 / 1,000 calls
($0.005/call); collections_search = $2.50/1k. Plus normal token cost on the
retrieved content. A real x_search ask ≈ $0.02–0.035 total (~3–5× a Perplexity
call, NOT 10–20×; most of it is the retrieved-post tokens, not the search fee)./v1/api-key returns key metadata only (name,
team, ACLs, blocked flags), no spend/balance. So xAI can’t be polled for a metered
budget like SGO/Odds-API; track cost via usage.cost_in_usd_ticks on each
Responses call or the xAI console.utils/x_fetch.py)video_fetch.exceeds_x_video_limit, #video-trim).utils/twitterio.py, the content curator’s X source, #curator)external_id, media THUMBNAIL URLs
(photo + video still, media_urls), a native-video flag (has_video, from the
media type/video_info — so a wire desk can re-share the clip via X’s own
.../status/<id>/video/1 deep-link, which X embeds as just the video, #video-upload),
author handle, engagement counts (like/retweet/reply/view), possibly_sensitive,
and the reposted author’s handle (reposted_handle, for repost-weighted adjacency
discovery, #1251); the @-mention display names (entities.user_mentions →
{handle: name}, JSON-encoded into extra[MENTIONS_KEY], read back by
source_providers.post_mentions, #2010); the post’s expanded external links
(link_urls, from entities.urls[].expanded_url — NOT the tweet body, which carries
only the t.co shortlink — so a wire post that links a video it announces
(“Watch: youtu.be/…”) hands the watch-link resolver a trusted rung-0 link,
#watch-link-source); per followings crawl — the account’s
followed handles + follower counts (fetch_followings, for the follow-graph half of
adjacency discovery, #1232 Phase 2); per follower-churn sweep — every account that
FOLLOWS us (fetch_followers → /twitter/user/followers, cursor-paginated 200 rows a
page, #x-churn): the X user id (the diff’s identity — rename-proof), userName,
display name, followers_count and statuses_count (the low-signal / follow-back-bot
read), plus the has_next_page/next_cursor pagination state, which is carried through
as FollowerFetch.complete because a truncated walk must never be diffed.favourites_count, can_dm, protected, verified and the account’s created_at.
The churn diff needs identity plus the two counts that separate a reader from a
follow-back bot; the rest is per-follower profile data we would be storing about
private individuals for no question it answers (data minimization, per the
constitution). Not available from ANY X API, so not omitted by choice: per-viewer
impression data. A post carries one aggregate impression count and no viewer list, and
twitterapi.io exposes no likers endpoint (/twitter/tweet/likers and
/twitter/tweet/favoriters both 404, verified 2026-09-14; /twitter/tweet/retweeters
does work). So “which posts did this account see” is unanswerable — the churn sweep
narrows it to the posts in the departure’s window, and the reply/mention/retweet record
is the only evidence that a specific account actually read a specific post.source_providers.plain_text, #2246).
X returns & as the entity & in text (and in a mention’s display name), and
nothing downstream decoded it. Measured on live wires: 4 of 60 recent posts across
@hiphopnumbers, @popbase and @chartdata carried one. It did not stay cosmetic — a raw
entity could REWRITE an artist’s name. The music newsroom shipped a milestone card
titled “We Still Don’t Trust You - Future & Metro Boomin”: music_news.wire_spelling
re-spells a classified artist the way the WIRE spells it, it strips punctuation per
token, and _TOKEN_STRIP contains ; — so “Future & Metro Boomin” offered the
candidate “Future & Metro Boomin”, which cleared corrected_name’s near-identical
bar (ratio 0.93, same first letter) and REPLACED the name the classifier had read
correctly. The Axiom trail shows the hop exactly: music_news_recover logged the right
name at 23:10:10 and every later event that slot logged the mangled one. So the decode
runs once, at ingest, where the classifier prompt, the compose, the art matchers, the
card title and the dedup key all read the same plain text. The same decode covers the
second X provider (scrapecreators.sc_tweet_to_post’s legacy.full_text, the same
payload through another vendor) and Reddit (reddit.reddit_search_result’s title +
selftext, escaped by the same convention). The general rule: normalize an upstream text
field at the ONE boundary that builds the shape, never at each reader — a reader that
matches on names can otherwise promote an encoding artifact into data. The class stays
open by construction: this fix is per-provider, so a source that starts escaping a
field would leak again. Deliberately NOT changed, on measurement — kworb’s Spotify chart
& and no entities (0 &, 4 raw & in link text, live
2026-08-09; its YouTube parser already unescapes), and TikTok desc + Instagram captions
return raw text. #2247 tracks the output-side detector that would close the class./video/1
deep-link rather than re-host the bytes, so the mp4 URL is never needed); the raw
full-metadata dump..@whamcbfw4's "Dead Fresh" enters the Top 20), so dropping the
entities left the newsroom unable to tell who a story was about — it shipped a card
reading “Dead Fresh - whamcbfw4” about Lil Baby, and “IS IT LOVE - Tylla” about Tyla.
X already returns the display name for every mention, so it costs no extra call. It
travels as EVIDENCE beside the handle, never as a substitution: measured across the
trusted wires, the display name is right often enough to matter (@whamcbfw4 = “Lil
Baby”, @Tyllaaaaaaa = “Tyla”, @ellalangleymsic = “Ella Langley”) and wrong often
enough that substituting it would break cases that already work (@Drake = “Drizzy”,
@tylerthecreator = “T”, @Latto = “BIG MAMA”, @MacMiller = “Mac”). The classifier gets
both forms and is told to trust neither alone.AsyncRateLimiter (proactive throttling under the free tier’s 1 req/5s) +
CircuitBreaker + retry_async, results DurableCached; @instrumented as
curator_fetch. Watched by the health cog (utils/health.py WATCHES, ops-only
fix-order on a crater) + the ops-monitor integration-health dashboard + a
/debug/integrations probe.GET /oapi/my/info
(get_usage, quota-EXEMPT + NOT breaker/rate-limiter guarded, so it answers even
while the paid reads fail) → recharge_credits + total_bonus_credits summed to
the remaining BALANCE (parse_twitterio_usage’s credits metric, limit=None).
With no fixed cap the generic pct-based quota_low can’t fire, so the ops-monitor
flags a dedicated balance-FLOOR finding (twitterio:budget, high) when the balance
drops under TWITTERIO_CREDITS_LOW — before it 401s the curator dark (“Credits is
not enough”), the same prepaid-wallet model as the Odds API budget finding.utils/x_poster.py, the @tootsiesbar crosspost target)POST /2/tweets (OAuth 1.0a user-context)
and consumes only the returned data.id (the created tweet id), which rides
telemetry (tweet_posted) and never reaches the model’s context.AsyncRateLimiter (courtesy
pacing under X’s write cap) + CircuitBreaker (integration=x_poster) +
retry_http; every write leg (post_tweet/repost/upload_media) is
@instrumented as tweet_posted; fail-open (any miss just skips the crosspost,
the room post is unaffected). Provisioning-gated on the 4 X_* OAuth env vars;
runtime behavior is the per-guild x_crosspost experiment.tweet_posted into
integration-health as x_poster (benign unprovisioned/empty_text excluded) and
flags integration_unhealthy on a sustained fail rate; a sustained crater also
trips the circuit_breaker (integration=x_poster). Latency p99 is gated by the
tweet_posted ceiling. No write-quota poll (deliberate): the drops volume is
far under the Free tier’s ~17 posts/day, and X exposes no cheap standalone
write-quota endpoint (only per-response rate headers), so a cap wall is unlikely and
would surface via the fail-rate rollup anyway.utils/scrapecreators.py, utils/social_search.py, #1272/#1273)One SCRAPECREATORS_API_KEY, one guarded REST client, TWO roles:
BackupXProvider over /v1/twitter/user-tweets,
routed behind twitterio.XProvider’s breaker via FailoverProvider (only while
the primary is degraded). Maps a tweet -> SourcePost (text, permalink,
external_id, media URLs, author handle, like/reply/retweet/view counts,
possibly_sensitive, reposted_handle). No X follow-graph endpoint, so
fetch_followings returns [] (adjacency degrades to the repost path).TikTokProvider): a SourceProvider
(platform tiktok) on the shared client for the curator’s seeding: fetch_recent
(/v3/tiktok/profile/videos → a seed’s own videos → SourcePost via the pure
tt_video_to_post: aweme_id, clean url permalink [tracking stripped], author
handle, cover image for has_media, digg/comment/share/play counts, and the clip’s
create_time [epoch seconds] normalized to ISO into extra["created_at"] so the
curator’s recency gate utils.curate.is_fresh can drop a stale pick — before this
the gate read nothing and failed open, so month-old TikTok clips shipped [#curator]),
and fetch_followings (/v1/tiktok/user/following → clean (unique_id, follower_count)
handles for adjacency). TikTok’s “repost” analog is the COLLAB: a video’s
collab_info.collaborators[].user_info.unique_id names the co-creator directly
(no resolve call), surfaced as SourcePost.endorsed_handles so a seed co-creating
with an account is a strong adjacency endorsement (weighted like an X repost).
TikTok has no readable repost feed, so is_repost is always False. Bare caption
text_extra @mentions are id-only (would need a paid resolve) and deliberately
NOT used — collabs are the free, high-signal endorsement. image_post_info.images[]
(TikTok photo mode) vs the item’s own video object decides
extra["media_kind"], the same vocabulary Instagram’s media_type feeds. Emits
curator_fetch provider=scrapecreators (phase=user|followings), with undated =
how many of the returned posts carried no create_time.InstagramProvider): a SourceProvider
(platform instagram) on the shared client, the third curator seed source beside
X + TikTok: fetch_recent (/v2/instagram/user/posts → a seed’s own posts/reels →
SourcePost via the pure ig_post_to_post). Surfaced: the SHORTCODE (code) as
external_id (so a channel-URL dedup matches curator.parse_external_id, not the
numeric id), the canonical /p/<code>/ permalink, user.username, caption.text,
the LARGEST image_versions2.candidates[] still for has_media + the vision judge,
like_count / comment_count / ig_play_count, taken_at [epoch seconds]
normalized to ISO into extra["created_at"] for the same recency gate as TikTok
(is_fresh), and media_type → extra["media_kind"] (1 = photo, 2 = video,
8 = carousel), which is what lets the artist curator’s caption name the post in
words when it carries none of its own. Omitted on purpose: carousel_media_count
— the caption quotes what a post SAYS rather than counting what it holds (owner
steer 2026-09-15), so a mapped count would have no reader.
WHY v2 AND NOT v1 (2026-09-15). v1 returns Instagram’s GraphQL posts[].node
shape and BOTH of its time fields went null upstream: taken_at_timestamp and
created_at were null on every post of every handle probed live (beyonce,
theweeknd, sza). Nothing raised — is_fresh reads an absent stamp as fresh
(fail-open) — so the 24h curation window stopped filtering Instagram altogether and
the artist curator shipped a SEVEN-MONTH-OLD Beyonce post (DUrhXmwEUF_, taken
2026-02-13) as a fresh pick. v2 costs the same one credit, carries a populated
taken_at, and states the media kind on top. The undated count on every
curator_fetch phase=user read is the tripwire for a repeat. IG’s
“repost”/collab analog is coauthor_producers (the
co-author usernames, right in the response) → SourcePost.endorsed_handles, the
same adjacency endorsement as a TikTok collab. Omitted (deliberate): IG exposes
no follow-graph or repost endpoint (/v1/instagram/user/following 404s), so
fetch_followings returns [] and adjacency leans entirely on the coauthor collab
(like TikTok’s no-repost-feed); is_repost is always False. The curator rewrites
the posted IG permalink to kkinstagram.com (to_fixup_link) so Discord unfurls
the media — native instagram.com links unfurl poorly, the reason the embed-fixer
mirrors exist (the mirror is swappable; confirm in the staging audition). Emits
curator_fetch provider=scrapecreators (phase=user).utils/video_fetch.py): TikTok +
Instagram video transcripts route through ScrapeCreators (yt-dlp resolves those
two hosts poorly from a datacenter IP). fetch_video tries
_fetch_scrapecreators FIRST for is_tiktok_url/is_instagram_url clips —
TikTok /v2/tiktok/video?get_transcript=true (WEBVTT transcript flattened by
the shared flatten_vtt + rich aweme_detail metadata: desc→title, author,
play/like/comment counts, duration ms→s via sc_tiktok_meta), Instagram
/v2/instagram/media/transcript (plain AI text). A transcript-less miss falls
through to the existing yt-dlp path (no regression). Emits video_transcribe
source=scrapecreators.utils/comments.py): the read_comments
/ask tool over the per-platform comments endpoints (/v1/tiktok/video/comments,
/v2/instagram/post/comments, /v1/youtube/video/comments,
/v1/reddit/post/comments) — plus X/Twitter REPLIES via twitterapi.io
(/twitter/tweet/replies, XProvider.fetch_tweet_replies), since ScrapeCreators
has no tweet-replies endpoint (confirmed against the catalog) and the reply
thread IS the audience reaction on X. Surfaced: each comment/reply’s text +
like count (normalized to one Comment(text, likes) shape, ranked loudest-first;
the X mapper reads likeCount). Omitted (deliberate): author handles (the
crowd’s WORDS are the signal, not who said them — avoids per-user attribution),
reply threads/timestamps (top-level reactions only), and any permalink (Reddit
stays link-free, owner steer). Emits comment_read.utils/social_profile.py): the
social_profile /ask tool over the per-platform profile endpoints
(/v1/tiktok/profile, /v1/instagram/profile, /v1/youtube/channel,
/v1/twitter/profile). Surfaced: handle, display name, verified badge,
follower count, following count, post/video count, bio (200 chars), account-age
string (normalized to one SocialProfile shape). Omitted (deliberate): avatar/
banner image URLs (not answer-valuable for a “how many followers” question), the
full follower/following lists (a stats read, not an enumeration), and platform-
specific extras (TikTok heartCount/room, IG business category, YouTube view
total) — the reach signal is the follower + post count, and the rest is noise for
the model’s use. The structure complement to social_search (which finds posts);
the model relays the real numbers, never invents a follower count. Emits
social_lookup.utils/reddit.py): the
search_reddit /ask tool over /v1/reddit/search (+ /v1/reddit/subreddit/search
when scoped). Surfaced: post title, subreddit, selftext (the discussion text,
280 chars), score + comment count (the traction signal). Omitted (deliberate,
owner steer): the permalink / any Reddit URL — Reddit is an INPUT source the
model synthesizes in her own voice, never a link she pastes (“not reddit-nerdy”);
RedditPost has no permalink field, and a unit test asserts the formatter never
emits a reddit link. Omitted (policy): NSFW posts (over_18 dropped) + author
handles (not answer-valuable, avoids per-user attribution). Ranked by score;
emits reddit_search.utils/social_search.py): the model-facing
search_socials (keyword: TikTok /v1/tiktok/search/top + YouTube
/v1/youtube/search + Instagram reels /v2/instagram/reels/search),
discover_trending (TikTok /v1/tiktok/get-trending-feed region-scoped +
YouTube /v1/youtube/shorts/trending), and search_hashtag (TikTok
/v1/tiktok/search/hashtag + YouTube /v1/youtube/search/hashtag). All map to
one SocialResult shape.SocialResult): platform, caption/title, permalink (the
answer payload — Discord unfurls it), author @handle/channel, verified badge
(TikTok author.custom_verify / IG owner.is_verified), a compact engagement
stat (TikTok digg/play; YouTube views; IG like/view), duration (TikTok
video.duration ms, IG video_duration s, YouTube lengthSeconds/durationMs/
formatted-string fallback), published date (TikTok create_time normalized from
unix-int-on-trending vs ISO-on-search), and a blended engagement magnitude for
ranking.music/clips_music_attribution_info — niche for topic discovery, a
transcript/read_media concern); author follower counts + full profile blobs (a
per-hit credibility number would bloat each line; verified is the compact signal
kept); thumbnails/cover art (Discord’s unfurl supplies the preview); TikTok Shop /
ads / age-gender / audience-demographics endpoints (out of scope for content
discovery). The dead /v1/tiktok/hashtags/popular catalog is deliberately NOT
wired (“TikTok took this page down” upstream — we SEARCH hashtags, not list them).AsyncRateLimiter + CircuitBreaker +
retry_http, @instrumented. Account endpoints answer content-type: text/plain
so the client reads resp.json(content_type=None). Fail-open per platform (a
source miss contributes nothing). Watched by the health cog + ops-monitor
integration-health + a /debug/integrations probe (all #1273).GET /v1/account/credit-balance
(get_usage, quota-exempt) → parse_scrapecreators_usage’s credits metric
(limit=None); the ops-monitor flags a balance-FLOOR finding
(scrapecreators:budget, high) under SCRAPECREATORS_CREDITS_LOW. Per-day spend
is ALSO surfaced (#1272): get_daily_usage() reads /v1/account/get-daily-usage-count
(a {usage_date, total_credits, request_count} list) and collect_usage folds the
most-recent day into an INFORMATIONAL day_credits burn-rate metric (period=day,
limit=None -> rendered in the quota table + /debug/usage, never flagged) so the
spend RATE is watchable alongside the balance floor. Emits usage_fetch
(source=scrapecreators_daily).utils/video_fetch.py)is_live flag (capped by duration regardless),
format/ext, raw full-metadata dump.utils/link_enrich.py)utils/gifs.py)rating (applied as a server-side pg-13
query filter); trending flag; creator username (privacy).utils/image_gen.py)revised_prompt (the safety-rewritten prompt,
to the caller — flagged-only in telemetry since it’s user-derived).created timestamp; n>1 (single image per prompt).utils/xai_image.py, #1388/#1389)The FAILOVER image backend: ImageClient retries a generate/edit on Grok when
OpenAI either safety-rejects the prompt (a content refusal on real named public
figures — permanent for that prompt) or is unavailable (429/5xx/timeout/
connection drop — an availability failure Grok covers on a separate quota),
mirroring the text provider-failover (#973). It does NOT fail over on a permanent
4xx (malformed request / bad key) or a decode miss. Gated on GROK_API_KEY;
fail-open. image_generated events carry provider="xai" (vs "openai") so each
backend’s health is separable; the failover emits provider_fallback
(requested=openai, fell_back_to=xai, reason=safety_reject|transient).
data[0].b64_json, or a hosted
data[0].url downloaded on the fallback path). Aspect maps to xAI’s
aspect_ratio (square→1:1, landscape→16:9, portrait→9:16). Edit passes the
source as a base64 data-URI in JSON (xAI is JSON-only; OpenAI’s edit is multipart).revised_prompt (xAI doesn’t return one); created;
n>1 (single image per prompt); multi-image edit (up to 3 refs supported upstream,
we send 1). Note the bytes come back JPEG, not PNG — fine for Discord delivery.utils/embeddings.py)index, echoed model id, usage tokens (not
answer-valuable; the vector is the product).utils/stt.py)utils/tts.py)GET /v1/user/subscription (get_subscription, a FREE
account read that doesn’t consume the character quota it reports) → the monthly
character cap (character_count / character_limit →
parse_elevenlabs_usage’s month_characters metric, the SGO-entities analog
for the voice budget) + tier + next_character_count_reset_unix. Polled →
quota events → quota_low ahead of the wall. Caveat: the read may need
the user_read key scope (the wall we hit on /v1/user with the STT key); a
scoped key gets a 401 and the report lands ok=False → quota_unreadable, so
the blind spot is visible, not silent.utils/gifs.py, #753)X-RateLimit-Limit-Day / -Remaining-Day) ride on every search response and
are captured PASSIVELY (last_requests_limit/-remaining, like Highlightly) →
giphy_usage_report’s day_searches metric. The daily window accumulates, so
it flags quota_low ahead of the wall (a beta key is ~100/day). The hourly
window is dropped (resets too fast to be a budget). Headers read
case-insensitively (Giphy’s casing is inconsistent).utils/embeddings.py + utils/image_gen.py, #753)/v1/organization/costs
+usage), which the owner declined (rate-headers-only), and the $ cap isn’t
machine-readable. So we capture the rolling per-minute request rate headers
(x-ratelimit-limit/remaining-requests) PASSIVELY off embeddings responses →
openai_usage_report’s minute_requests metric (period="minute"). This is
INFORMATIONAL: the ops-monitor renders it but NEVER flags quota_low on it
(a per-minute window resets constantly — it can’t fill ahead of time; the real
hard signal is the 429, already in integration-health). Our embedding workload
sits near 0% of the RPM cap, so it just makes a runaway burst visible.402 on exhaustion); only
X-RateLimit-Remaining (no limit header → no pct) rides on responses./debug/integrations), not polled — there’s nothing pollable.utils/github.py, check_fixes)state_reason (completed/not_planned/reopened), created/updated/closed relative
ages, comment count, html_url; labels via the query.GET /rate_limit (get_rate_limit, FREE — calling it
doesn’t count against the limit) → the REST core hourly budget
(resources.core.used/limit → parse_github_rate_limit’s hour_core_requests
metric; the real limit is read live — 5,000/hr for a PAT, 1,000/hr for an
Actions GITHUB_TOKEN — never hardcoded). The bucket the /order
issue-filing + reconciliation path spends against, so a runaway burn flags
quota_low. Search/graphql buckets omitted (the bot’s client is REST-core only).utils/railway.py, check_deploys)These are the only answer-valuable fields a user could plausibly want that we still drop. Each is a conscious low-priority deferral, recorded here so it’s explicit:
| Integration | Field | Why deferred |
|---|---|---|
| Deezer | track-level release date / explicit | Not in Deezer’s track-search response shape (an upstream limitation; the album catalog path carries the date). |
(API-Sports venue + referee, previously listed here, are now surfaced via GAME CONTEXT; only the Deezer upstream limitation remains.)
Everything else is either surfaced or an intentional omission listed above.
The music desk’s source watcher (docs/ARCHITECTURE.md) must see a new edition within minutes, so four reads got a faster re-read. Field coverage is unchanged.
utils/billboard.py): the WAITING re-read, used only while a chart flip is due, is 2 minutes (was 20).utils/hits.py): the CURRENT projection document is held 5 minutes (was 6 hours), because HITS edits a document in place under the same tracking date; the WAITING re-read is 2 minutes (was 20); the radio documents are held 5 minutes (was 90).utils/riaa.py): the newest-awards read (recent) is held 10 minutes (was 2 hours).utils/pollstar_charts.py): the current-issue lookup is held 10 minutes in its own cache (was 6 hours, shared with the bodies). The bodies stay 6 hours, keyed by issue.