tootsies

a discord bot for the tootsies server. ask, recap, discuss, ship features by typing.


Project maintained by mejasonmejason Hosted on GitHub Pages — Theme by mattgraham

Upstream integration field coverage

What every external API hands us vs. what we surface to the model. This is the “are we dropping anything” ledger: each integration lists the fields we extract and the fields we deliberately omit (with the reason). The rule is either 100% of the answer-valuable fields, or an explicit entry here saying what’s missing and why — so an omission is a documented decision, never a silent gap.

Scope note: “surface to the model” means the field reaches Claude’s context (a rendered block or a tool result), not telemetry. The constitution’s data-minimization rule governs telemetry + her outbound replies, NOT the in-context input we feed the model, so enriching context is in-spec. A handful of fields are withheld from output by the HARD RULES (no personal-info disclosure) even when fetched — those are noted.

Every surfaced field is also self-describing in its rendered block or tool description (a header legend or inline label), so the model knows what each value means and what to do with it — not just that it exists.


Music / reference

Industry news feeds (utils/industry_feeds.py, #reporting-oversight)

The trade press read directly via public RSS (no key), so a first-party number is in hand rather than chased through a lagging web-confirm. Five sources in FEEDS: Pollstar news, IQ Magazine, Music Business Worldwide, Billboard music-news, and Billboard Chart Beat (added 2026-09-28: Billboard’s own chart announcements, which the music-news feed does not carry). One namespace-driven parser. The cache holds a feed 90 seconds, because the newsroom’s feed watcher reads every 2 minutes (see docs/ARCHITECTURE.md).

RIAA Gold & Platinum (utils/riaa.py, #2839)

RIAA’s own certification database, read directly (no key). Four views: the recent feed (every award newest first, 30 rows a page, paged by the site’s own “show more” JSON call), one award’s timeline (its full level history with dates), the top-tallies list (the 100 highest-certified albums, certified units in millions, 4 pages) and the awards-by-artist search (one row per credited act: certified units, Gold / Platinum / Multi-Platinum / Diamond counts, SOLO titles only per the site’s own note). The site holds no singles tally list (probed 2026-09-02: every sub-tab value answers the album list).

Pollstar charts (utils/pollstar_charts.py, #reporting-oversight)

Pollstar’s public chart JSON (data.pollstar.com, no key), the live-music numbers Billboard does not publish. Anon free-tier only — no auth is ever sent, and only the preview rows the endpoint returns are taken (top ~10 concert / ~50 Mediabase). One parser for every chart in CHARTS.

Billboard (utils/billboard.py)

The first-party chart read (server-rendered HTML, no key). One parser covers every chart in CHARTS; past weeks are addressable by date.

Chart-history store (utils/chart_history.py, cogs/chart_history.py, 2026-09-20)

Not a new upstream: the DURABLE copy of three integrations below plus Nora Music’s X posts, written to Postgres so the history is read from the store, never re-fetched.

HITS Daily Double (utils/hits.py)

The industry first-week projections + Mediabase radio feeds, read off HITS’ public Sanity GROQ API (no key; access reasoning in docs/BILLBOARD_SOURCES.md §5b).

kworb YouTube video charts (utils/kworb.py, #youtube-chart-art)

The “Today’s Most Viewed Music Videos on YouTube” pages — the global top-500 and the anglophone top-300, read as one merged list. Their read is CARD ART, not model context: nothing on these rows reaches Claude, so the ledger question here is only whether we drop something the art path could use. (The other kworb pages — Spotify, Apple, iTunes, radio — have no entry yet; that is a documented gap, not a decision.)

kworb HTTP response headers (utils/kworb.py, #3191)

Every kworb page serves ETag and Last-Modified, and until 2026-09-12 the client read neither. They are not chart content, so they reach no model and no card; they are read for two jobs.

kworb Spotify track page (utils/kworb.py, 2026-09-19)

The per-track chart-history page (spotify/track/<id>.html), read by track_streams. Since 2026-09 it renders TWO sections in the same table shape: <div class="weekly"> (one row per chart week) and then <div class="daily"> (one row per day); each has a Total and a Peak row and a per-country column set.

kworb Spotify daily chart page date (utils/kworb.py, research round 4, 2026-09-21)

The Spotify daily chart pages (spotify/country/us_daily.html, global_daily.html) name the chart DATE in their title (Spotify Daily Chart - United States - 2026/09/19): the day the streams are FOR. kworb publishes a day one or two days after it (read 2026-09-21, the US page showed 2026/09/19), so the read’s date is not the streams’ date.

kworb Spotify artist / listeners rankings (utils/kworb.py, owner ask 2026-08-10)

The all-artists ranking (spotify/artists.html) and the monthly-listeners ranking (spotify/listeners.html), consumed by /ask’s spotify_chart standing line, the watch sweep, and now the artist leaderboard boards (utils/chart_boards.py).

kworb US weekly YouTube chart (utils/kworb.py, owner ask 2026-08-31)

The per-country YouTube chart for the US (youtube/insights/us.html), behind the us_video_top + us_video_new boards. A DIFFERENT page and metric from the most-viewed video charts above: those are worldwide, daily, and by VIEWS; this is US, weekly, and by STREAMS. It is the only YouTube page we read that says whether a title is new, and since 2026-09-14 it is the only YouTube page that feeds a BOARD at all. The page size flips: kworb has served it at 100 rows and at 20, and the boards need it whole, so kworb._US_VIDEO_CHART_SIZES holds both measured sizes and a read matching neither darkens both boards with a chart_fetch shape_change.

kworb YouTube artist ranking (utils/kworb.py, owner ask 2026-08-30)

The all-time most-viewed YOUTUBE ARTISTS leaderboard (youtube/archive.html — kworb files it under “archive”, but its own nav calls it Artists). ~1,732 acts. It backs the market card’s ranking strip on a YouTube market, so the scale is read on the same platform the card’s number is on. No consumer since 2026-09-29: the artist standing strip it backed was removed from every card (owner steer). The client read and parser stay, but nothing calls them.

kworb per-artist songs page (utils/kworb.py, album card, owner ask 2026-08-16)

The per-artist songs page (spotify/artist/<id>_songs.html, parse_artist_songs): every tracked song of one artist with its cumulative total + latest daily. It backs the fresh-drop album-streams card (chart_boards.album_streams_board) — match an album’s iTunes tracklist to these rows and each track’s daily comes free.

kworb per-artist albums page (utils/kworb.py, owner ask 2026-08-28)

The per-artist ALBUMS page (spotify/artist/<id>_albums.html, parse_artist_albums): every tracked ALBUM of one artist with its cumulative total + latest daily – the same shape as the songs page above, one level up. It is the PREFERRED source for an album’s lifetime streams (utils/stream_totals.py), because it is Spotify’s own published per-release figure: the older tracklist sum only approximates it (a track’s own total carries the plays it took as a single, so summing over-counts a release that has one), and iTunes does not return every act’s albums at all, so for some acts the sum could never resolve (measured 2026-08-28: an iTunes album search for Bad Bunny returns his singles and features and none of his albums).

Two feeds kworb publishes that the client did not read, added 2026-09-06 after a survey of every path the site exposes.

Genius (utils/genius.py)

KTT2 forum (utils/ktt2.py, #ktt2)

The music FORUM, read as discourse (search_ktt2) and as a breaking signal (ktt2_breaking). No key and no API: every section page server-renders its Apollo GraphQL cache into __NEXT_DATA__, so one GET yields typed JSON entities. Access reasoning + the three pre-ship checks are in the module docstring.

Apple Music / iTunes (utils/apple_music.py)

Deezer (utils/deezer.py)

Deezer — new-releases board — REMOVED (owner steer 2026-09-05, “never use Deezer”)

Apple Music Top Albums RSS (utils/apple_releases.py) — the releases board’s source

iTunes Store Top Albums RSS (utils/itunes_charts.py, #3469) — the iTunes ALBUMS chart

Spotify Web API — new releases (utils/spotify.py) — DORMANT (#2447)

Songstats (utils/songstats.py, utils/streaming_stats.py, utils/milestones.py)

The RapidAPI aggregator — /tracks/stats, /artists/stats, /tracks/historic_stats. Each payload’s stats array carries ~20 platform blocks (spotify, apple_music, amazon, deezer, youtube, tiktok, instagram, shazam, soundcloud, tidal, itunes, beatport, traxsource, tracklist, facebook, twitter, songkick, bandsintown, radio).

MusicBrainz (utils/musicbrainz.py)

Wikidata (utils/wikidata.py)

Wikipedia chart parser (utils/chart_data.py)

Wikipedia reader (utils/reference.py)

Steam store art (utils/steam_art.py, #2741)

The catalog rung for a VIDEO GAME subject, the way TMDB is film’s and Deezer is music’s. Public store endpoints, no key. It exists because the Wikipedia lead-image rung structurally cannot serve a game: pageimages EXCLUDES non-free files, and a game article leads with its copyrighted box art (measured 2026-08-31 — every game tested returns no image, while free-licensed subjects return one).


Sports

SGO — SportsGameOdds (utils/markets.py, utils/sportsdata/board.py)

Polymarket (utils/markets.py)

Kalshi (utils/markets.py)

The Odds API (utils/the_odds_api.py)

The SGO-down resilience backstop (#725 epic). Lower tier — 20,000 credits/month, cost = #markets × #regions per /odds or /event-odds call; /sports + /events are FREE, /scores = 1 (2 with daysFrom). x-requests-remaining is captured on every call (client.credits_remaining) so the budget is observable.

Endpoints covered (all live v4 endpoints):

API-Sports (utils/api_sports.py)

Highlightly (utils/highlightly.py)

ESPN scoreboard (utils/espn.py, #1113)

The FREE, keyless last-resort score/finals backstop for the SGO-only INDIVIDUAL sports — tennis (SGO is its only source; no API-Sports tennis host) and MMA/UFC (API-Sports /fights is its only settler; The Odds API /scores is 0-completed for MMA). Public site.api.espn.com site API, no auth. Wired as EspnScoreProvider, the actual LAST provider in the hub chain (after API-Sports + SGO; the metered Odds API scores provider left the chain in #3552), so first-provider-wins dedup makes it a pure fill-in. Since #3569 it is the free WORKHORSE (owner steer 2026-09-21: “the free source is the workhorse; the metered feeds are accuracy references, not dependencies”): the capability matrix’s _ESPN_SETTLES names every team sport plus tennis, bot.py derives the always-on gate from it, so ESPN’s live, upcoming and finals reads run for all of them whatever the primaries say. API-Sports and SGO rows still win the dedup when they exist. Only MMA keeps the api_sports.degraded gate (API-Sports /fights is its settler). Polite by construction (“a free service, not a resource”): a day slate with no game in progress and no kickoff inside 10 minutes is cached 5 minutes (_ESPN_TEAM_TTL_QUIET), a slate with something moving 30 seconds (_ESPN_TEAM_TTL), and concurrent live/upcoming/finals reads share one in-flight fetch per (league, day). Measured before the flip: ~800 ESPN calls an hour, almost all from the bet slate refresh; the TTLs are what keep the always-on flip at or below that.

Endpoints covered:

The STATS reads (#2268). The same host publishes, free and keyless, the team-league data the sports stats boards are built on. Three more endpoints, each its own market_fetch source so one dead read is separable from a healthy one:


Media / AI / enrichment

Box Office Mojo (utils/box_office.py, the cinema desk’s box-office source)

A scrape, not an API, so “the fields” are the chart page’s COLUMNS. BOM prints three chart layouts and they do not share a column ORDER, which is why parse_chart reads the <th> row and maps each column by NAME (#2266). A header-less snippet falls back to the original positional heuristics.

page columns
weekend /weekend/<YYYY>W<NN>/ Rank, LW, Release, Gross, %± LW, Theaters, Change, Average, Total Gross, Weeks, Distributor
daily /date/<YYYY-MM-DD>/ TD, YD, Release, Daily, %± YD, %± LW, Theaters, Avg, To Date, Days, Distributor
year /year/<YYYY>/ Rank, Release, Genre, Budget, Running Time, Gross, Theaters, Total Gross, Release Date, Distributor

TMDB (utils/tmdb.py, the cinema desk’s release + enrichment source)

OMDb (utils/omdb.py)

The cinema desk’s critical scores. One endpoint (omdbapi.com/?t=<title>), read by the scores story lane, the critics scoreboard board lane, and utils/reference_movie.py.

Rotten Tomatoes (utils/rotten_tomatoes.py)

The critics scoreboard’s RT column. There is no API — this reads RT’s public pages, which is an owner decision (2026-08-11) and not an agreement. It sits outside their terms of service and can break with no warning, so every path fails open and the rt_fetch event carries the fill rate.

Perplexity (utils/perplexity.py)

xAI Grok x_search (utils/grok_search.py, #1390)

Real-time X (Twitter) grounding via Grok’s server-side x_search Agent tool (Responses API on api.x.ai), the sibling of Perplexity. Wired into: the on-demand /ask search_x tool (alongside the ScrapeCreators social tools) AND a live X-pulse grounding block on the scheduled trend/take surfaces — discourse (one source among Perplexity/markets), music (a fresh-hit signal: “what tracks/artists are people posting about on X”; she names the real track, never pastes an X link into the links-only channel), recap (what X is saying about what the room’s on; same not quiet gate as Perplexity, #880), the trending-clip reactions (clip-scoped: what people are saying around the pool’s trending moments, so a “did you see this” lands on WHY it’s buzzing – context for her take, never a “who said what” relay), and the market surfaces market_drop + market_alert (the topic-scoped X CROWD READ via the shared market_topic_pulse helper, so the take plays crowd-vs-reality the bare %s can’t – an alert can even explain WHY a market just moved; sentiment only, the compose still quotes numbers ONLY from the market blob). Two lanes are carved OUT of that wiring, all structurally (withhold the input, do not only ban the output). The rule: a card that REPORTS does not get the crowd read (owner steer, 2026-08-31). The crowd read serves the crowd-vs-reality angle, and that angle needs a live question to be about. Three kinds of card have none:

  1. A FORECAST of a number (#2745) – market_alert’s forecast_move plus its closing range / ladder ending_soon branches, AND market_drop’s forecast / forecast_rank projection angles. The post IS a projected number. This block is what turned a clean figure into “…forecast at ~7.1M, well below the streaming dominance the timeline keeps citing” (live 2026-08-31).
  2. A settled or locked RESULT – market_alert’s settled_out report, and market_drop’s >= 95% report lock. The question is answered.
  3. A confirmed LEADER – market_alert’s closing rank-leader read, and market_drop’s dominant frontrunner. The card states who is out front; the crowd does not get a vote on that.

The gate flag (forecast_lane, which drives the self-gate’s opinion-clause hard fail) stays scoped to (1) alone, because its wording is forecast-specific. Which INPUT the compose gets and which GATE rubric applies are two separate questions. The withheld input is the WHOLE fix, and that is measured, not assumed. An added “do not editorialize” instruction shipped alongside it at first and was then cut: on the real blob, n=8 per arm, crowd-read-out scored 0/8 opinion tails with the rule and 0/8 without, while crowd-read-in scored 3/8 without. Either lever alone fixes it, so the instruction was pure prompt cost – and the enumerated-ban shape #1140 warns about. tests/test_market_alert.py ratchets it: a test now FAILS if the ban is added back.

Every carve-out is known BEFORE the fetch, so the Grok call (and Perplexity, on the market_drop path) is skipped outright, not paid for and discarded. SELECTIVE by design (pricier than Perplexity, so only where live X sentiment is the point — not blanket-wired). Default model grok-4.3 (GROK_SEARCH_MODEL-overridable): a live dry-run (2026-07) measured it at ~3 x_search calls / ~8-11k input tokens / ~$0.03 per call with 7-9 real X citations (~5-8× Perplexity), vs the frontier grok-4.5 which drives the agentic loop ~14 tool calls / ~275k tokens / ~$0.3-1.2 for the SAME all-X citation quality — a 10-20× cost delta with no quality gain, so the cheap model is the default (max_tool_calls does NOT rein grok-4.5 in — model choice is the only cost lever). grok-4.6 was measured the same way (2026-09, #model-upgrade) and stays OFF the default for the same reason: on the same two grounding queries it ran 4-6× slower (48-69s vs 12s), read 1.5-4× the input tokens (12.0k/42.2k vs 8.0k/10.2k), and returned the SAME or FEWER citations (5 vs 7 on one, 4 vs 4 on the other) — on top of a rate 1.6×/2.4× above 4.3. Citation yield is the only thing a grounding block is bought for, so the newer model is a straight loss here. Gated on GROK_API_KEY; fail-open. Emits grok_search.

xAI API reference (verified live against api.x.ai, 2026-07-15). The single place we keep the confirmed xAI facts, so a future session (image failover #1389, text backend #1391, video #1392) doesn’t re-derive them:

X / fxtwitter (utils/x_fetch.py)

twitterapi.io (utils/twitterio.py, the content curator’s X source, #curator)

X API v2 write (utils/x_poster.py, the @tootsiesbar crosspost target)

ScrapeCreators (utils/scrapecreators.py, utils/social_search.py, #1272/#1273)

One SCRAPECREATORS_API_KEY, one guarded REST client, TWO roles:

yt-dlp video (utils/video_fetch.py)

Giphy (utils/gifs.py)

OpenAI image (utils/image_gen.py)

xAI / Grok Imagine (utils/xai_image.py, #1388/#1389)

The FAILOVER image backend: ImageClient retries a generate/edit on Grok when OpenAI either safety-rejects the prompt (a content refusal on real named public figures — permanent for that prompt) or is unavailable (429/5xx/timeout/ connection drop — an availability failure Grok covers on a separate quota), mirroring the text provider-failover (#973). It does NOT fail over on a permanent 4xx (malformed request / bad key) or a decode miss. Gated on GROK_API_KEY; fail-open. image_generated events carry provider="xai" (vs "openai") so each backend’s health is separable; the failover emits provider_fallback (requested=openai, fell_back_to=xai, reason=safety_reject|transient).

OpenAI embeddings (utils/embeddings.py)

ElevenLabs STT (utils/stt.py)

ElevenLabs TTS (utils/tts.py)

Giphy usage (utils/gifs.py, #753)

OpenAI usage (utils/embeddings.py + utils/image_gen.py, #753)

Perplexity / Genius usage (no machine-readable budget, #753)


Ops (model-facing via tools)

GitHub (utils/github.py, check_fixes)

Railway (utils/railway.py, check_deploys)


Known real gaps (documented, not yet closed)

These are the only answer-valuable fields a user could plausibly want that we still drop. Each is a conscious low-priority deferral, recorded here so it’s explicit:

Integration Field Why deferred
Deezer track-level release date / explicit Not in Deezer’s track-search response shape (an upstream limitation; the album catalog path carries the date).

(API-Sports venue + referee, previously listed here, are now surfaced via GAME CONTEXT; only the Deezer upstream limitation remains.)

Everything else is either surfaced or an intentional omission listed above.

Source-watcher cache speeds (2026-09-28)

The music desk’s source watcher (docs/ARCHITECTURE.md) must see a new edition within minutes, so four reads got a faster re-read. Field coverage is unchanged.