tootsies

a discord bot for the tootsies server. ask, recap, discuss, ship features by typing.


Project maintained by mejasonmejason Hosted on GitHub Pages — Theme by mattgraham

Upstream integration field coverage

What every external API hands us vs. what we surface to the model. This is the “are we dropping anything” ledger: each integration lists the fields we extract and the fields we deliberately omit (with the reason). The rule is either 100% of the answer-valuable fields, or an explicit entry here saying what’s missing and why — so an omission is a documented decision, never a silent gap.

Scope note: “surface to the model” means the field reaches Claude’s context (a rendered block or a tool result), not telemetry. The constitution’s data-minimization rule governs telemetry + her outbound replies, NOT the in-context input we feed the model, so enriching context is in-spec. A handful of fields are withheld from output by the HARD RULES (no personal-info disclosure) even when fetched — those are noted.

Every surfaced field is also self-describing in its rendered block or tool description (a header legend or inline label), so the model knows what each value means and what to do with it — not just that it exists.


Music / reference

Billboard (utils/billboard.py)

The first-party chart read (server-rendered HTML, no key). One parser covers every chart in CHARTS; past weeks are addressable by date.

HITS Daily Double (utils/hits.py)

The industry first-week projections + Mediabase radio feeds, read off HITS’ public Sanity GROQ API (no key; access reasoning in docs/BILLBOARD_SOURCES.md §5b).

kworb YouTube video charts (utils/kworb.py, #youtube-chart-art)

The “Today’s Most Viewed Music Videos on YouTube” pages — the global top-500 and the anglophone top-300, read as one merged list. Their read is CARD ART, not model context: nothing on these rows reaches Claude, so the ledger question here is only whether we drop something the art path could use. (The other kworb pages — Spotify, Apple, iTunes, radio — have no entry yet; that is a documented gap, not a decision.)

kworb Spotify artist / listeners rankings (utils/kworb.py, owner ask 2026-08-10)

The all-artists ranking (spotify/artists.html) and the monthly-listeners ranking (spotify/listeners.html), consumed by /ask’s spotify_chart standing line, the watch sweep, and now the artist leaderboard boards (utils/chart_boards.py).

Genius (utils/genius.py)

KTT2 forum (utils/ktt2.py, #ktt2)

The music FORUM, read as discourse (search_ktt2) and as a breaking signal (ktt2_breaking). No key and no API: every section page server-renders its Apollo GraphQL cache into __NEXT_DATA__, so one GET yields typed JSON entities. Access reasoning + the three pre-ship checks are in the module docstring.

Apple Music / iTunes (utils/apple_music.py)

Deezer (utils/deezer.py)

Songstats (utils/songstats.py, utils/streaming_stats.py, utils/milestones.py)

The RapidAPI aggregator — /tracks/stats, /artists/stats, /tracks/historic_stats. Each payload’s stats array carries ~20 platform blocks (spotify, apple_music, amazon, deezer, youtube, tiktok, instagram, shazam, soundcloud, tidal, itunes, beatport, traxsource, tracklist, facebook, twitter, songkick, bandsintown, radio).

MusicBrainz (utils/musicbrainz.py)

Wikidata (utils/wikidata.py)

Wikipedia chart parser (utils/chart_data.py)

Wikipedia reader (utils/reference.py)


Sports

SGO — SportsGameOdds (utils/markets.py, utils/sportsdata/board.py)

Polymarket (utils/markets.py)

Kalshi (utils/markets.py)

The Odds API (utils/the_odds_api.py)

The SGO-down resilience backstop (#725 epic). Lower tier — 20,000 credits/month, cost = #markets × #regions per /odds or /event-odds call; /sports + /events are FREE, /scores = 1 (2 with daysFrom). x-requests-remaining is captured on every call (client.credits_remaining) so the budget is observable.

Endpoints covered (all live v4 endpoints):

API-Sports (utils/api_sports.py)

Highlightly (utils/highlightly.py)

ESPN scoreboard (utils/espn.py, #1113)

The FREE, keyless last-resort score/finals backstop for the SGO-only INDIVIDUAL sports — tennis (SGO is its only source; no API-Sports tennis host) and MMA/UFC (API-Sports /fights is its only settler; The Odds API /scores is 0-completed for MMA). Public site.api.espn.com site API, no auth. Wired as EspnScoreProvider, the actual LAST provider in the hub chain (after API-Sports + SGO + The Odds API), so first-provider-wins dedup makes it a pure fill-in. Fetched per-sport ONLY while that sport’s primary is degraded (tennis ← sgo.degraded, mma ← api_sports.degraded) — free, so the gate is consistency + politeness to a no-SLA host, not cost.

Endpoints covered:

The STATS reads (#2268). The same host publishes, free and keyless, the team-league data the sports stats boards are built on. Four more endpoints, each its own market_fetch source so one dead read is separable from a healthy one:


Media / AI / enrichment

Box Office Mojo (utils/box_office.py, the cinema desk’s box-office source)

A scrape, not an API, so “the fields” are the chart page’s COLUMNS. BOM prints three chart layouts and they do not share a column ORDER, which is why parse_chart reads the <th> row and maps each column by NAME (#2266). A header-less snippet falls back to the original positional heuristics.

page columns
weekend /weekend/<YYYY>W<NN>/ Rank, LW, Release, Gross, %± LW, Theaters, Change, Average, Total Gross, Weeks, Distributor
daily /date/<YYYY-MM-DD>/ TD, YD, Release, Daily, %± YD, %± LW, Theaters, Avg, To Date, Days, Distributor
year /year/<YYYY>/ Rank, Release, Genre, Budget, Running Time, Gross, Theaters, Total Gross, Release Date, Distributor

TMDB (utils/tmdb.py, the cinema desk’s release + enrichment source)

OMDb (utils/omdb.py)

The cinema desk’s critical scores. One endpoint (omdbapi.com/?t=<title>), read by the scores story lane, the critics scoreboard board lane, and utils/reference_movie.py.

Rotten Tomatoes (utils/rotten_tomatoes.py)

The critics scoreboard’s RT column. There is no API — this reads RT’s public pages, which is an owner decision (2026-08-11) and not an agreement. It sits outside their terms of service and can break with no warning, so every path fails open and the rt_fetch event carries the fill rate.

Perplexity (utils/perplexity.py)

xAI Grok x_search (utils/grok_search.py, #1390)

Real-time X (Twitter) grounding via Grok’s server-side x_search Agent tool (Responses API on api.x.ai), the sibling of Perplexity. Wired into: the on-demand /ask search_x tool (alongside the ScrapeCreators social tools) AND a live X-pulse grounding block on the scheduled trend/take surfaces — discourse (one source among Perplexity/markets), music (a fresh-hit signal: “what tracks/artists are people posting about on X”; she names the real track, never pastes an X link into the links-only channel), recap (what X is saying about what the room’s on; same not quiet gate as Perplexity, #880), the trending-clip reactions (clip-scoped: what people are saying around the pool’s trending moments, so a “did you see this” lands on WHY it’s buzzing – context for her take, never a “who said what” relay), and the market surfaces market_drop + market_alert (the topic-scoped X CROWD READ via the shared market_topic_pulse helper, so the take plays crowd-vs-reality the bare %s can’t – an alert can even explain WHY a market just moved; sentiment only, the compose still quotes numbers ONLY from the market blob). SELECTIVE by design (pricier than Perplexity, so only where live X sentiment is the point — not blanket-wired). Default model grok-4.3 (GROK_SEARCH_MODEL-overridable): a live dry-run (2026-07) measured it at ~3 x_search calls / ~8-11k input tokens / ~$0.03 per call with 7-9 real X citations (~5-8× Perplexity), vs the frontier grok-4.5 which drives the agentic loop ~14 tool calls / ~275k tokens / ~$0.3-1.2 for the SAME all-X citation quality — a 10-20× cost delta with no quality gain, so the cheap model is the default (max_tool_calls does NOT rein grok-4.5 in — model choice is the only cost lever). Gated on GROK_API_KEY; fail-open. Emits grok_search.

xAI API reference (verified live against api.x.ai, 2026-07-15). The single place we keep the confirmed xAI facts, so a future session (image failover #1389, text backend #1391, video #1392) doesn’t re-derive them:

X / fxtwitter (utils/x_fetch.py)

twitterapi.io (utils/twitterio.py, the content curator’s X source, #curator)

X API v2 write (utils/x_poster.py, the @tootsiesbar crosspost target)

ScrapeCreators (utils/scrapecreators.py, utils/social_search.py, #1272/#1273)

One SCRAPECREATORS_API_KEY, one guarded REST client, TWO roles:

yt-dlp video (utils/video_fetch.py)

Giphy (utils/gifs.py)

OpenAI image (utils/image_gen.py)

xAI / Grok Imagine (utils/xai_image.py, #1388/#1389)

The FAILOVER image backend: ImageClient retries a generate/edit on Grok when OpenAI either safety-rejects the prompt (a content refusal on real named public figures — permanent for that prompt) or is unavailable (429/5xx/timeout/ connection drop — an availability failure Grok covers on a separate quota), mirroring the text provider-failover (#973). It does NOT fail over on a permanent 4xx (malformed request / bad key) or a decode miss. Gated on GROK_API_KEY; fail-open. image_generated events carry provider="xai" (vs "openai") so each backend’s health is separable; the failover emits provider_fallback (requested=openai, fell_back_to=xai, reason=safety_reject|transient).

OpenAI embeddings (utils/embeddings.py)

ElevenLabs STT (utils/stt.py)

ElevenLabs TTS (utils/tts.py)

Giphy usage (utils/gifs.py, #753)

OpenAI usage (utils/embeddings.py + utils/image_gen.py, #753)

Perplexity / Genius usage (no machine-readable budget, #753)


Ops (model-facing via tools)

GitHub (utils/github.py, check_fixes)

Railway (utils/railway.py, check_deploys)


Known real gaps (documented, not yet closed)

These are the only answer-valuable fields a user could plausibly want that we still drop. Each is a conscious low-priority deferral, recorded here so it’s explicit:

Integration Field Why deferred
Deezer track-level release date / explicit Not in Deezer’s track-search response shape (an upstream limitation; the album catalog path carries the date).

(API-Sports venue + referee, previously listed here, are now surfaced via GAME CONTEXT; only the Deezer upstream limitation remains.)

Everything else is either surfaced or an intentional omission listed above.