The stack
From podcast audio to graded predictions, every claim traceable to the second it was said.
Take Makers listens to sports podcasts, finds every prediction the hosts make, figures out who made it, and grades it when the outcome is known. Each take links to the exact moment in the audio, the verbatim quote, the evidence for the speaker, and the data or source that settled it.
17
episodes processed
849
predictions extracted
366,041
words transcribed
100%
of quotes timed to the audio
97%
speaker recognition on unseen episodes
164
takes settled · 59% came true
Live from the site’s data. Speaker accuracy is measured on episodes held out of the voiceprints.
How it works
Eight stages, each a small, testable step that can resume where it left off.
- 1
Ingest
Each supported show's RSS feed is read and compared against what's on the site. One click (or one command) imports an episode.
RSS · Apple Podcasts lookup API
- 2
Own the audio
We download our own copy of the MP3, so every timestamp refers to audio we control, not a third party's copy with different ad breaks.
ffmpeg · range-request streaming
- 3
Transcribe
Word-level speech-to-text runs locally on the Apple Silicon GPU. Every word gets a start and end time; nothing leaves the machine.
Whisper large-v3-turbo via MLX
- 4
Recognize speakers
A 192-number voiceprint of every two seconds of speech is matched against known hosts' voiceprints, learned from human-confirmed quotes. Unclear stretches stay unlabeled instead of guessed.
SpeechBrain ECAPA · PyTorch (MPS)
- 5
Extract
Claude reads the full, speaker-labeled transcript and drafts the summary, every prediction (betting picks included) and the line guesses, then does a second pass specifically hunting for what the first pass missed.
Claude Opus 5.5 · streaming Messages API
- 6
Verify
Every quote must appear verbatim on its transcript line, every speaker must be on the episode, every team must exist. Anything that fails is dropped, never 'fixed up'.
Deterministic validators · checker
- 7
Grade
Single-game takes settle from saved ESPN box scores and closing lines; season-long takes from saved standings; everything else from web research that must cite a page the search actually returned.
ESPN APIs · Claude + web search
- 8
Review
Imports land as drafts. A review queue surfaces takes likely to be wrong (voice disagrees with the credited speaker, low confidence, hard to grade) with the exact editor that fixes each.
Human in the loop
What it’s built with
Application
Next.js 16
App Router, React Server Components, static generation of every episode, take, person and team page
React 19
Client components only where interaction lives: the audio player, editors, live progress
Tailwind CSS 4
Utility-first styling, no component library
Data
Data as code
Episodes, takes, sources and standings are plain JS data files in git: every change is a diff, every edit is reviewable and reversible
Golden-file tests
The grade of every take and pick is snapshotted; any change to logic, data or saved games that flips a grade fails CI until accepted on purpose
Saved sports data
Final ESPN scores, box scores, closing lines and standings are snapshotted, so grades are reproducible and never depend on a live API
AI
Claude Opus 5.5
Extraction (two passes), drafting takes from highlighted text, grading plans, and web-researched grades
Structured outputs
Every AI answer is JSON validated against the transcript and the data model; streaming keeps long extractions reliable
Web search with provenance
Grades must cite a source URL that the search actually returned; anything else is rejected and stays open
Speech
MLX Whisper
Local, GPU-accelerated transcription with word timestamps (a 2-hour episode in about 6 minutes)
SpeechBrain + scikit-learn
Speaker voiceprints and recognition, with held-out evaluation
Exact audio sync
Each transcript word is aligned to our audio, so 'play this quote' starts on the first word
Operations
Node.js scripts
Import, align, snapshot, assign ids and check: each step idempotent and resumable
Python (venv)
Speech models, isolated from the web app
GitHub Actions
Tests, data checks and a production build on every push
Why you can trust a grade
Verbatim or nothing
A quote that isn't word-for-word on its transcript line is rejected. The site can always play the exact moment it was said.
No invented sources
A web-researched grade is only accepted if its source is one of the pages the search really returned.
Voice over guesswork
Speaker attribution comes from voice recognition, validated on episodes the model never saw. Where the voice disagrees with the record, a person decides.
Deterministic first
Anything a box score or the standings can settle is settled by code, not a model. AI handles only what needs judgment.
Reproducible grades
Every grade comes from saved data and is pinned by a golden-file test, so a grade can't silently change.
Human in the loop
Imports are drafts until reviewed. Every editor writes a reviewable diff to git, never a hidden database row.
By the numbers
- Transcript lines
- 16,966
- Speaker-attributed takes
- 847
- Bets among the takes
- 111
- Line guesses tracked
- 79
- People in the dataset
- 19
- Voiceprints
- 9 (10 min of confirmed speech)
- Cited sources
- 10
- Grades pinned by tests
- 849
See it in action: shows and their episodes, any episode’s transcript, or a single take with its audio, speaker evidence and grade.