The stack

From podcast audio to graded predictions, every claim traceable to the second it was said.

Take Makers listens to sports podcasts, finds every prediction the hosts make, figures out who made it, and grades it when the outcome is known. Each take links to the exact moment in the audio, the verbatim quote, the evidence for the speaker, and the data or source that settled it.

17

episodes processed

849

predictions extracted

366,041

words transcribed

100%

of quotes timed to the audio

97%

speaker recognition on unseen episodes

164

takes settled · 59% came true

Live from the site’s data. Speaker accuracy is measured on episodes held out of the voiceprints.

How it works

Eight stages, each a small, testable step that can resume where it left off.

  1. 1

    Ingest

    Each supported show's RSS feed is read and compared against what's on the site. One click (or one command) imports an episode.

    RSS · Apple Podcasts lookup API

  2. 2

    Own the audio

    We download our own copy of the MP3, so every timestamp refers to audio we control, not a third party's copy with different ad breaks.

    ffmpeg · range-request streaming

  3. 3

    Transcribe

    Word-level speech-to-text runs locally on the Apple Silicon GPU. Every word gets a start and end time; nothing leaves the machine.

    Whisper large-v3-turbo via MLX

  4. 4

    Recognize speakers

    A 192-number voiceprint of every two seconds of speech is matched against known hosts' voiceprints, learned from human-confirmed quotes. Unclear stretches stay unlabeled instead of guessed.

    SpeechBrain ECAPA · PyTorch (MPS)

  5. 5

    Extract

    Claude reads the full, speaker-labeled transcript and drafts the summary, every prediction (betting picks included) and the line guesses, then does a second pass specifically hunting for what the first pass missed.

    Claude Opus 5.5 · streaming Messages API

  6. 6

    Verify

    Every quote must appear verbatim on its transcript line, every speaker must be on the episode, every team must exist. Anything that fails is dropped, never 'fixed up'.

    Deterministic validators · checker

  7. 7

    Grade

    Single-game takes settle from saved ESPN box scores and closing lines; season-long takes from saved standings; everything else from web research that must cite a page the search actually returned.

    ESPN APIs · Claude + web search

  8. 8

    Review

    Imports land as drafts. A review queue surfaces takes likely to be wrong (voice disagrees with the credited speaker, low confidence, hard to grade) with the exact editor that fixes each.

    Human in the loop

What it’s built with

Application

Next.js 16

App Router, React Server Components, static generation of every episode, take, person and team page

React 19

Client components only where interaction lives: the audio player, editors, live progress

Tailwind CSS 4

Utility-first styling, no component library

Data

Data as code

Episodes, takes, sources and standings are plain JS data files in git: every change is a diff, every edit is reviewable and reversible

Golden-file tests

The grade of every take and pick is snapshotted; any change to logic, data or saved games that flips a grade fails CI until accepted on purpose

Saved sports data

Final ESPN scores, box scores, closing lines and standings are snapshotted, so grades are reproducible and never depend on a live API

AI

Claude Opus 5.5

Extraction (two passes), drafting takes from highlighted text, grading plans, and web-researched grades

Structured outputs

Every AI answer is JSON validated against the transcript and the data model; streaming keeps long extractions reliable

Web search with provenance

Grades must cite a source URL that the search actually returned; anything else is rejected and stays open

Speech

MLX Whisper

Local, GPU-accelerated transcription with word timestamps (a 2-hour episode in about 6 minutes)

SpeechBrain + scikit-learn

Speaker voiceprints and recognition, with held-out evaluation

Exact audio sync

Each transcript word is aligned to our audio, so 'play this quote' starts on the first word

Operations

Node.js scripts

Import, align, snapshot, assign ids and check: each step idempotent and resumable

Python (venv)

Speech models, isolated from the web app

GitHub Actions

Tests, data checks and a production build on every push

Why you can trust a grade

Verbatim or nothing

A quote that isn't word-for-word on its transcript line is rejected. The site can always play the exact moment it was said.

No invented sources

A web-researched grade is only accepted if its source is one of the pages the search really returned.

Voice over guesswork

Speaker attribution comes from voice recognition, validated on episodes the model never saw. Where the voice disagrees with the record, a person decides.

Deterministic first

Anything a box score or the standings can settle is settled by code, not a model. AI handles only what needs judgment.

Reproducible grades

Every grade comes from saved data and is pinned by a golden-file test, so a grade can't silently change.

Human in the loop

Imports are drafts until reviewed. Every editor writes a reviewable diff to git, never a hidden database row.

By the numbers

Transcript lines
16,966
Speaker-attributed takes
847
Bets among the takes
111
Line guesses tracked
79
People in the dataset
19
Voiceprints
9 (10 min of confirmed speech)
Cited sources
10
Grades pinned by tests
849

See it in action: shows and their episodes, any episode’s transcript, or a single take with its audio, speaker evidence and grade.