claudetools

Author	SHA1	Message	Date
Mike Swanson	c760e430c0	radio: bumper detection in diarizer + full archive download script Adds a transcript-driven bumper filter to the diarization pipeline. When a transcript segment matches qa_extractor's promo/bumper signatures, the overlapping audio windows are labeled BUMPER and the WavLM cosine match is skipped. Prevents music/promo from being matched against speaker profiles (the failure mode Mike caught in 2018-s10e18 @ 09:20-10:05). Code changes: - src/voice_profiler.py: identify_speakers() takes optional skip_ranges parameter; windows whose midpoint falls in a skip range get labeled "[bumper]" and skip cosine match - src/diarizer.py: diarize() takes optional transcript_path; pre-computes bumper time ranges via qa_extractor._is_promo_or_bumper, passes to identify_speakers; adds BUMPER speaker label - benchmark.py: passes transcript_path to diarize() Aggregate impact across 9-episode test set: Tara attribution: 4880s -> 3680s (-1200s / -25%) Q&A pairs: 17 -> 19 (+2) (bumper-flagged segments had been disrupting conversation detection in 2017-s9e30 and 2018-s10e18) CALLER total: 1320s -> 1190s (bumpers previously labeled CALLER moved) Per-episode bumpers caught: 1-8, total ~165 bumper segments across set Remaining Tara false positives are real callers acoustically similar to Tara (Christopher in 2018, Kay in 2012, William and Charles in 2015) and guest Clay in 2015-s7e19 — those need profile rebuild + Clay profile, not bumper filtering. Adds download_full_archive.py — resumable mirror-style downloader that walks IX server's /home/gurushow/public_html/archive/{year}/ and copies all MP3s to archive-data/episodes/. Run is in progress (~589 files, ~10-15GB). Used to source clean profile windows for the remaining co-hosts (Tara rebuild, Clay, Tony, Rob, Randall, producers). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-04-27 16:17:50 -07:00
Mike Swanson	e9ac607500	radio show: co-host voice profile, Q&A extraction fixes, archive index - Build Tom (co-host) voice profile (44 embeddings, 0.698 similarity to Mike) - diarizer.py: add CO-HOST speaker label for cohost-role profiles - voice_profiler.py: emit "Cohost: <name>" label for cohost role - qa_extractor.py: overlap resolution at load time (midpoint boundary split), 4s CALLER-preference threshold, turn-based caller-intro lookback (2 HOST turns), _preceded_by_caller_intro() helper, _PHONE_GREETING pattern, 751-1041 + "we'll get your problem solved" promo signatures - benchmark.py: use src.transcriber.transcribe with batch_size=16 - add index_test_episodes.py and build_cohost_profile.py scripts - add .gitignore (exclude episodes, transcripts, *.db, .venv) - session log: 2026-04-27-qa-extraction-cohost-indexing.md Result: 2016-s8e43 drops from 12 false-positive Q&A pairs to 2 real caller pairs. archive.db: 6 episodes, 762 segments, 10 Q&A pairs, FTS5 search verified. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-27 14:41:04 -07:00
Mike Swanson	79abef9dc9	radio: diarization pipeline fixes, benchmark setup, test episode set - Fix voice_profiler threshold bug (HOST label overwrote Unknown unconditionally) - Audio preload optimization: single ffmpeg per episode, 149.5x realtime on 5070 Ti - WavLM threshold raised to 0.85 (Mike 0.90-0.99, callers 0.46-0.83) - Promo/bumper filter: weighted signature scoring, 42->27 clean Q&A pairs - Text-only Q&A fallback for episodes with no CALLER diarization labels - TRANSFORMERS_OFFLINE=1 to skip HuggingFace freshness checks - Add diarize_2018.py for targeted re-run + FTS5 rebuild - Add benchmark.py + BENCH_SETUP.md for GURU-BEAST-ROG (RTX 4090) comparison - Commit 9-episode training diarization.json outputs - Session log: 2026-04-27-diarization-pipeline.md Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-27 13:20:40 -07:00
Mike Swanson	a1e0442d8b	Add radio show audio processor and post-show workflow - Audio processor CLI tool with 6-stage pipeline: transcribe (faster-whisper GPU), diarize (pyannote), detect segments (multi-signal classifier), remove commercials, split segments, analyze content (Ollama) - Post-show workflow doc for episode posts, forum threads, deep-dive blog posts - Training plan for using 579-episode archive for voice profiles and commercial detection - Successful test: 45min episode transcribed in 2:37 on RTX 5070 Ti - Sample transcript output from S7E30 (March 2015) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-03-21 11:51:59 -07:00

4 Commits