# Build Log

Tell Me a Story records the bedtime stories my daughter and I make up. It turns them into transcripts that know who's speaking, and tries to get her invented names down right. Everything runs on my own computer; no family audio leaves the house.

*Claude Code writes these from each session's work; Saurin Choksi scans and gives notes. Entries he has hand-edited stay his.*

**Five words you'll meet:**

- **Session**: one recorded bedtime storytelling. Press record, tell stories, stop.
- **Universe**: the story-world a tale belongs to (the Mahabharata, Thomas & Friends). Our own invented stories live under a universe called Home. Earlier entries say "world"; same thing.
- **Fixer**: one of three tools that notice likely name mistakes and ask about them. Fixers only ask; they never change anything.
- **Watching mode** (the shadow phase): the current training-wheels period. The machine only asks and records my answers. Not one transcript word changes, not even from my own yes taps, until one future, deliberate apply step.
- **Name Book**: the one place my rulings on names live, one file per universe. Nothing writes in it except my own taps.

---

## 2026-07-03 — Our own stories get the name treatment too

The name fixer used to sit out any night it couldn't recognize the story world. That meant it skipped our own invented stories entirely, and those are exactly the nights with our characters in them. Now it re-listens to those nights with our registered characters in Whisper's ear, the same way a Mahabharata night gets the Mahabharata cast, and whatever it hears becomes a question for me to judge like any other.

## 2026-07-02 — The name fixer now watches instead of writing

I rebooted the whole name-fixing side into a watching mode: the fixer still listens back to the audio around every suspect name, but every opinion is now a question on a review screen, grouped by why the machine believed it. To start clean I rebuilt every recording back to exactly what Whisper first heard, and nothing touches a transcript until a whole group of questions has earned trust; then everything applies at once, in one deliberate step.

## 2026-07-02 — One book for every name ruling

My decisions about names used to live in three different files that three different parts of the system read. Now there is one Name Book, one file per story world, holding my yeses ("`Bushma` means Bhishma") and my nos ("`Jammus` is her own engine, leave it alone"). A no stops the question from being asked, but its spots stay viewable in a folded section with a button to bring the question back if a night ever proves me wrong.

## 2026-07-02 — The new name-fixer listens instead of guessing

Instead of guessing what a garbled name should be from its spelling, I built a new name fixer that gives Whisper another listen with the story world's own cast in its ear. This fixes the vast majority of misspellings and garbled transcriptions. On a Mahabharata night it turned `Bushma` into "Bhishma" right through the war. The few the new name fixer is not sure about wait on a review screen.

## 2026-06-30 — Re-timed every word so I can click one and hear it

Whisper's word timings were often off in a consistent way: it marked a word as starting slightly before the actual speech. The opening "Okay," was tagged at 1.70 seconds, but my daughter doesn't say it until 2.06. Using a technique called forced alignment, I fixed it so now clicking a word in the reviewer lands right on it, instead of a beat early.

## 2026-06-24 — Overrule wrong name flags

Sometimes a detector flags a name it shouldn't, so I added a way to set it straight. It marked "Jammus" (a train engine my daughter and I made up) as a misspelling of "James" (a real Thomas & Friends character), because the two sound almost identical. They aren't the same character. We invented one that just sounds like a real one. Now I can correct that false flag right on the monitor page, and it stays corrected through every future scan.

## 2026-06-19 — Moved the story tools onto Qwen 3.5 4B

I wanted the story work to run on one small model instead of juggling multiple. So I tested whether Qwen 3.5 4B could take over what Gemma 4 E4B was doing, and it won: Qwen splits a recording into its stories more cleanly and catches more misspelled names. The splitter, the world-recognition step, and the canon name-checker all moved onto Qwen.

## 2026-06-18 — Built the per-story canon name-checker

For each story, a model lists the characters of whatever world the story is set in, then flags any the transcriber spelled wrong. In a Mahabharata story the model knows the cast includes Bhishma, so when the transcript says `Bishma`, the model catches the misspelling.

## 2026-06-17 — Got world-recognition working

Built a "is this a known story-world, or just made up?" classifier. For each story the classifier makes one of three calls. The story might be from a world the model knows (Thomas & Friends, the Mahabharata, Steven Universe). The story might be wholly improvised. Or the story might be a mix: a known world with additional invented characters. Knowing the world is what lets the name-checker tell a misspelled real character from an invented new one.

## 2026-06-15 — Split each recording into its separate stories

For numerous features, including character name checking, I need to segment each recording into the individual stories it holds. Simple cues help determine where the breaks are: a long silence (three seconds or more), or telltale lines in the transcript like "once upon a time," "the end," or "start the story." Then a small model reads through and confirms the real breaks. A last pass fixes over-splitting: if a long pause in the middle of one story (someone got up for milk) made it look like two, it stitches the halves back together when they share the same characters and plot.

## 2026-06-12 — Added a second name detector

It catches a made-up character whose name gets written several different ways in one story, so I can still follow that character through the telling. Plain text-matching catches most. The ones it misses are invented names that are also ordinary words. Those go to a small local model that reads a few lines and decides whether the word is being used as a name.

## 2026-06-11 — Took the name Monitor live

The first detector now runs over every recording and shows its catches on a new Monitor screen. It flags the family's own names when Whisper mangles them, like my daughter's name heard as an ordinary word. The screen lists every flag across all sessions, each with a play button to hear the exact moment. It only points at the transcript, never changes it.

## 2026-06-01 — Tested the family-name detector on unseen recordings

Built a family-roster-name detector and checked it on two recordings it had never seen, against answers I marked by ear. It caught every mis-transcribed name a fresh listen turned up, including spellings of my daughter's name it had never been shown.

## 2026-05-28 — Built sweeps to find the failures I'd miss by hand

Three read-only passes that hunt for failures: speech the transcriber dropped entirely, words stamped at the wrong moment, and stretches where it got stuck repeating one little word (`Right. Right. Right.`).

## 2026-03-03 — Built cross-session speaker ID

Taught the pipeline to recognize who's speaking and carry that across recordings. From a single voice sample of each of us, it correctly picked out me and my daughter on a session it had never heard. When it's unsure (a silly voice, someone far from the mic) it says "suggested" rather than force a confident match.

## 2026-02-18 — Ran the pipeline end to end from one command

Now I can drop a recording in a folder, run one script, and get back a transcript split by who's speaking, with a timestamp on every word.

## 2026-02-07 — Got name correction working on the real recording

Ran the whole pipeline on the real Mahabharata recording for the first time, name correction included, and it fixed the mangled Sanskrit names cleanly. What made it work was handing the model the whole transcript at once instead of one line at a time. Line by line, with no surrounding context, it had invented wrong fixes like turning "dad" into a Sanskrit name. I'm giving the model all the names manually so this is not the final fix but it points me in the direction.

## 2026-02-07 — Flag Whisper's made-up words for review

Whisper sometimes writes words that were never said. The validation tool flags two kinds for me to review: a word Whisper itself barely believes, where its confidence on that word is near zero, and a word sitting where the speaker detector heard only silence. These just point me at the suspect spots; they don't change the transcript.

## 2026-02-03 — Save the raw transcript, compute the rest on demand

The pipeline keeps three files: the raw transcript with the words as Whisper first heard them, the speaker data, and a combined transcript with corrections enriched on top. The raw one is never overwritten, so if a correction later turns out wrong, I can always see what Whisper actually heard, like the original `fondos` before it was fixed to "Pandavas."

## 2026-01-27 — Built the validation tool

Used Claude Code to whip up a "validation tool" where I can listen to a recording and view the transcript running alongside. I mark where the transcript has errors, segment by segment.

## 2026-01-24 — Settled how to handle Whisper's hallucinations

If someone really spoke but Whisper couldn't make out the words, we keep the spot and mark it unintelligible. If Whisper invented words out of pure silence, aka hallucinations, we delete them.

## 2026-01-22 — The first transcript that reads as a conversation

Combined Whisper's output with speaker detection and got the artifact that feels like a real transcript.

> SPEAKER_01: Dad, why do the Fondos and the Goros want to be king?
> SPEAKER_00: Uh-huh. Well, so the oldest brother of the Goros, his name was, do you remember?
> SPEAKER_01: Durioden.

## 2026-01-21 — Child speech needs the large Whisper model

Compared Whisper's small, medium, large, and large turbo models on the same stretch of audio. For catching a quiet, young child's speech I need Whisper large (not turbo). The other models frequently output nothing even though I can hear her talking.

## 2026-01-20 — Day one!

I hooked up Whisper and ran my first transcription: a ~5 min bedtime convo with my daughter about the Mahabharata. It mostly worked. Looking at the transcript felt a lil magical, but I noticed mangled Sanskrit names throughout. Like "Yudhishthira" became `you this there`. "Pandavas" became `fondos`. 🤔
