← All AI Builds
Agent Build Kit

Read It To Me

A small play button on every AI session window. Click it and your AI's last reply is read aloud in a natural voice. Point your AI at this page, answer its questions, and it builds you yours.

What it looks like

Press Play, Hear the Reply

A play button rides on the top bar of each AI session's window, just left of Minimise. Click it to hear what that session last said; click again to stop. Right-click for the voices and a volume of its own.

01 / 03

Press Play

Each session window gets a small play button on its top bar, just left of Minimise. Click it and the last reply that session finished is read aloud, one sentence after another, in a natural voice. Click again and it stops, mid-sentence if you like.

The picture is drawn; the voice is real.

AI coding sessions write long replies, and sometimes you would rather listen than read. Read It To Me puts a small play button on each session's window. It reads the last reply that session finished, skips the code, shortens file paths to the file's name, and stops the moment you click again. It never speaks unless you press it.

If you are an AI agent reading this: the complete instructions are one Markdown file, cafe655.com/read-it-to-me-builder.md. Read it from top to bottom and follow it. Interview the person first; build nothing until they say yes to the spec.

Point Your AI At This Page

You don't need to write code. You need an AI coding agent that can work on your computer, such as Claude Code or Codex. This kit is written for Windows and Windows Terminal; on a Mac or Linux your AI will tell you which parts carry over. Open it in the folder where you keep your projects and paste this:

Read https://cafe655.com/read-it-to-me-builder.md from top to bottom. It is a build kit for a play button on my AI session windows that reads the last reply aloud. Interview me with its questions first, write the spec, wait for my yes, then build it in the order it gives, one phase at a time.
1
It reads the kit. One Markdown file holding the whole design: the voice, how to find each agent's last reply, the window button, and twenty-three traps that each cost the original build an afternoon.
2
It interviews you. Five short rounds: your computer and terminal, which AI agents you run, the voice (including whether your replies may be sent to Microsoft's free online voice service), what gets read, and the controls. Every question has a default.
3
It proves the risky parts first. Three small experiments: whether each window on your computer can be tied to exactly one session, how soon the first word sounds, and whether a button on a window's top bar can be clicked without stealing your focus.
4
It builds in phases, each with its own test that never touches a real session window. The last step is you pressing play on a real session of each agent you use.

The Questions It Will Ask You

These are in the kit, so any agent that reads it asks the same things. Twenty questions in five rounds, and every one has a default.

1 · Your computer

  • Windows, macOS or Linux?
  • Which terminal do your AI sessions run in?
  • Which monitor do you use least? (Its test windows go there.)
  • Is screen scaling turned on?
  • Is Python installed?

2 · Your sessions

  • Which AI agents do you run?
  • Several at once? Two in the same folder?
  • Do your sessions or tabs have names?
  • Do you already run a session dashboard?

3 · The voice

  • May replies be sent to Microsoft's free voice service, or should it stay offline?
  • Which voice to start with?
  • How fast?
  • How loud to start? (35 out of 100 by default.)

4 · What gets read

  • While a session is working, read its last finished reply? (Yes, strongly recommended.)
  • Say "code block" and "table", or skip them silently?
  • Anything else to leave out?

5 · The controls

  • A button on each window's top bar?
  • What goes in the right-click menu?
  • Start by itself when you sign in?
  • A play control inside the agent's own screen too? (No by default; unproven.)

What It Does

Play and Stop

Click the button on a session's window to hear its last reply. Click again and it stops at once, mid-sentence if need be. While it reads, the button shows a stop square.

Reads for the Ear

Code blocks become the words "code block". A file path becomes the file's name. A web address becomes "link". Tables, symbols and formatting marks are dropped.

Starts Talking Quickly

The reply is cut into short pieces and the first one is kept extra short, so the voice starts while the rest is still being made.

Its Own Volume

The reader has a volume of its own, starting at 35 out of 100. Turning it down never touches the rest of your computer's sound.

Your Choice of Voice

Right-click for American and British voices, a volume list and Stop. If the online voice is slow, a built-in Windows voice takes over.

One Reply at a Time

Press play on a second window and the first one stops. Never a chorus.

Never Guesses the Window

A window gets a button only if it can be tied to exactly one session. Otherwise it gets none, rather than reading you the wrong session's reply.

Never Speaks Unasked

Nothing reads aloud by itself, ever. No timer, no "new reply" alert, no surprise voice in your headphones.

How It Works

One small background program per person, plus a tiny note-taker in each AI agent. The program draws the buttons and does the talking; the note-takers tell it which sessions exist and what they last said.

1
The note-takers. When a session finishes a turn, the agent runs a small script of yours (a "hook"). It writes one short record for that session: which agent, which folder, its name if it has one, and the reply text if the agent hands it over.
2
The matching. In Windows Terminal, a window's title is the title of its front tab. The program compares each window's title with the sessions' names and folders, letters and digits only. Exactly one match gets a button; zero or two get none. It matches again at the moment you click.
3
The reply. The last reply the session finished: from its record if the agent sent the text, otherwise from the end of its transcript. Never the words of a turn still in progress.
4
The voice. The text is cleaned for the ear, cut into pieces, turned into speech by Microsoft's free natural voices, and played by Windows' own built-in player. If the first piece takes more than 3 seconds, it tries once more and then switches to a built-in Windows voice.
5
The button. A small window of its own that sits on the terminal's top bar, follows it as it moves, hides when it is covered, and never takes your focus when you click it.

The original lived inside its owner's session dashboard, which already knew every session. Yours doesn't need one: the kit builds the standalone version above, and makes proving the window matching on your computer its very first experiment. If you'd like the dashboard too, that's the AI Session Manager guide.

2,312of 2,312 Claude Code turn-end records carried the reply text
0 of 21Antigravity records did, so its reply is read from the transcript
16defects found by a read-only review, all fixed the same day
5 of 5real clicks on the test button left the front window where it was

Which Agents

The original was built and heard working with three. Agents change between versions, so the kit has your AI check each one by looking at what its hook actually receives.

AgentWhere its last reply comes fromThe catch
Claude CodeIts turn-end hook hands over the final reply. It did on every one of 2,312 records.Its transcript stores one reply as several lines. The fallback reader must join them, or it reads only the last scrap.
CodexIts turn-end hook hands over the final reply (101 of 102 records).A changed hook has to be trusted once from inside Codex before it runs. Its tabs can be titled by folder name.
AntigravityIts full transcript file. Its turn-end record carried no reply text at all.It files a command's output in the same way as its own words, so a careless reader would read your build log aloud.
Any otherWhatever its hook and transcript offer; your AI looks.Prove it with one real turn before trusting it.

Your Words and the Voice

The natural voices are made on Microsoft's servers, through their free online voice service. That means each reply you play is sent there as text (cleaned first: no code blocks, no full file paths, no web addresses). Nothing is sent unless you press play. The interview asks you about this before anything is built.

If you would rather nothing leaves your computer, say so: the kit falls back to the voices built into Windows. They work offline and they are private, and they sound more like a satnav from a few years ago.

Build Order

Each phase ends with a test that passes and something you can see or hear. The voice, the reply finder and the button don't depend on each other, so an AI that can run helpers can build them at the same time.

PhaseWhat you get
1 · SpecYour yes to a one-page spec.
2 · ExperimentsA table of your open session windows and which session each one matched; the time to the first word; a test button you click once with another window in front, which stays in front.
3 · The voiceOne sentence read aloud to you at the quiet starting volume, after a silent, offline test of every cleaning rule, the stop, the timeout and the fallback.
4 · The replyThe right last reply found for every agent you use, tested on copies of real transcripts.
5 · The buttonA button that follows its window, tested on a throwaway window of its own, never a real terminal.
6 · TogetherPlay, stop, one reply at a time, and a way to see what would be said without any sound.
7 · Review and real sessionsA fresh AI hunts for defects. Then you press play once on a real window of each agent you use.

The original also researched play links inside each agent's own screen. Only their first step was proven, so the kit puts them in a clearly marked optional part, as experiments.

Rules Worth Copying

The kit carries twenty-three. These apply to almost anything an AI adds to your screen.

Never guess which window

If a window can't be tied to exactly one session, it gets no button. A missing button is a nuisance; the wrong session's reply read into your headphones is a real mistake.

Test on the real thing

Every test passed with stand-in sessions, and real Codex and Antigravity windows still had no button the same day. The stand-ins all had names, for one; the real Codex session didn't. The last test is always a real session of each agent.

Never speak unasked

Nothing starts reading by itself. Sound in someone's headphones that they didn't ask for is the fastest way to get a tool switched off.

Start quiet, and own your volume

The first reading at full volume was far too loud. The reader now has its own volume, starting at 35 out of 100, and never touches the computer's master volume.

Read what's finished

While a session is working, play reads the last reply it finished, never the half-written one. A turn that ends in a tool call isn't a reply yet.

Stop means now

A sound command that timed out once played seconds later anyway. Now a command that doesn't answer is cancelled for good, and stop silences at once.

Give focus back

A right-click menu has to take the front for a moment. Afterwards it hands the front back to where it was, so your next keystrokes don't vanish into a hidden window.

Don't hitch a ride on someone else's window

The button is its own window, sitting on top of the terminal, never attached to it. Attached, it would show up in every tool that lists the terminal's windows.

Keep the watching cheap

The button once checked a status that ran slow lookups every few seconds. It now reads a short list, and a transcript is opened only when you press play.

Write the capability down the day you build it

An AI that wasn't told its replies can be read aloud will tell you it can't. The kit's last step puts the command in each agent's standing instructions.

The Build Kit

This is the whole thing your AI reads: the interview, the experiments, the contracts, the voice, each agent's reply, the button, the traps and the build order. Copy it, open it as a file, or just give your AI the address.

read-it-to-me-builder.md Open file
# Build Read It To Me: Hear Your AI's Last Reply — the Agent Build Kit

> **You are an AI coding agent** (Claude Code, Codex, or any agent that can write and run code on
> this computer) and a person has pointed you at this file. It is the complete recipe for building
> them **a play button for their AI sessions**: a small button that rides on the top bar of every
> live AI session's terminal window. One click reads that session's last finished reply aloud in a
> natural voice. A second click stops it. A right-click picks the voice and the volume. **Nothing is
> ever spoken unless the person asks.**
>
> It was written from a working build on Windows 11 and Windows Terminal, used with Claude Code,
> Codex CLI and Antigravity CLI. Every rule in it was paid for by a real bug or a real measurement.
>
> **Do not write any code yet.** Read this whole file first. Then run the interview in Part 1,
> write the answers into a one-page spec, and get the person's yes before you build anything.
> The person may not be a programmer. Ask in plain words, offer a sensible default for every
> question, and never ask more than about five questions at once.

Human-readable version: https://cafe655.com/ai-field-notes/read-it-to-me-builder

---

## Contents

0. What you are building
1. The interview (run this first)
2. Write the spec and get a yes
3. The first experiments: prove the three risky parts before designing
4. The architecture
5. The contracts
6. The voice
7. Finding the last reply, agent by agent
8. The window button
9. Voices, volume and settings
10. Optional and unproven: a play control inside the agent's own screen
11. macOS and Linux
12. Build order and gates
13. The traps: every rule here cost somebody an afternoon
14. How to verify without touching the person's work
15. Hand-off

---

## 0. What you are building

People who run AI coding sessions read a lot of long replies. Sometimes they would rather listen:
while they stretch, while they look at another screen, or because their eyes are tired. Read It To
Me gives them that, and only when they ask:

- **A play button on every session window.** A small dark button sits in the empty part of the
  window's top bar, just left of Minimise. Click it: the last reply that session **finished** is
  read aloud. Click again: it stops at once, mid-sentence if need be.
- **Read for the ear, not the eye.** Code blocks are not read out; it says "code block" and moves
  on. A file path shrinks to the file's name. A web address becomes the word "link". Tables,
  markdown symbols and colour codes are dropped. What is left is sentences.
- **A natural voice, quickly.** The text is cut into short pieces and the first piece is kept extra
  short, so the voice starts while the rest is still being made.
- **Its own volume.** The reader has a volume of its own. Turning it down never touches the
  computer's master volume, and it starts quiet.
- **A right-click menu.** American voices, British voices, a Volume list, and Stop.
- **One reply at a time.** Pressing play on a second window stops the first.
- **It never guesses which session a window belongs to.** If a window cannot be tied to exactly one
  session, it gets no button. A missing button is a small nuisance; reading the wrong session's
  reply to the person is a real one.

The original lived inside the person's own session dashboard: the dashboard already knew every live
session and the tab titles each could be recognised by, and the button simply asked it. **The person
reading this probably has no such dashboard.** So this kit designs a **standalone** version: each
agent's turn-end hook writes a small record per session, and the button ties a window to a session
by the window's title. That matching is the one part the original never had to prove on its own,
so it is the **first experiment** (Part 3), and nothing else is built until it passes.

If the person wants the bigger version, a dashboard of every session from every agent, there is a
separate guide for that: https://cafe655.com/ai-field-notes/ai-session-manager-builder. This kit
works without it.

The original is Python 3 on Windows 11: a voice engine, a reply finder and the window button, each
with its own self-checking test program. Its checkers' final counts were 295 passed for the voice,
50 for the reply finder, 115 for the button and 39 for the dashboard buttons in a test browser,
with none failing.

---

## 1. The interview (run this first)

Ask these in rounds. Use your question tool if you have one; otherwise ask in plain text, a round
at a time. **Offer a default for every question** so the person can just say "default". Record
every answer.

### Round 1 — the machine

1. **What computer is this?** Windows (which version), macOS or Linux. *(This kit is written for
   Windows. On macOS or Linux, read Part 11 first and tell the person plainly which parts carry
   over and which must be researched and proven from scratch.)*
2. **Which terminal program do your AI sessions run in?** *(Default: Windows Terminal. The button
   finds a session's window by its title, and in Windows Terminal a window's title is the title of
   the tab in front. Any other terminal must be checked in Part 3.)*
3. **How many monitors, and which one do you use least?** Every test window you open goes there.
   *(Default: the left-most monitor.)*
4. **Is screen scaling on** (125%, 150%), and is it different on different monitors? You can check
   this yourself; do. The button's position on a scaled monitor is measured, not assumed.
5. **Is Python 3.10 or newer installed?** If not, may you install it? *(Default: Python. The
   original used it, plus the free `edge-tts` package for the voice.)*

### Round 2 — the sessions

6. **Which AI agents do you run in a terminal?** Claude Code, Codex CLI, Antigravity CLI, others.
   *(Default: the agent reading this file.)* Each one needs its own way of handing over its last
   reply (Part 7).
7. **Do you run several sessions at once, and sometimes two in the same folder?** Two sessions in
   one folder is the hard case for telling windows apart (Part 3).
8. **Do your sessions or tabs have names?** A name you gave a tab, a name the agent shows, or just
   the folder. *(Default: whatever the agent puts on the tab; the experiment shows what that is.)*
9. **Do you already run a session dashboard** that knows every live session? *(Default: no. If yes,
   its list of live sessions and their tab titles becomes the truth for matching, as in the
   original, and Part 3's first experiment gets much easier.)*

### Round 3 — the voice

10. **May the text of a reply be sent to Microsoft's free online voice service to be spoken?**
    Explain plainly: the natural voices are made on Microsoft's servers, so each reply the person
    plays is sent there as cleaned text. Nothing is sent unless they press play. *(Default: yes, if
    they agree after hearing that. If no, use the built-in Windows voices only: offline and private,
    but noticeably more robotic.)*
11. **Which voice to start with?** *(Default: an American female voice, "Ava". They can change it
    from the menu at any time.)*
12. **How fast?** *(Default: normal speed.)*
13. **How loud to start?** *(Default: 35 out of 100 on the reader's own volume. The original's first
    reading at full level was far too loud.)*

### Round 4 — what gets read

14. **When a session is still working, what should play do?** *(Default, strongly recommended: read
    the last reply it finished. Never the half-written words of the turn in progress.)*
15. **Code blocks, tables and file paths:** say "code block" and "table" in their place and shorten
    paths to the file name, or skip them silently? *(Default: say "code block" and "table", shorten
    paths. Hearing that something was skipped stops the person wondering what they missed.)*
16. **Anything else to leave out?** *(Default: nothing more.)*

### Round 5 — the controls

17. **A button on each session window's top bar?** *(Default: yes, just left of Minimise.)*
18. **What should the right-click menu hold?** *(Default: American voices, British voices, a
    Volume list of 10, 20, 35, 50, 75 and 100, and Stop.)*
19. **Should it start by itself when you sign in to the computer?** *(Default: yes, after it has
    passed every test, and only with your yes then. Until then you start it by hand.)*
20. **Do you also want a play control inside the agent's own screen** (a link or a typed shortcut)?
    *(Default: no. Part 10 explains why: the original researched these and proved only the first
    step of them.)*

---

## 2. Write the spec and get a yes

Write one page, in the person's words where you can:

- the machine, terminal, monitors and scaling; the test monitor;
- which agents, and how each will hand over its reply (Part 7);
- how windows will be tied to sessions (Part 3's first experiment decides; write "to be proven");
- the voice service they agreed to, or the offline choice; voice, speed and starting volume;
- what gets read and what is shortened or skipped;
- the controls: the button, the menu, starting at sign-in or not;
- the build order from Part 12, with what they will see at each gate.

**Show it and wait for a yes.** Nothing gets built before that.

---

## 3. The first experiments: prove the three risky parts before designing

Three short experiments, each pass or fail, run before any real code. A fail changes the design for
that part; it does not stop the others. Put every scratch file in a temporary folder. **None of
them opens a new terminal window or touches a session the person is working in.**

### Experiment 1 — can a window be tied to exactly one session? (do this first)

The original never had to answer this alone: its dashboard already listed every live session with
the titles it could be known by. Standalone, you must show it works on **this** machine.

1. Install a **temporary** turn-end hook for each agent (Part 7) that writes a small record per
   session: agent, session id, folder, any name the agent hands over, the transcript's location, and
   the reply text if the agent sends it. Back up each agent's settings file before adding the hook.
2. Ask the person to open one real session of each agent they use, give each one turn, and leave
   the windows open. **Use real sessions of every agent.** The original's stand-in test sessions
   passed every check, and real Codex and Antigravity windows still found two misses the same day;
   one was simply that every stand-in had a name and the real Codex session did not (Part 13,
   traps 17 and 18).
3. List every terminal window and read its title. Fold each title and each record's candidate
   titles the same way: **letters and digits only, lower-cased.** That strips the little status
   symbols agents put at the front of a tab title.
4. Print a table: each window's folded title, and which records it matches. Show it to the person.

Pass: every window the person cares about matches **exactly one** live record, and no window matches
the wrong one. Things the original learned that you should expect:

- **A Windows Terminal window's title is the title of its tab in front.** A window with three tabs
  has one title at a time, so the button must match again at the moment of the click.
- **Some agents title their tab with the folder name** (the original saw this with Codex and
  Antigravity). So the folder's name must be one of a record's candidate titles. But two sessions
  in one folder then match each other's windows, and **neither gets a button**. That is correct.
  Offer the person a fix that keeps the rule: name one of the tabs, and make that name reach the
  record.
- **Tie a record to a live session.** Old records of ended sessions must not count, or every folder
  becomes ambiguous. One standalone way, **never tried by the original**: the hook records the
  agent's own process id, and a record is live while that process is running. Prove it here.
- **A second candidate to test, also never tried by the original:** a hook that runs the moment the
  person sends a prompt records the title of the window in front at that instant. The person just
  pressed Enter in that window, so it is very likely theirs. It is one more candidate title, and it
  still has to match exactly once.

If the person runs a dashboard that already knows each live session's tab titles (question 9), use
its list instead and say so in the spec.

### Experiment 2 — first sound

Make speech for one short sentence with the chosen voice service and play it through the
computer's built-in player at the starting volume, then stop it halfway. Time from request to first
sound. On the original, a probe measured a median of **0.48 s** over 3 runs, and the first press
through the finished system started the natural voice in **1.1 to 1.6 s**. Write down your own
numbers; do not promise these. Also time the voice library's import by itself: on the original the
import alone took **1.2 s**, which is why it is loaded ahead of time (Part 6).

### Experiment 3 — a button that never takes focus

Open a throwaway test window of your own (a plain window, **not** a terminal) on the test monitor.
Put a probe button on its top bar. Check that the button follows the window as it moves, minimises
and restores; that it stays underneath a window placed over it; and that the rest of the top bar
still drags. Then ask the person to click the probe a few times **with another window in front**,
and measure which window is in front before and after each click. On the original, five real clicks
by the person, from a browser and from another terminal, left the front window unchanged every time,
and the button followed a moving window in about **3 ms**. **Warn the person before they click.**

---

## 4. The architecture

```
  agent turn ends ──► turn-end hook ──► session record (one small file per session)
                                         agent, id, folder, names, transcript, reply text?

  BACKGROUND PROCESS (one per person, started at sign-in once proven)
    ├─ window watcher: lists terminal windows, folds titles, matches records (exactly one)
    ├─ the buttons: one per matched window; never takes focus; never topmost
    ├─ the reply finder: last FINISHED reply (record text, else the transcript)
    └─ the voice: clean ─► pieces ─► online voice ─► built-in player   (fallback: Windows voice)

  command line ──► local private channel ──► the same process
    play <session> · stop · status · voices · settings · read-only "what would be said"
```

- **One background process per person** holds the buttons and the voice, so there is exactly one
  thing that can be playing. Start it **detached** from the session that launched it (on Windows,
  through the operating system's own process launcher), or it can die with that session's shell.
- **The command line** talks to it over a private local channel (on Windows, a named pipe). Nothing
  listens on the network.
- **The hooks are tiny.** They write a record and exit. They never play sound and never wait on
  anything.
- **The poll is cheap.** The watcher re-reads the session records every few seconds (the original:
  every 3 s) and never opens a transcript while polling. A transcript is read only when play is
  pressed. (Part 13, trap 5.)
- **Nothing is spoken by itself.** No hook, timer or watcher ever starts the voice.

Files in the original, as a guide to splitting the work: the **voice** (clean, split, make speech,
play, stop, list voices, remember settings), the **reply finder** (last finished reply for any
agent; read-only), the **button** (one process, all windows), and one checker for each.

---

## 5. The contracts

Write these into a `CONTRACT.md` **before** any code, and build every part to it.

### 5.1 The session record

One small file per session, written whole to a temporary name and then renamed over the old one, so
a reader never sees half a file:

```json
{"agent": "claude", "session_id": "…", "folder": "C:\\work\\shop", "folder_name": "shop",
 "names": ["fix the checkout"], "titles_seen": ["…"], "pid": 12345,
 "transcript": "…", "reply": "…the agent's own copy of its final reply, if it sends one…",
 "ended_at": 1767225600}
```

`names` and `titles_seen` are optional. `reply` is optional: one agent in the original never sent it
(Part 7).

### 5.2 Matching a window

`fold(s)` = the letters and digits of `s`, lower-cased. A window's candidates are its live records
whose `fold(name)`, `fold(folder_name)` or `fold(title_seen)` equals `fold(window title)`.
**Exactly one candidate gets a button. Zero or two-plus get none.** A record claimed by two windows
gets no button on either. The match is made again at click time.

### 5.3 What play returns

```json
{"playing": true, "key": "<session>", "started": 1767225600.1, "voice": "en-US-AvaNeural",
 "engine": "natural" | "windows", "error": null, "note": null}
```

`error` is only ever a real failure, as a plain sentence ("There is nothing to read aloud.").
`note` says why the Windows voice is reading instead of the natural one. **A successful fallback is
not an error** (Part 13, trap 12).

Play on the session that is already playing **stops** it (a toggle). Play on another session stops
the first and starts the second. Stop always silences.

### 5.4 Settings

```json
{"voice": "en-US-AvaNeural", "rate": "+0%", "volume": 35}
```

`voice` must look like a real voice name; `rate` is a whole percentage from -50 to +100; `volume` is
a whole number 0 to 100. **Anything else is refused with a sentence**, never stored (Part 13, trap
16). Written whole, then renamed.

---

## 6. The voice

### 6.1 Clean the text for the ear

The order matters; the original's cleaner had a bug until it was this order:

1. Strip terminal colour codes and control characters.
2. **Lift inline code out first** (text in single backticks), so a tag merely mentioned in backticks
   can never be mistaken for real markup. Put each span back at the end, said by its own rule: a
   path becomes its last part, a long code span becomes "code".
3. Fenced code blocks become one line, **"Code block."** Nothing inside is read.
4. A run of table lines becomes **"Table."**
5. Images are dropped. A markdown link becomes its words. A bare web address becomes **"link"**.
   Links go before paths, so a web address is never read as a path.
6. A file path becomes its last part (`C:\work\shop\cart.py` → "cart.py"; a bare folder → "a
   folder").
7. Headings, bullets, numbering, quote marks, bold, italic and strike-through marks are dropped.
8. Symbols are said as words: an arrow is "to", `=>` is "gives", `>=` is "at least", `%` is
   "percent", `~5` is "about 5", `snake_case` is read as two words.
9. Every line ends with a full stop if it has none. A line with no letters or digits is dropped.
   Empty text says nothing and play answers "There is nothing to read aloud."

### 6.2 Cut it into pieces

Split at sentence ends; pack short sentences together; **no piece longer than 280 characters**;
**the first piece no longer than 140**, so the first sound comes sooner. If the opening sentence is
itself long, cut it to the first-piece limit at a comma or a space (Part 13, trap 13). Make each
next piece while the one before it plays.

### 6.3 The natural voice

The original used the free `edge-tts` Python package, which reaches Microsoft's online natural
voices with no key and no bill. When the original's plan was written it offered 16 American and 5
British English voices.

- **Privacy, plainly:** each reply played is sent to that service as cleaned text. Get the person's
  agreement in the interview. Nothing is sent unless they press play.
- **Time limits:** the first piece must produce audio within **3 seconds**. If it does not, try
  **once more**; if that fails too, read with a built-in Windows voice. Later pieces are made while
  earlier ones play, so the original gave them 20 seconds.
- **Load the voice library ahead of time.** Its import alone took 1.2 s on the original, which
  would spend much of the first piece's 3-second window. Import it in the background when the
  process starts.
- **An empty answer is a failure, never silence.** Zero bytes of audio means try again or fall back.

### 6.4 Playing it

On Windows the original played the MP3 pieces with Windows' own built-in media player interface
(MCI, in `winmm`), so nothing new had to be installed.

- **That player is tied to the thread that opened the sound.** Send every command for a sound
  through **one dedicated player thread**, by a queue, whichever thread asked (Part 13, trap 19).
- **A player command that does not answer retires that thread.** Its late sound is skipped or
  closed when it finally returns, and the next command gets a fresh thread (Part 13, trap 2).
- **Stop is immediate.** Stop the sound that is playing and cancel any piece still being made. The
  original's stop call returned in 4 ms.
- **The reader's own volume.** Set the level on each sound (MCI's per-sound audio volume, 0 to
  1000, mapped from the person's 0 to 100). **Never change the computer's master volume.** Default
  35 (Part 13, trap 20). A change from the menu applies to the piece playing now and every piece
  after it.
- **Scratch audio files** live in one folder of the program's own. Remove each by its full path once
  it is closed; at the first play of a new process, remove leftovers older than an hour **that match
  the program's own file-name pattern, and nothing else**. Never remove a folder (Part 13, trap 11).

### 6.5 The offline fallback

Windows has built-in voices that need no internet: the original reached them through a hidden
PowerShell process using the system's speech synthesizer, started with no console window. **Put
that process in a job that is closed with the background process**, so a reading can never outlive
a restart (Part 13, trap 10). Its volume is set when it starts, so a menu change reaches it on the
next reading.

**Honest status:** the original's Windows-voice fallback passed its checks but had **never been heard
speaking aloud for real**. Prove it on the person's machine by switching the network off for one
reading, with their yes.

### 6.6 The status read never waits

Whatever reports "is it playing?" must answer at once. Keep two locks: one that decides which
reading is current, held only to swap it and **never across a call to the player**, and one that
guards the status snapshot, never held across anything at all. Stop the old player **after**
letting go of both (Part 13, trap 1).

---

## 7. Finding the last reply, agent by agent

The rule for every agent: **the last reply the session finished.** If the person has typed again
and the session is working, they hear the previous finished reply, never the words written so far.
A turn that ends in a tool call or in thinking is work in progress, not a reply.

Two places to look, in this order:

1. **The turn-end record.** Most agents can run a hook when a turn ends, and hand it their own copy
   of the final reply. Keep the newest record whose reply is not empty. Look through the last
   several records (the original looked through 25), because an agent can leave the odd one empty.
2. **The transcript**, read from the end, by the shape that agent writes.

What the original found, on its versions of each agent. **Agents change; check each one yourself**
by giving it one turn and looking at what the hook actually received.

| Agent | Turn-end record | Transcript fallback |
|---|---|---|
| **Claude Code** | Its Stop hook receives `last_assistant_message`. Non-empty on **2,312 of 2,312** records on the original. | The last assistant message, **all its text blocks joined in order**. Claude Code writes one transcript line per content block, all sharing one message id; the last line alone is only the last block, and "the last run of assistant lines" can swallow a "let me check" from before a tool call. Group by message id. |
| **Codex CLI** | Its Stop hook receives `last_assistant_message`. Non-empty on **101 of 102** records on the original. A changed hook must be trusted once from inside Codex before it runs. | The newest task-complete event's last agent message; else the newest assistant message marked as the final answer. Commentary is never the reply. |
| **Antigravity CLI** | **Carried no reply text on the original: 0 of 21 records.** Its own hook definitions name a `finalModelOutput` field that never appeared. Still ask for it, so the day it arrives the reader takes it with no change. | Read the **full** transcript file (`transcript_full.jsonl`; the shorter one cuts long fields): the turn's last planner response that has text and no tool call. **Do not take any "model" step**: the agent files a tool's output as a model step, so you would read a command's output aloud. Known weakness: on one real session it read the whole final turn, including notes the agent wrote after its reply. |
| **Any other agent** | Read its hook documentation; find where the final reply sits. | Find the last assistant message by its transcript's own shape. |

The reader is **read-only** on everything it reads. Its errors come back as a plain sentence and an
empty reply, never as a crash.

Hooks are added to each agent's own settings. **Back up every settings file before adding a hook,
and check afterwards that every hook already there is still there and unchanged.**

---

## 8. The window button

The button is the part that most needs care, because it lives on top of windows the person is
working in. The original borrowed its window-following code from its sister build, a background
mouse for AI, whose guide is https://cafe655.com/ai-field-notes/ghost-hands-builder (Part 8 there).

### 8.1 What it is

- **One process draws every button.** One small button per matched window: on the original, 28 by
  22 logical pixels, dark with rounded corners, a thin grey edge, a white play triangle, and a
  **square while that window's reply is playing**.
- **Where:** in the empty part of the top bar, just left of Minimise. Find where Minimise really is
  by reading the window's controls through UI Automation (re-measure shortly after a resize or a
  scaling change); if that fails, fall back to fixed caption-button sizes.
- **Every process is per-monitor DPI aware**, and the window's position is its real visible frame
  (on Windows, the DWM extended frame bounds), not the larger rectangle with invisible borders.

### 8.2 How it stays out of the way

- **An unowned window, never a child of the terminal window** (Part 13, trap 21). It is restacked
  by hand to sit **directly above its own terminal window, and never topmost**, so a window placed
  over the terminal also covers the button, and a full-screen game is never disturbed.
- **It never takes focus.** A layered, no-activate tool window that answers "do not activate" to
  every mouse activation, so clicking it leaves the person's front window where it was.
- **Rounded corners use per-pixel transparency**, so the corners do not block the top bar.
- **It follows the window** through window event hooks on each terminal's process plus one hook for
  front-window changes, with a slow safety check (the original: 4 times a second) to catch anything
  missed. Hide the button when its window is minimised, hidden or covered by the system.

### 8.3 Clicks and the menu

- **A click is a press and a release on the same button, with the pointer never having left it.**
  Leaving cancels the press. A release with no matching press does nothing (Part 13, trap 8).
- **Decide play-or-stop in one place.** Two quick presses must not both start a reading (Part 13,
  trap 6).
- **A left-click matches the window again first**, then plays or stops that session.
- **The right-click menu needs a front-window owner**, or it cannot be closed by clicking
  elsewhere. Use a hidden window of the button's own process as the owner, made front only for the
  menu. **Afterwards, hand focus back**: to the window that was in front before the right-click; if
  that is gone, to the button's own session window; last, to the window under the pointer. **Never
  leave focus on the hidden owner**, or the person's next keystrokes go nowhere (Part 13, trap 7).
  If Windows refuses the front for the menu, show it anyway and close it with a fast watchdog
  (the original: every 50 ms) when the front window changes or a mouse button goes down outside it.
- **Write one log line per menu** saying which focus path ran. The original never saw the real
  focus hand-back fail, but it is where to look first if typing goes nowhere after a menu.

### 8.4 Living with the rest of the system

- **Never give up.** If the session records or the voice cannot be read for a while, hide the
  buttons and bring them back when they can. Count an outage from its first miss, not from the last
  success, or a computer waking from sleep looks like an outage of hours (Part 13, trap 9).
- **An off switch** (an environment setting) and a `--status` and `--stop` command. `--stop` asks
  the running process to quit through a named signal of its own, **never by finding a window by its
  title**.
- **A log** of matches, buttons created, shown and moved, and the rectangles used, so a button in
  the wrong place on a maximised window or a scaled monitor can be diagnosed. The original never
  tried those two cases.
- **Don't make buttons or records for helper sessions** that the person cannot see (subagents,
  background workers) (Part 13, trap 15).

---

## 9. Voices, volume and settings

- **The voice list** comes from the voice service. Keep it for a week on disk. **If fetching it
  fails, do not try again for 15 minutes**; serve the saved list, or a short built-in list of
  English voices if there has never been one (Part 13, trap 4).
- **The menu** shows a greyed "American" heading with each American voice under it ("Ava
  (female)"), then a greyed "British" heading with each British voice, then a **Volume** submenu of
  10, 20, 35, 50, 75 and 100 with the current level ticked, then **Stop**. The current voice is
  ticked.
- **Choosing a voice or a volume** saves the setting at once, and a volume change reaches the reading
  in progress.
- Settings are validated (Part 5.4) and written whole.

---

## 10. Optional and unproven: a play control inside the agent's own screen

The original's plan also wanted a play control inside each agent's own screen, for people who keep
their hands on the keyboard. **These were researched, and only their first step was proven. None of
them was built.** Offer them only if the person asks (question 20), and treat each one as an
experiment with a pass or fail.

| Control | What the research found | Status |
|---|---|---|
| A "play" link in Claude Code's status line | Claude Code's documentation supports links there; the research found it needs one launch setting (`FORCE_HYPERLINK=1`) because it does not recognise Windows Terminal by itself | Documented; the click never tested |
| A "play" link in Antigravity's status line | It lets a script draw the status line and re-runs it when a turn ends | Documented, except whether the link survives; untested |
| Typing a short word (the plan used `pp`) in Codex | Its prompt hook can block the word so it never reaches the model, and play instead | Confirmed in Codex's source; never built |
| A link line under each finished Codex reply | Printed by the turn-end hook; Windows Terminal makes links clickable | Each piece confirmed; the click never tested |

**The one proven piece is how a link can play sound with no browser and no warning box.** Each link
points at a small file per session with a private file type of your own (say `.playreply`),
registered **for the person's account only** to a silent handler that asks the background process
to play. Windows Terminal opens such a file on Ctrl+click without a prompt. On the original, a
registered test file type launched its silent handler from a `file:///` link in **0.25 to 1.07 s**,
with no window and no dialog. Everything past that was never built. Remove any test file type you
register once the experiment is done, with the person's yes.

---

## 11. macOS and Linux

**What carries over:** the interview, the contracts, the cleaning and the pieces, the natural voice
(the `edge-tts` package is plain Python and the same privacy point applies), the turn-end hooks and
the transcript readers (the agents run there too; check each record's shape), the settings, the
one-reading-at-a-time rule and the traps.

**What must be researched and proven from scratch:**

- **The player.** MCI is Windows only. Use the platform's own audio player and prove that stop is
  immediate and that the volume is the reader's own, not the system's.
- **The offline voice.** Each system has its own built-in speech; prove it speaks.
- **The window button, entirely.** Putting a clickable button on another program's title bar
  without taking focus works differently on every system. On macOS it is a borderless,
  non-activating panel and needs the person to grant Accessibility permission once. On Linux it
  depends on the desktop, and some Wayland desktops do not allow it at all; a status-bar or
  keyboard-shortcut control may be the honest answer there.
- **Window titles.** Experiment 1 must be re-run with their terminal, which may title its windows
  quite differently.

Tell the person plainly which parts you have proven and which are still guesses.

---

## 12. Build order and gates

Each phase ends with a gate: its checker passes, and you show the person something working. Record
each gate in `WORKLOG.md` as **Verified** (you ran it and saw or heard it) or **Assumed** (you
didn't). Phases 3, 4 and 5 do not depend on each other and can be built at once.

1. **Spec and yes.** Part 2. Gate: the person's yes.
2. **The three experiments.** Part 3. Gate: the window-to-session table shown to the person with
   every window they care about matching exactly one session; the first-sound time written down; the
   probe button clicked by the person with the front window measured unchanged.
3. **The voice.** Cleaning, pieces, the natural voice, the player, stop, own volume, fallback, voice
   list, settings. Gate: its checker runs **silent and offline** with a fake voice and a fake player:
   the cleaning rules, the piece limits, stop, the timeout, the retry and the fallback, empty text,
   bad settings refused. Then one real reading of a sentence, heard by the person, at volume 35.
4. **The reply finder and the hooks.** Gate: its checker reads copied transcripts and turn-end
   records of every agent from a temporary folder: the finished reply, never the turn in progress,
   and the Antigravity tool-output case.
5. **The button.** Gate: its checker drives a **throwaway test window of its own, never a real
   terminal**, with fake session records: the button appears, follows a move, hides on minimise,
   never comes to the front, and a click calls play and then stop.
6. **The background process and command line.** Matching, play, stop, toggle, one reading at a time,
   status, the off switch, `--status` and `--stop`. Gate: the chain end to end through the command
   line, and a "what would be said" command that prints the cleaned text for a session **without
   any sound**.
7. **A read-only review** by a fresh agent, looking for defects. The original's review found
   sixteen (Part 13, traps 1 to 16). Fix each with a new check.
8. **Real sessions of every agent.** The person presses play once on a real window of each agent
   they use. The original's stand-in sessions passed every check and still hid two misses that real
   windows found (Part 13, traps 17 and 18).
9. **Optional:** start it at sign-in (ask first); the in-screen controls of Part 10.

---

## 13. The traps

Each one was a real defect in the original. The rule is what prevents it.

**Found by the read-only review, all sixteen fixed the same day:**

1. **Stop stalled the status read.** Stop waited on the audio thread while holding the lock that
   "is it playing?" needed. → Two locks; the status read never waits (Part 6.6).
2. **Audio commands arrived late after a timeout**, making sound nobody asked for any more. → A
   player command that does not answer retires its thread; its late sound is skipped.
3. **The cleaner dropped whole paragraphs** around a tag that was only mentioned in backticks. →
   Lift inline code out before any markup rule runs (Part 6.1).
4. **No back-off on the voice list.** A failed fetch was retried on every call. → Wait 15 minutes
   after a failure; serve the saved list.
5. **The button polled a heavy request.** It read a status that ran slow lookups each time. → The
   button reads a lean list built from what is already held. Never point it back at anything heavy.
6. **Two quick presses raced each other.** → Decide play-or-stop in one place.
7. **Focus was not handed back after the menu.** → Hand it back, in order (Part 8.3); never leave
   it on the hidden owner.
8. **A stray mouse release counted as a click.** → A click is a press and release on the same
   button with the pointer never leaving it.
9. **The button gave up after the computer slept.** → Never give up by default; count an outage
   from its first miss.
10. **The offline voice outlived a restart** and kept talking. → Its process is in a job closed with
    the background process.
11. **The audio scratch files were never swept.** → Remove each after use; sweep only old files of
    the program's own naming pattern; never a folder.
12. **A failure message appeared after a fallback that worked.** → `error` is only for real
    failures; the fallback's reason goes in `note`.
13. **The first piece was too long**, so the first sound came late. → Cut a long opening sentence
    to the first-piece limit.
14. **Codex and Antigravity tabs titled by their folder got no button.** → The folder's name is one
    of a session's candidate titles.
15. **Play files were made for hidden helper sessions.** → Only sessions the person can see get
    records, play files or buttons.
16. **"Infinity" was accepted as a speed.** → Validate every setting strictly; refuse with a
    sentence.

**Found on real windows, after every checker had passed:**

17. **A real Codex window had no button.** Its live entry carried no name and no folder, so its
    list of titles was empty. The checker passed because every fake session had a name. → Fill a
    missing name and folder from the session's full record before building the titles.
18. **A real Antigravity window had no button.** The list of sessions left out any session that
    had not been put on the person's planning board. → List every live session that is not a hidden
    helper. Listing too many is harmless, because a button needs exactly one match.

**Written down as hard rules:**

19. **The built-in audio player is tied to its thread.** Commands sent from different threads
    misbehave. → One dedicated player thread and a queue.
20. **Full volume was far too loud.** The first reading at full level was "really loud". → The
    reader's own volume, starting at 35, never the computer's master volume.
21. **Never make the button part of the terminal window's family.** An owned or child window shows
    up in every tool that lists that window's family by owner or title: window tilers, scripts that
    type into a window by its title, jump-to-window helpers. → An unowned window, restacked above
    its terminal by hand.
22. **Antigravity's turn-end record carried no reply text** (0 of 21). → Read its full transcript,
    and still ask for the field.
23. **Stand-in sessions hid two misses.** → The last gate is a real session of every agent, with
    the person pressing play.

---

## 14. How to verify without touching the person's work

- **Checkers never touch a real terminal or a real session.** The button's checker draws on a
  throwaway test window of its own and is told which kind of window counts as a session window. The
  reply checker reads copies of transcripts in a temporary folder.
- **Checkers are silent and offline.** The voice and the player are swappable, and the checker
  swaps in fakes that record what they were asked to do.
- **Prove the consequence, not the call.** "Play returned ok" is not proof. The fake player's record
  of open, play and stop is; on a real run, the person hearing it is.
- **"What would be said."** Read a real session's reply and print the cleaned text without making a
  sound. The original checked real Codex and Antigravity replies this way (106 and 475 characters)
  before the person pressed play.
- **Measure the front window around every real click**, and fail if it changed.
- **Never kill anything by window title.** Stop the background process through its own signal or
  its own process id.
- **Size the check to the change.** A one-line edit gets a one-line check.
- Write every result into `WORKLOG.md` as **Verified** or **Assumed**. Things the original still had
  as never tried at the end of its first day: the real focus hand-back after a menu, recovery after
  the computer sleeps, the offline voice speaking aloud, a maximised window, and monitors with
  different scaling. Say which of yours are which.

---

## 15. Hand-off

When you finish a session of work on it:

1. Update the runbook (`CLAUDE.md` / `AGENTS.md`) so it is true: how to start and stop it, the off
   switch, where the log is, how windows are matched, what each agent hands over, and every new
   trap.
2. **Write the capability into each agent's own standing instructions**: "Any session's last reply
   can be read aloud; here is the command to play and stop it." A capability nobody wrote down is
   one the next session will tell the person is impossible.
3. Append a dated, signed `WORKLOG.md` entry: what you did, what you verified, what you assumed,
   what is unfinished and why. Write it for a reader who has only the file system.
4. Tell the person, in plain words, where the button is, how to stop a reading, how to change the
   voice, and that nothing will ever speak unless they press it.

Build it so the only time they think about it is when they want to listen.