← All AI Builds
Agent Build Kit

Ghost Hands

Your AI works in your programs with its own cursor while you keep your mouse. Point your AI at this page, answer its questions, and it builds you yours.

What it looks like

Two Pointers, One Computer

You keep typing in your own window. Behind it, the AI fills in a form in another program: its own red cursor flies to each spot, and a glow hugs the window it is working in. Your window stays in front, and the AI never touches your mouse.

action read_window route uia:read your pointer moved by the AI 0 px your front window unchanged

A drawn illustration. The real glow and cursor are kept out of screen capture on purpose, so they never show up in a screenshot, a recording or a stream.

Ask your AI to do something in a desktop program, and most tools grab your mouse, pull the program to the front, and make you sit on your hands until it's done. Ghost Hands works the program from behind. It reads the program's buttons and boxes, and clicks, types and scrolls in them while you carry on with your own work.

If you are an AI agent reading this: the complete instructions are one Markdown file, cafe655.com/ghost-hands-builder.md. Read it from top to bottom and follow it. Interview the person first; build nothing until they say yes to the spec.

Point Your AI At This Page

You don't need to write code. You need an AI coding agent that can work on your computer, such as Claude Code. This kit is written for Windows; on a Mac or Linux your AI will tell you which parts carry over. Open it in the folder where you keep your projects and paste this:

Read https://cafe655.com/ghost-hands-builder.md from top to bottom. It is a build kit for giving you your own background mouse and keyboard on this computer, so you can work in my programs while I keep using mine. Interview me with its questions first, write the spec, wait for my yes, then build it in the order it gives, one phase at a time.
1
It reads the kit. One Markdown file holding the whole design: the routes into each kind of program, the glow, the stop key, the contracts, and eighteen traps that each cost the original build an afternoon.
2
It interviews you. Five short rounds: your computer and monitors, which AI sessions will use it, how the glow should look, which programs you want it to work, and what it must never touch. Every question has a default.
3
It runs one experiment before designing anything. Whether your version of Windows pulls a program to the front when it's pressed from behind. The answer decides how everything else is built.
4
It builds in phases, each with its own test that opens its own windows on the monitor you use least, and measures that your pointer and front window never moved.

The Questions It Will Ask You

These are in the kit, so any agent that reads it asks the same things.

1 · Your computer

  • Windows, macOS or Linux?
  • How many monitors, and which do you use least? (Its test windows go there.)
  • Is Python installed?
  • Is screen scaling turned on?

2 · Who will use it

  • Which AI agents should have it?
  • Do you run several AI sessions at once?
  • Does each session already wear a colour or name somewhere?

3 · The look

  • Glow colour for each session or agent
  • What the pill on the window should say
  • A soft breathing glow or a steady line
  • A quiet tick on each click, or silence

4 · What it may touch

  • The three programs you most want it to work
  • Anything it must never touch
  • Do you already have a tool for web pages?

5 · Safety

  • Which key stops everything? (Esc by default.)
  • When it meets a program it can't work from behind, should it stop and tell you? (Yes, strongly recommended.)
  • May it open programs by itself?

What It Does

Sees a Program

Lists its buttons, boxes, lists and menus as a numbered outline. The AI aims at a number instead of guessing pixels from a screenshot.

Works It From Behind

Click, double-click, right-click, type, press keys, scroll, pick from a list. The program stays where it is and your pointer stays where you left it.

Proves It Every Time

Every action measures your pointer and your front window before and after, and reports the numbers. Not "it shouldn't move": "it moved 0 pixels".

Shows Who Is Working

A drawn cursor flies to each spot and pulses on a click. A glow in the session's own colour hugs the window, with a pill naming who is working and how to stop it.

Several AIs at Once

Each session gets its own colour, cursor and pill, so two AIs working two programs are told apart at a glance.

Esc Stops Everything

Your physical Esc key stops every AI that is working. Click or type in the window it is using and that AI pauses, then looks again before carrying on.

Says No Out Loud

It turns down terminals, password and security prompts, administrator windows and anything on your own list, with a plain sentence. It never quietly borrows your mouse.

Lends Its Glow

Any other script that already drives a program can light the same glow with two lines, so you always know when something is working a window.

How It Works

One small background service per person, and three ways in. The service starts the first time anyone asks for it and goes away by itself after a quiet quarter of an hour.

1
The service. It owns the glows, the cursors, the Esc key and the action engine, and listens on a private local channel. Actions run one at a time.
2
Read, then act. The AI reads a window and gets a numbered outline. Each number comes with a token. Reading the window again makes every older token stale, so the AI can never click a button that has since moved or vanished.
3
The cursor lands first. The engine works out where the action will happen, flies the drawn cursor there, and only then acts, so what you see is what happened.
4
The glow follows the window. Four thin strips sit around the window's edge, one layer above it and below everything above it. They follow it as it moves within a few milliseconds and stay out of screen capture.
5
Three front doors. A command line any agent can run, an MCP server so agents get it as native tools, and a two-line library call for any other script.
0 pxyour pointer moved, measured on every action
144 msmedian action on an ordinary program
3–14 msfor the glow to follow a moving window
~1%of one CPU core while the glow breathes

Which Programs It Can Work

Programs are built in different ways, and each kind needs a different way in. The engine picks the route from the kind of window.

Kind of programHow it is worked
Classic Windows programsThrough each control's own messages: the same notice a button sends when it is pressed. Typing goes in at the cursor and leaves the rest of the box alone.
Newer Windows programsThrough an older accessibility interface that presses a control without pulling its window forward.
Desktop apps built like web pagesDiscord, Slack, Spotify and many more. Opened once with a developer channel, then worked through it, the same way a browser add-on works a page.
Web pages in a browserNot this tool. A browser automation tool does that better, and this one says so.
Games and programs that draw everything themselvesNone. There are no buttons to read, so it stops and tells you instead of taking your mouse.

Build Order

Each phase ends with a test that passes and something you can see. Phases 2 to 4 don't depend on each other, so an AI that can run helpers can build them at the same time.

PhaseWhat you get
1 · Spec and experimentYour yes to a one-page spec, and the measured answer to "does pressing a button from behind pull the window forward here?"
2 · The glowA glow, pill and cursor round a test window, following it as it moves.
3 · Seeing and actingReading and working ordinary programs, each action checked against the test window's own record of what happened to it.
4 · Web-style appsThe developer-channel route, tested against a browser window it opens itself.
5 · The serviceThe command line, the MCP tools, the Esc key and pausing when you touch the window, end to end.
6 · Review and demoA fresh AI hunts for defects, then a live demo while you keep using your own mouse, then your three programs, one at a time.

Rules Worth Copying

The kit carries eighteen. These apply to almost anything you'll ever let an AI do on your computer.

Measure before you design

The obvious way to press a button from behind turned out to pull the window to the front first. One short experiment found it; building on the assumption would have wasted the whole build.

Prove it with a number

Every action reports how far your pointer moved and whether your front window changed, measured around the action. A claim the code makes about itself is not proof.

Never guess which window

A name that matches two windows gets both listed back, not a coin toss. Clicking in the wrong program is worse than asking.

Stale means stale

Reading a window again cancels every older token, so the AI can't click where a button used to be.

Stop means now

The first version let one queued action run after Esc. Stop now clears the queue, and a timed-out action is cancelled so it can't arrive seconds later in whatever is on screen by then.

Yield to the human

If you touch the window the AI is working in, it pauses and must look again. It never carries on blind into a screen you've changed.

Refuse out loud

A program it can't work from behind gets a plain sentence and a stop. The tempting workaround, borrowing your real mouse for a second, is exactly what it exists not to do.

Put back what you change

The only thing it ever changes on a program is the colour of the thin line round its window, and it saves the old colour first, even surviving a crash.

Tests make their own windows

Every test opens its own test windows on the monitor you use least, closes only what it opened, and talks to a private copy of the service so it never disturbs one that's in use.

Write the capability down the day you build it

An AI that wasn't told it has its own mouse will tell you it can't work your programs. The kit's last step puts the command in the AI's standing instructions.

The Build Kit

This is the whole thing your AI reads: the interview, the experiment, the contracts, every route, the glow, the stop key, the traps and the build order. Copy it, open it as a file, or just give your AI the address.

ghost-hands-builder.md Open file
# Build Ghost Hands: Your AI's Own Mouse — the Agent Build Kit

> **You are an AI coding agent** (Claude Code, or any agent that can write and run code on this
> computer) and a person has pointed you at this file. It is the complete recipe for building them
> **a background mouse and keyboard for their AI**: a small service that lets any AI session open a
> normal desktop program, read its buttons and boxes, and click, type and scroll in it **while the
> person keeps using their own mouse**. A cursor drawn by the tool shows where it is working, and a
> glow in the session's colour hugs the one window it is working in.
>
> It was written from a working build on Windows 11. Every rule in it was paid for by a real bug or
> a real measurement.
>
> **Do not write any code yet.** Read this whole file first. Then run the interview in Part 1,
> write the answers into a one-page spec, and get the person's yes before you build anything.
> The person may not be a programmer. Ask in plain words, offer a sensible default for every
> question, and never ask more than about five questions at once.

Human-readable version: https://cafe655.com/ai-field-notes/ghost-hands-builder

---

## Contents

0. What you are building
1. The interview (run this first)
2. Write the spec and get a yes
3. The first experiment: measure before you design
4. The architecture: one service, three front doors
5. The contracts
6. Seeing a program
7. The routes: how each kind of program is worked
8. The glow, the pill and the drawn cursor
9. Stopping it, and yielding to the person
10. What it turns down
11. Opening programs without stealing focus
12. The automatic glow (optional)
13. macOS and Linux
14. Build order and gates
15. The traps: every rule here cost somebody a day
16. How to verify without touching the person's work
17. Hand-off

---

## 0. What you are building

Most "computer use" tools for AI work one of two ways, and both are bad for a person sitting at
the same computer:

- **They take the real mouse.** The pointer moves by itself, the program jumps to the front, and
  the person has to sit on their hands until the AI is done. Some hide the real pointer behind a
  drawn one so it *looks* like the AI has its own, but the real one is still being moved.
- **They work invisibly.** Things change in a window and nothing on screen says which AI did it
  or how to stop it.

Ghost Hands does neither. It is:

- **Background first, and measured.** The person's real pointer is never moved and their front
  window is never changed. Every action **measures** both, before and after, and reports the
  numbers. "It shouldn't move" is not the claim; "it moved 0 pixels" is.
- **Visible.** A cursor drawn by the tool flies to each spot and pulses on a click. A thin glow in
  the AI session's own colour hugs the window being worked, with a small pill on its top edge:
  "CLAUDE is working in Notepad · Esc to stop". Several AI sessions at once each get their own
  colour and cursor. The glow and cursor are kept out of screen capture, so they never appear in
  a recording, a stream or a screenshot.
- **Stoppable.** The physical Esc key stops every AI session that is working. If the person
  clicks or types in the window the AI is working in, that AI pauses and must look again before
  carrying on.
- **Honest about what it can't do.** Games and programs that draw everything themselves have no
  controls to read. The tool says so and stops; it never quietly borrows the real mouse.

Shape: **one small background service per user**, which owns the glow, the cursors, the Esc key
and the action engine, and listens on a local named pipe. On top of it, three front doors that
all speak to the same service:

1. a command-line tool (`hands windows`, `hands state <window>`, `hands click …`),
2. an MCP server, so any AI session gets the tools natively (`list_windows`, `read_window`,
   `click`, `type_text`, `press_key`, `scroll`, `set_value`, `glow_on`, `glow_off`),
3. a two-line library call any other script can wrap around its own window work to light the
   same glow.

The original is Python 3 on Windows 11: about 7,500 lines, plus five self-checking test programs
of about 3,600 more.
Its measured results: the real pointer moved **0 px** across every action; actions took a median
of **144 ms** on ordinary windows and **18 ms** on the web-app route; the glow followed a moving
window within **3 to 14 ms**; the service used about **1% of one CPU core** while glowing.

---

## 1. The interview (run this first)

Ask these in rounds. Use your question tool if you have one; otherwise ask in plain text, a round
at a time. **Offer a default for every question** so the person can just say "default". Record
every answer.

### Round 1 — the machine

1. **What computer is this?** Windows (which version), macOS or Linux. *(This kit is written for
   Windows. On macOS or Linux, read Part 13 first and tell the person plainly which parts carry
   over and which must be researched and tested from scratch.)*
2. **How many monitors, and which one do you use least?** Every test window you open goes on that
   one, far from where they work. *(Default: the left-most monitor, top-left corner.)*
3. **Is Python 3.10+ installed?** If not, may you install it? *(Default: Python. The original
   used it, plus pywin32, comtypes, Pillow and numpy, all free.)*
4. **Is screen scaling on** (125%, 150%)? It changes how a click point is worked out. You can
   check this yourself; do.

### Round 2 — who will use it

5. **Which AI agents should be able to use it?** Any agent that can run a command can use the
   command-line front door. Agents that speak MCP (Claude Code and many others) get the tools
   natively. *(Default: the agent reading this file.)*
6. **Do you run several AI sessions at once?** If yes, each gets its own colour and cursor, and
   you need a way to tell which session is asking (Part 5.6).
7. **Does each session already have a colour or a name somewhere** (a terminal tab colour, a
   dashboard)? The glow should use the same one, so the person knows at a glance whose hands
   these are. *(Default: one colour per agent, picked in Round 3.)*

### Round 3 — the look

8. **Glow colour** per session or per agent. *(Default: the agent's brand colour.)*
9. **What should the pill say?** *(Default: "<NAME> is working in <Program> · Esc to stop".)*
10. **How strong a glow?** A soft breathing edge, or a steady line. *(Default: soft, slow
    breathing, about 6 px wide.)*
11. **A sound on each click?** *(Default: none. If yes, make it a very quiet tick. A full-volume
    beep in headphones is painful; the original learned this the hard way.)*

### Round 4 — what it may touch

12. **Which programs do you most want it to work?** Name three. These become the first real
    tests after the checkers pass. *(Ask whether each is an ordinary program, a Microsoft Store
    app, or a desktop app built on web technology, such as Discord, Slack or Spotify. If they
    don't know, find out yourself.)*
13. **Anything it must never touch?** The defaults in Part 10 (terminals, password and security
    prompts, administrator windows) always apply; add theirs.
14. **Web pages in a browser:** these are better worked by a browser automation tool than by this
    one. Do they already have one? *(Default: this tool turns browser page windows down and says
    which tool to use instead.)*

### Round 5 — safety

15. **Which key stops everything?** *(Default: Esc. Only the physical key counts.)*
16. **When the AI meets a program it can't work in the background** (a game, a program that draws
    everything itself), what should happen? *(Default, strongly recommended: stop and tell the
    person, and never borrow the real mouse.)*
17. **May it open programs by itself?** *(Default: yes, on the monitor from question 2, without
    taking focus.)*

---

## 2. Write the spec and get a yes

Write one page, in the person's words where you can:

- the machine, monitors and scaling; the test monitor;
- which agents use it, and how each session is told apart (name, colour);
- the look: colours, pill text, glow style, sound;
- the first three programs to prove it on, and their kind;
- the never-touch list and the stop key;
- which front doors to build (default: all three);
- the build order from Part 14, with what they will see at each gate.

**Show it and wait for a yes.** Nothing gets built before that.

---

## 3. The first experiment: measure before you design

**Do this before writing the action engine.** It decides the whole design, and it differs between
operating system versions.

The obvious way to press a button in another program on Windows is UI Automation: find the
button, call its Invoke pattern. **On the original machine, UI Automation's own action calls
(Invoke, Toggle, SelectionItem, ExpandCollapse, Value, RangeValue, Scroll) brought the target
window to the front before acting.** So did the accessibility calls on classic and WinForms
controls. That is exactly what this tool exists to avoid.

The experiment, on a test window you open yourself on the test monitor:

1. Note the front window and the pointer position.
2. Call one action pattern on a control in your test window while another window is in front.
3. Measure the front window and pointer again.
4. Repeat for each pattern, and for each kind of window (classic Win32, WinForms, WPF).

What the original found, and therefore built:

- **Classic Win32 and WinForms controls are worked with their own window messages**: the click
  notification a button sends its parent, the check, list, drop-down and slider messages,
  replace-selection for edit boxes, the wheel message to the child under the point, posted keys.
- **WPF and other modern programs are worked through the older accessible interface** (its
  default action, select and set value), which did **not** activate the window.
- **The modern UI Automation action calls are used only when the window is already in front.**
  Otherwise the answer is `background_unavailable`.

Reading is different: reading a window's controls through UI Automation activates nothing.

**Warn the person before running the experiment:** while you measure, their front window may be
taken for a fraction of a second. Run it on your own test window, never on theirs, and never
while they are typing something important.

---

## 4. The architecture: one service, three front doors

```
  AI session ──► CLI ──┐
  AI session ──► MCP ──┼──► named pipe ──► SERVICE (one per user)
  any script ──► lib ──┘                     ├─ overlay thread: glows, pills, cursors
                                             ├─ action engine: one action at a time
                                             ├─ input listener: physical Esc, touches
                                             └─ idle timer: quits after 15 quiet minutes
```

- **One service per user**, started by the first request and quitting by itself after 15 minutes
  with no requests. Start it **detached from the session that asked** (on Windows, a WMI process
  launch); a child process started from an AI's shell can be killed when that shell ends.
- **A named pipe**, newline-delimited JSON, one request and one answer per line. Cap a request
  line (the original: 1 MB) and the number of connections served at once (8; the next waits).
- **One action engine.** Actions are run one at a time. Every action computes its route and the
  screen point **before** acting, flies the drawn cursor there, then fires.
- **The overlay runs on its own thread** with its own message loop, inside the service.
- **The three front doors are thin.** All the logic is in the service; the CLI, the MCP server and
  the library only build requests and print answers.

Files in the original, as a guide to splitting the work: `windows` (listing, identity, refusals),
`uia` (reading controls), `capture` (a picture of one window), `act` (the engine and ordinary
routes), `cdp_route` (the web-app route), `overlay` (glow, pill, cursor), `service` (pipe, engine
queue, listener), `client` library, `cli`, `mcp`, and one checker per part.

---

## 5. The contracts

Write these into a `CONTRACT.md` **before** any code, and build every part to it. If several
agents build parts in parallel, none of them edits the contract; a change is raised and agreed.

### 5.1 Coordinates

Every process is **per-monitor DPI aware v2**. All coordinates are **physical screen pixels**,
except `wx, wy`, which are window-relative with the origin at the top-left of the window's real
visible frame (on Windows, the DWM extended frame bounds, not `GetWindowRect`, which includes an
invisible border).

### 5.2 A window's identity

```json
{"hwnd": 132456, "pid": 8812, "exe": "notepad.exe", "title": "Untitled - Notepad",
 "class": "Notepad", "rect": [l, t, r, b], "minimized": false, "elevated": false,
 "kind": "win32" | "uwp" | "electron" | "chromium-browser" | "terminal" | "other"}
```

A window can be named by number, by program (`notepad` or `notepad.exe`), by `pid:<n>`, or by part
of its title. **A name that matches two windows is never guessed at**: the answer is `not_found`
and it lists both.

### 5.3 The snapshot and its tokens

Reading a window returns a numbered outline of its controls and a snapshot id:

```
- [12] Button "OK" id=btnOk actions=[invoke]
  - [13] Edit "File name" value="report.txt" actions=[set_value, text]
```

Each actionable control gets a token, `s1a2b3c4d:12` (snapshot id and number). **A new snapshot of
the same window makes every older token for it `stale`.** That one rule stops the AI clicking a
button that has since moved or vanished.

Action vocabulary: `invoke, toggle, select, expand, collapse, set_value, set_range, scroll, text`.

### 5.4 The result of every action

```json
{"ok": true, "route": "post:click", "effect": "confirmed" | "unverifiable" | "none",
 "evidence": "BN_CLICKED to parent", "cursor_moved_px": 0, "foreground_changed": false,
 "at": [x, y], "error": null}
```

`cursor_moved_px` and `foreground_changed` are **measured** around the whole call, including the
cursor's flight. They are never filled in from what the code believes it did.

### 5.5 Error codes

`stale`, `not_found`, `refused`, `elevated`, `minimized`, `background_unavailable` (this program
can't be worked without the real mouse; the message names it, and the caller stops and tells the
person), `timeout`, `stopped` (the stop key), `paused_user` (the person touched this window; read
it again before continuing), `internal`.

### 5.6 Who is asking

Every request carries a session key, a display name and a colour. Work them out once, in the
client library: the agent's own session id if it publishes one, a name the person gave the
session, the colour the session already wears elsewhere. Let environment variables override all
three. An explicit colour on a request always wins.

---

## 6. Seeing a program

- Read the control tree through UI Automation in **one cached fetch per window**, not one call per
  control. Limits: 4,000 nodes, depth 30, 4 seconds; say which limit was hit if one was.
- **Drop off-screen controls.** Drop nodes with neither a name nor a value and lift their
  children. Cap values at 200 characters.
- Hit-test a point with the rectangles from the last snapshot plus one element-from-point call,
  never a pixel search.
- A picture of one window (`--screenshot`) uses PrintWindow, with a check for an all-black frame
  (some programs draw nothing through it; say so rather than returning a black picture).
- **Read, then act.** Tokens come from a read. After anything that changes the screen, read again.

---

## 7. The routes: how each kind of program is worked

The engine picks the route from the kind of window. Whatever the route, the window is not brought
forward and the real pointer is not moved.

| Kind of program | Route | Notes |
|---|---|---|
| Classic Windows controls and WinForms | The control's own window messages | Typing uses replace-selection, so it goes in at the caret and leaves the rest of the field alone |
| WPF, UWP and other modern programs | The older accessible interface (default action, select, set value) | Does not activate the window |
| A click with no control action (right-click, double-click, a bare spot) | A mouse message posted to the exact child window under the point | Scroll goes to the child under the point, not the top-level window |
| Desktop apps built on web technology (Electron / Chromium: Discord, Slack, Spotify and many more) | The DevTools channel: the app is opened with a remote debugging port, and clicks, typing and wheel go in through it | Only when the app has a debugging port. Opening an app that way needs the person's yes, and the app must be closed first (Part 11) |
| Web pages in a browser | **Not this tool.** A browser automation tool does this better | Browser page windows are turned down with a sentence saying which tool to use |
| Games and programs that draw everything themselves | None | `background_unavailable`, naming the program. Stop and tell the person |

**On the web-app route, which page is in the window?** An app can have several pages on its
debugging port. Choose the visible one; if you cannot tell which page fills the window, answer
`not_found` and list the candidates. Never guess. Map page coordinates to screen coordinates so
the drawn cursor lands where the click lands.

**Programs that are not DPI-aware:** convert a click point under that window's own DPI setting, so
it lands where the drawn cursor shows on a scaled monitor.

**A timed-out action can never be delivered late.** Give each action a limit (the original: 10
seconds). When it runs out, cancel it so that nothing it had queued is sent afterwards, and answer
`timeout`.

---

## 8. The glow, the pill and the drawn cursor

- **The glow hugs the one window being worked, not the whole screen.** Four thin click-through
  strips sit round the window's real frame. They follow it through move, resize and minimize by
  **window event hooks, not polling**, and sit **one slot above the target** in the stacking
  order, so they never cover the person's other windows.
- **The pill** sits on the top edge: who is working, in which program, and how to stop it.
- **The window's own frame line** (on Windows 11, the DWM border colour attribute) is set to match
  the glow as a hairline. **Save each window's own border colour before changing it and put it
  back afterwards.** Keep the record in a file, so windows left coloured by a service that died
  are restored when the next service starts.
- **Keep everything out of screen capture** (on Windows, display affinity "exclude from capture").
  Make it switchable, but only for the checker, which must sample pixels.
- **The drawn cursor**, one per session, moves on a gentle curve and pulses on a click. The action
  fires **when it lands** (the engine waits; cap the flight at about 450 ms).
- **The glow fades** 6 seconds after a session's last action unless held. A held glow ends after an
  hour regardless, so nothing can stay lit forever.
- **Several sessions at once** each get their own strips, pill and cursor, keyed by session.
- Breathing should be cheap: the original measured about 1% of one core while glowing.

---

## 9. Stopping it, and yielding to the person

- **The physical stop key stops everything.** A low-level keyboard listener watches for it but
  **never swallows a key**, and **ignores keys made by software**, so only the person's own key
  counts. It stops every session that worked in the last five minutes. A stopped session's next
  calls answer `stopped` until it deliberately wakes again.
- **Stop means stop now**, including an action already waiting in the queue. (The original's
  first version let a queued action run after the stop.)
- **The listener has a heartbeat.** Check every 2 seconds that the hooks are alive; Windows can
  drop them silently. Put them back, restart a dead listener, and report the listener's state in
  `status` so an agent can check it before relying on the key.
- **If the person touches the window being worked, the AI pauses.** A real click or keypress in
  that window makes that session's next action answer `paused_user`, and it stays paused until the
  session reads the window again. It never carries on blind into a screen the person has changed.
- **A `stop` command** does the same as the key from the command line, with a `--mine` option that
  stops only the asking session.

---

## 10. What it turns down

These come back as `refused` (or `elevated`) with a plain sentence. **A refusal means stop and ask
the person. Never work around it, and never borrow their real mouse to get past it.**

| Turned down | Why |
|---|---|
| Terminals and console windows | The AI sessions themselves live in them. Typing into the wrong one is how an AI sends instructions to another AI by accident |
| Browser page windows | A browser automation tool works web pages better and more safely |
| Administrator (elevated) windows | A normal process cannot safely drive them. If the person needs this, build a separate helper they approve once |
| The lock screen, password and security prompts, any secure desktop | Nothing is worked there, ever |
| Anything on the person's own never-touch list | Their call |

---

## 11. Opening programs without stealing focus

- **Open programs without activating them**, and **place them on the test monitor** (or wherever
  the person chose) by moving the window without activation and without changing the stacking
  order. Un-maximize first if needed. Some programs jump back to a remembered spot a moment later:
  watch for about 1.5 seconds and put them back.
- **Microsoft Store apps** are started through an alias that a plain process launch refuses (on
  the original, WMI answered error 8). Fall back to the shell's own start command, run hidden, and
  say in the answer which route was used.
- **Opening a web-technology app with a debugging port** (`--remote-debugging-port=N`) is a
  separate, explicit command, run only when the person asked. **If the app is already running,
  refuse and start nothing**: a second copy hands itself to the running one and brings it to the
  front. Match "already running" by program file name, so a Store alias matches the running Store
  copy.
- Some programs bring themselves to the front on their own (closing an "upgrade available" offer
  did this on the original). That is the program, not the tool; the measured result will show it,
  and the answer should say so.

---

## 12. The automatic glow (optional)

If the person has other scripts that already drive programs (scripts the AI runs to work a music
app, a scanner, a dashboard), the glow can light round those programs too, with no change to the
scripts:

- **The library call:** `with operating(window, label="Reading the sensors"): …` lights the glow
  for the duration of the block. **It must never break the work it wraps**: if the service cannot
  be reached, it does nothing and the code inside runs as normal.
- **A hook in the AI agent's own settings** (for Claude Code, PreToolUse and PostToolUse hooks)
  runs a tiny matcher before and after every tool call. A list of programs maps each one to the
  scripts and tools that work it. A matching call lights that program's glow until the call ends.
  - Only **running** a program's script counts; reading or searching its files does not.
  - Common script names only count from that program's own folder.
  - A call that matches nothing must exit fast (the original: about 40 ms) and load nothing heavy.
  - Use a separate session key for the automatic glow, so it never ends a glow the session lit
    itself and never undoes a stop.
  - Cap it (the original: 2 minutes) so a missed "call ended" can never leave a glow stuck.
  - **Back up the agent's settings file before adding the hook**, and check afterwards that every
    older hook is still there and unchanged.

---

## 13. macOS and Linux

The design carries over: one service, a pipe, three front doors, read-then-act with stale tokens,
measured pointer and front window, a glow, a stop key, and the same refusals. The mechanics do
not. On macOS the accessibility API (AXUIElement and its press and set-value actions) is the
place to start, the overlay is a borderless, click-through, non-activating panel, and the app needs
the person to grant Accessibility and Screen Recording permission once. On Linux it depends on the
desktop: AT-SPI for reading and acting, and an overlay that may not be possible under every
Wayland compositor.

**Run Part 3's experiment on their system before designing anything.** Whether an action call
brings the window forward is the question the whole tool rests on, and it has to be measured
there, not assumed from this file. Tell the person plainly which parts you have proven and which
are still guesses.

---

## 14. Build order and gates

Each phase ends with a gate: its checker passes, and you show the person something working.
Record each gate in `WORKLOG.md` as **Verified** (you ran it and saw it) or **Assumed** (you
didn't). Phases 2, 3 and 4 do not depend on each other and can be built at once.

1. **Spec, experiment and contract.** Part 2, then Part 3 on your own test window, then
   `CONTRACT.md`. Gate: the person's yes, and the experiment's numbers written down.
2. **The overlay.** Glow strips, pill, drawn cursor, window tracking, capture exclusion, several
   sessions. Gate: its checker lights a glow round its own test window, samples the pixels (with
   capture exclusion switched off in the test only), moves the window and measures how fast the
   glow follows, and measures CPU while it breathes.
3. **Seeing and acting.** Reading the tree, the window picture, the ordinary routes, typing at the
   caret, scrolling. Gate: its checker opens its own classic window and its own WinForms form and
   works every control in them, **checking each action against the test window's own event log**
   (not against what the tool says about itself), with pointer movement and front-window change
   measured per action. It also looks up a terminal and confirms it is refused.
4. **The web-app route.** Gate: its checker opens a **visible** browser window with a temporary
   profile and its own debugging port, on the test monitor, and reads, clicks by token and by
   point, toggles, types at the caret, scrolls, selects, sets a value, refuses a stale token, and
   checks the cursor mapping against a screenshot. It closes only the browser it started.
5. **The service, the CLI, the MCP server, the stop key and yield-on-touch.** Gate: the whole chain
   end to end through the CLI's JSON output: read, click, type, scroll; the glow on and its fade;
   hold; yield-on-touch; a simulated stop; `--dry-run` touching nothing; exit codes; the MCP server
   listing its tools over a real stdio handshake; a clean shutdown.
6. **A read-only review** of the whole thing by a fresh agent, looking for defects. The original's
   review found fifteen (Part 15 lists them). Fix each with a new check.
7. **The live demo.** Open a harmless program (Character Map is a good one) on the test monitor
   and work it while the person uses their own mouse elsewhere. Show them the measured numbers.
8. **The person's three programs**, one at a time, with them at the computer.
9. **Optional:** register the MCP server for every session (ask first; it is a global change), and
   the automatic glow (Part 12).

Use `--dry-run` on every action while building: it says which window, which route and where the
cursor would land, and touches nothing.

---

## 15. The traps

Each one was a real defect in the original. The rule is what prevents it.

1. **The obvious action call brought the window forward.** → Part 3. Measure it on the person's
   system before designing the engine.
2. **Typing replaced the whole field.** Setting a value overwrites what was there. → Type at the
   caret (replace-selection, or a text-pattern insert). Keep "set the whole value" as a separate,
   explicit action.
3. **Scrolling went nowhere.** The wheel message was sent to the top-level window. → Send it to the
   child window under the point.
4. **Other tools "worked in the background" by disabling the target window** or flipping its
   styles while acting. → Never disable, re-style or activate a target. The only permitted change
   is its frame colour, saved and restored.
5. **The stop key didn't stop a queued action.** → Stop clears the queue and cancels the action in
   flight.
6. **An ambiguous window name was guessed at.** "Notepad" matched two windows and the tool picked
   one. → Two matches is `not_found` with both listed.
7. **A timed-out action was delivered seconds later**, into whatever was on screen by then. →
   Cancel on timeout; nothing queued is sent afterwards.
8. **A blocked browser slipped through the "open a program" command.** → Apply the refusal list to
   launching as well as acting.
9. **Glows stuck on after a missed end.** → Every glow fades by itself; a held glow has a hard
   limit; the automatic glow has its own cap.
10. **Clicks landed in the wrong spot on a scaled monitor.** → Per-monitor DPI awareness
    everywhere, and convert under the target window's own DPI setting.
11. **One giant request could jam the pipe.** → Cap request size and simultaneous connections.
12. **The checker talked to the shared service** other sessions were using. → Every command takes
    a `--pipe NAME`; checkers start their own private service on their own pipe.
13. **A Store app wouldn't open.** The plain process launch refused its alias. → Fall back to the
    shell's start command, hidden, and report the route.
14. **Opening an app with a debugging port while it was running** handed control to the running
    copy and brought it to the front. → Refuse if it is already running.
15. **A service started from an AI's shell died with that shell.** → Start it detached, through
    the operating system's own process launcher.
16. **Windows silently dropped the keyboard hook,** and the stop key stopped working. → A
    heartbeat that re-hooks, and `status` reports whether the listener is live.
17. **A loud click sound in headphones.** → If there is a sound at all, a short quiet tick.
18. **Testing took the person's front window.** While the action calls were being measured, test
    runs briefly took the person's front window several times. → Test windows on the monitor they
    use least, warn them first, and tell them the same turn if it happens.

---

## 16. How to verify without touching the person's work

- **Checkers open their own test windows** on the test monitor, and touch nothing else. They close
  only what they opened.
- **Prove the consequence, not the call.** A test window keeps its own log of what happened to it
  (button pressed, text changed, item selected). The checker reads that log. "The tool returned
  ok" is not proof.
- **Measure the person's pointer and front window around every single action**, and fail the
  check if either changed because of the tool.
- **A private service on its own pipe** for every checker run, never the shared one.
- **Print the numbers every run**: pointer pixels moved, follow latency, CPU. Re-run a checker
  rather than quoting an old number as current.
- **Size the check to the change.** A one-line edit gets a one-line check.
- Write every result into `WORKLOG.md` as **Verified** or **Assumed**. Things the original still
  had as assumed after its first day: other web-technology apps beyond the first one proven,
  scrolling in WPF, the physical stop key pressed by a real person, and whether a streaming
  program really leaves the excluded glow out of a live stream. Say which of yours are which.

---

## 17. Hand-off

When you finish a session of work on it:

1. Update the runbook (`CLAUDE.md` / `AGENTS.md`) so it is true: how to use each front door, the
   routes, the refusals, how to stop it, and every new trap.
2. **Write the capability into the agent's own standing instructions** (for Claude Code, the
   user's `CLAUDE.md`): "You can work any normal program in the background with your own cursor;
   here is the command." A capability nobody wrote down is one the next session will tell the
   person is impossible.
3. Append a dated, signed `WORKLOG.md` entry: what you did, what you verified, what you assumed,
   what is unfinished and why. Write it for a reader who has only the file system.
4. Tell the person, in plain words, what now works, how to stop it, and what to try first.

Build it so they forget it's there until they see the glow.