Operating a run unattended
"Unattended" here does not mean nobody decides anything. It means nobody has to be watching a terminal for the run to keep moving, and the moments that genuinely need a person reach that person wherever they are — instead of scrolling past in a window nobody has open.
Four things are still a person's, always: a new product decision, a budget ceiling going up, work that leaves the boundary the What cited, and the final merge. The whole design here is about making those four reachable in seconds, not about removing them.
The two ways to drive a run
A host session — tldrx run attend host. A lock, not an engine: it sets one field on the run and from then on the framework never spawns on it. Every turn is a tldrx next --prepare / tldrx next --commit handshake with a session you drive, so you get that session's judgement and its own tools, and it can already reach you because you are talking to it. What you give up is the metering — those turns are billed to your session, not measured per stage — and the parallelism: a host session drives one turn at a time. This is the mode tldrx drive writes a mandate for.
The engine — tldrx run auto. A headless loop that calls next over and over, spawning a metered sub-agent stage after stage. You get a per-stage USD meter, an enforced model, and stories inside one build wave running in parallel. What it does not have is a way to reach anybody: it announces an open question or a gate by exiting 4 and printing to stdout. Everything below is how you give it one.
The two do not compose, and mixing them is refused rather than guessed at: run auto on a run marked attended_by: host exits 1, before the event log is opened.
Declaring the hook
One optional block in .tldrx/workspace.yml. A workspace without it notifies nothing and behaves exactly as it did before the key existed.
notify:
command: "bin/notify-owner" # a single argv line, like `commands:` — no shell is opened
events: [question.raised, gate.requested, run.failed] # optional; omitted means every kind
timeout_s: 30 # optional; one invocation's ceilingcommandis argv, never a shell line. It is split and executed directly, exactly like every other command a workspace declares, and a bare shell metacharacter (|,&,;,<,>,$, backtick, parens, braces,*,?,~) is refused rather than shelled. If you need a pipeline, put it in a script and declare the script. The payload never touches the command line either — a question's own title could otherwise become shell syntax.eventsfilters by kind. Omitting it means every kind, deliberately: someone who declared a command wants to hear about the run, and a default that silently subscribed to nothing would be a configured hook that never fires and never says why. An unknown kind is a validation error.timeout_sbounds one invocation (30 seconds by default). A notifier is a message, not a job.- A failing notifier never changes the run. A command that will not split, a binary that is not there, a non-zero exit, a timeout — each is written down as a
notify.failedevent carrying the reason and then dropped. "The owner was not told, and here is why" is a fact about the run; "the side channel was down, so the run failed" would make a side channel load-bearing. A delivered one isnotify.sent, with the kind, the child's exit code and its duration. Both cost$0.00.
tldrx init writes the block commented out, with a line saying what it is for. It does not guess a command: who gets woken up is not a thing to detect.
The payload
One JSON object on stdin, version: 1, with the same nine top-level keys every time: version, kind, at, run, root, stage, summary, command, detail.
summary is one paragraph written to be read on a lock screen. command is the exact line to type, run id already in it — or null, honestly, when there is nothing to do; an invented next command would be the framework guessing at an intention. stage is <phase>/<stage>, or null when the notification is about the run as a whole. root is the absolute workspace root, because a script is not otherwise told where the run lives. detail is per-kind and always an object.
question.raised — the one the whole feature exists for
{
"version": 1,
"kind": "question.raised",
"at": "2026-09-07T18:20:04.117Z",
"run": "260907-checkout",
"root": "/Users/alan/code/checkout",
"stage": "01-what/what",
"summary": "260907-checkout stopped at 01-what/what on 1 open question(s): Q1 · Should an abandoned hunt count toward the leaderboard?. The run is parked until one is answered; nothing is being spent while it waits.",
"command": "tldrx answer Q1 \"…\" --run 260907-checkout",
"detail": {
"questions": [
{
"id": "Q1",
"title": "Should an abandoned hunt count toward the leaderboard?",
"why_asked": "no rule for abandoned hunts exists in memory [src: absent:.tldrx/memory/facts.yml]",
"options": [
{ "letter": "A", "text": "count them" },
{ "letter": "B", "text": "drop them" }
],
"recommendation": { "option": "B", "why": "matches how players talk about it", "src": "01-what/handoff.md:22" },
"answer_command": "tldrx answer Q1 \"…\" --run 260907-checkout"
}
]
}
}The top-level command is the first question's answer command — a payload has one command slot, and questions are answered one at a time. A script that wants to render a button per question reads detail.questions[]. The options arrive as {letter, text} rather than a rendered A) … line, so building buttons out of them does not mean re-parsing a string the framework had already parsed.
recommendation comes from one of two places, and it is null when neither carried one — never a manufactured one. An agent gate's evidence note (recommend:) wins; otherwise the question block's own optional Recommended: line does, which the asking stage writes:
- A) count them
- B) drop them
Recommended: B — matches how players talk about it [src: 01-what/handoff.md:22]
[Answer]:That line exists because only an agent gate ever writes a note, so questions parked at an auto gate used to arrive with no guidance at all — while the stage that raised them was the one thing in the run that knew the trade-off. A Recommended: line the parser cannot read is ignored, never refused: it is guidance, so a typo costs the guidance and not the gate.
What you will see when an auto gate is waiting
Two things changed in the order and content of what reaches you, and both are about an auto gate held only by open questions:
- The questions arrive first, and the gate may not arrive at all. When the ONLY thing holding an auto gate is its open questions, the gate is downstream of them rather than a second ask — so
question.raisedis delivered first and thegate.requestednotification is held back. You get "here is what to decide", not "sign this" followed by "and here is why". Thegate.requestedEVENT is still appended to the run's log either way; only the notification waits. - The gate closes itself once you answer. Under
--wait-gates, each poll re-runs the seven auto conditions, and the moment every one holds the loop signs the gate through the sametldrx approvedoor and carries on to the next stage. So the sequence you actually see is: the questions, your answers from your phone, and thenstage.donefor the NEXT stage. No approve tap at all.
If something OTHER than the questions is holding the gate — an unverified citation, a stage over its ceiling — both notifications go out, questions first, and the gate's summary names the condition: "…is waiting at an auto gate that did not close by itself — a person signs it. It is held by: claim-sources=1 unverified citation(s) — …". Before this, that sentence said only "did not close by itself" and named nothing.
A Build gate names the story outcomes. Whoever signs it — human, agent or auto — the Build gate's notification says what the stage actually delivered before it says what it cost:
"260909-scoring finished 04-build/build for $1.78 and is waiting at a human gate — a person signs it. It 0 of 3 stories delivered, S1 blocked (npm run test exited 127…), S2 not started. Nothing runs after it until the gate is approved or rejected."
The counts and the first blocked story's own reason ride on the gate.requested payload as stories, blocked_story and blocked_reason, so a script can route on them. The reason is the handoff's own sentence, never a paraphrase. Two real runs were approved from a phone over a summary that said $1.78 and one green check while every story was blocked — the counting existed and ran for auto gates alone.
And the run's own end says it too. A run whose Build delivered nothing does not reach you as a plain done: run.finished reads "…the loop finished with exit 0 (ok), $1.78 spent by this loop. The run: nothing delivered: 0 of 3 stories; S1 — npm run test exited 127." The same sentence is on run.yml as outcome:, in tldrx run status, on the dashboard and in the tldrx ship PR body — and tldrx ship refuses such a run outright rather than opening a PR whose "What shipped" section is empty.
status — the heartbeat
{
"version": 1,
"kind": "status",
"at": "2026-09-07T18:30:04.002Z",
"run": "260907-checkout",
"root": "/Users/alan/code/checkout",
"stage": "02-how/design",
"summary": "260907-checkout is still running at 02-how/design. Nothing is waiting on you — this is the periodic heartbeat `--notify-every` asked for.",
"command": "tldrx run status 260907-checkout",
"detail": { "status_text": "…what `tldrx run status` prints, verbatim…", "waiting_on": [] }
}When waiting_on is not empty the heartbeat changes what it says: it names the open questions and its command becomes the literal tldrx answer line. A heartbeat that went on saying "nothing is waiting on you" while the run sat on somebody's answer would be worse than silence, because a heartbeat is believed. Parked-ness is decided by the same predicate --wait-answers polls and next parks on, never a second opinion.
A run parked on a gate had exactly the same hole, and it is closed the same way. While a signature is pending the payload grows two keys:
"detail": {
"status_text": "…what `tldrx run status` prints, verbatim…",
"waiting_on": [],
"waiting_on_gate": "01-what/what",
"gate_policy": "human"
}…the summary says the run is waiting for a person to sign that stage, and command becomes tldrx approve --run <id>. waiting_on_gate is a sibling of waiting_on, not a member of it: an adapter maps every id in waiting_on to tldrx answer <id>, and a stage id there would make it build a command nobody can type. Both keys are absent when no gate is pending, so an adapter written before this existed sees the payload it always saw.
The nine kinds
The enum is closed — a kind that arrives from nowhere is a branch nobody wrote — so a switch on kind with a default is a complete adapter.
| kind | fires when | command carries | detail |
|---|---|---|---|
question.raised | the loop parked on an open question | the first question's tldrx answer line | questions[] — id, title, why_asked, options[] as {letter, text}, recommendation (option, why, src) or null, answer_command |
question.timeout | --wait-answers lapsed and the loop is about to exit 4 | the same answer line | the same questions[], plus waited_ms |
gate.requested | a stage finished and a person must sign it — deferred, and possibly never sent, when an auto gate is held only by open questions | tldrx approve --run <id> | cost_usd, approve_command, reject_command, gate_policy, and one of held_by (an auto gate's failing conditions) / signer_held (an agent signer's reasons) — absent when nothing looked |
gate.timeout | --wait-gates lapsed and the loop is about to exit 4 | the same approve line | approve_command, reject_command, gate_policy, waited_ms, and cost_usd only when this loop is the one that saw the gate raised |
stage.done | a stage finished and the loop moved on | null — the loop is already running the next stage | cost_usd |
run.finished | the loop ended with exit 0 | null | exit_code, exit_family, spent_usd |
run.failed | the loop ended with any non-zero exit, refusals included | tldrx run status <id> | exit_code, exit_family, spent_usd |
budget.warned | a ceiling is close | tldrx budget show --run <id> | spent_usd, ceiling_usd |
status | every --notify-every <duration> while the loop runs | tldrx run status <id>, or the answer line when parked on a question, or the approve line when parked on a gate | status_text — what tldrx run status prints, verbatim — waiting_on, the blocking open question ids ([] when none), and waiting_on_gate + gate_policy only while a gate is pending |
A truncated input rides in the summary, and adds no kind. When a stage's inputs_max_bytes could not fit a declared input whole, the stage.done, run.failed and status summaries end with one extra sentence — "1 input truncated: facts.yml 169 KB → 87 KB (cap 96 KB)." — so the switch you already wrote keeps working, and you learn that a sub-agent read a prefix rather than the file while the run is still going rather than from .agent/<stage>/prompt.md after it failed.
exit_family is the exit code in words, so a notification on a phone says "refused — a budget ceiling or a gate said no" rather than "exit 2". The questions, their options and their recommendation are the same card run auto --gate-agent prints, so a notification and a terminal can never disagree about what was asked.
An adapter, in about thirty lines
The framework ships no integration with any messaging service, and it is not going to. Who reaches you is different for every workspace, and a built-in integration would be this framework deciding whose product everybody's run depends on — one operator's setup becoming a dependency of everyone else's. The console is the default because it is the one surface every run has; anything past that is yours, and tldrx defers to it without knowing what it is.
So the adapter is the piece you own. Node, no dependencies, no framework knowledge beyond the payload:
#!/usr/bin/env node
// bin/notify-owner — reads one notify payload on stdin.
// Replace this with whatever reaches you: an HTTP POST, an email, a ticket, a phone.
async function sendToMyChannel(title, body, action) {
console.log([title, body, action].filter(Boolean).join("\n"));
}
let raw = "";
process.stdin.on("data", (chunk) => { raw += chunk; });
process.stdin.on("end", async () => {
const p = JSON.parse(raw);
const action = p.command ? `Run this to unblock it:\n${p.command}` : null;
switch (p.kind) {
case "question.raised":
case "question.timeout":
// Every question, with its options and its own answer command.
for (const q of p.detail.questions) {
const options = q.options.map((o) => `${o.letter}) ${o.text}`).join("\n");
await sendToMyChannel(`${p.run} · ${q.id}: ${q.title}`, `${options}\n\n${q.why_asked}`, q.answer_command);
}
break;
case "gate.requested":
case "gate.timeout":
case "budget.warned":
case "run.failed":
await sendToMyChannel(`${p.run} · ${p.kind}`, p.summary, action);
break;
default: // stage.done, run.finished, status — reports, nothing to type
await sendToMyChannel(`${p.run} · ${p.kind}`, p.summary, action);
}
});Then chmod +x bin/notify-owner and declare it. The loop closes when a person reads that message and runs the command it carried — tldrx answer Q1 "B — rankings are global" --run 260907-checkout, from a laptop, a phone over SSH, or a button in your own script that shells out on their behalf. It is an ordinary tldrx answer; the loop never answers its own question.
Running it
tldrx run auto 260907-checkout --notify-every 10m --wait-answers 4h --wait-gates 4h --retry-failed 2All three flags take a duration: 30s, 10m, 2h, or a bare number of seconds. A value that is not a duration is refused with exit 1.
--notify-every <duration>adds the periodicstatuspayload. Off by default, and it does nothing at all unless anotify:command is declared. It exists because the period when you most want to know a run is alive is the twenty minutes it is inside one stage.--wait-answers <duration>changes where the loop stops on a QUESTION. Instead of exiting4the moment a stage parks on one, it polls the run's question files for that long and resumes by itself if somebody answers. Nothing is spent while it waits. When the wait lapses it sends onequestion.timeoutand then exits4with the same lines it always did — the run is intact, nothing was lost, andtldrx run autopicks it up again once the question is answered.--wait-gates <duration>does the same for a GATE, the other half of exit4. It is a sibling flag rather than a wider--wait-answersbecause the two parks are closed by different verbs:tldrx answerfor one,tldrx approve/tldrx rejectfor the other, and calling a signature an "answer" would be the flag name lying about what you did. Approve inside the window and the loop carries on to the next stage; reject and it stops, printing your note; let it lapse and it sends onegate.timeoutand exits4. Nothing is spent while it polls.It waits FOR a signature and produces one only where the run already said it could: an
autogate is re-evaluated on every poll and signed the moment its seven conditions hold (below). Forhumanandagentit produces none — and by the time it is waiting on an agent gate, the engine's own signer has already had its turn (below), so what is left to wait for is a PERSON. Approving an agent-policy gate yourself is a recorded override and is always allowed. The heartbeat and thegate.requestedpayload both name the policy, so you know which of the two you are doing.
Who closes a gate under the engine
Three policies, three different things happen when a stage finishes:
human— the loop stops and a person signs it:tldrx approve, ortldrx reject --note "…". With--wait-gatesthe loop waits for that signature instead of exiting on the spot.agent— the engine spawns one bounded gate signer of its own: the stage's model and effort, a quarter of the stage's per-agent ceiling, allowed to read anything and to write exactly one file,.agent/<stage>/evidence.md. That note then goes through the ordinarytldrx approve --as-agentpath — the same validator a person's note goes through.verdict: signwith every condition holding and every claim carrying a[src: …]closes the gate under the note's ownby:, and the loop walks on. Anything else —refuse,sign-with-fixlist, a note that does not validate, a signer that wrote nothing — leaves the gate pending for you, with the reasons on thegate.requestedpayload. The turn is recorded asagent.spawned/agent.resultwithrole: gate-signerand appears intldrx cost. There is no flag:gates_policy: agentis already your recorded decision that an agent may close it.auto— no signer and no note: seven measured conditions, and the gate closes only if all seven hold. Otherwise it falls to a person with the failing ones named, on thegate.requestedpayload'sheld_byas well as on stdout. And it keeps the offer open: under--wait-gatesthe seven are re-measured on every poll, so a gate held by an open question closes itself as soon as the question is answered. Onlyauto— the run already granted that authority — and your ownapproveorrejectoverrides it at any moment.
Both wait flags may be given together — that is the shape of a fully unattended launch: --wait-answers 4h --wait-gates 4h.
Retrying a stage that failed
--retry-failed <n> is the one flag here that is a COUNT, not a duration: how many times in a row the loop may run a failed stage again before it stops. 0 is the default, and it is what every invocation before this got — one attempt, then exit 5.
A retry is the same tldrx next you would have typed. The stage is on disk as failed with its reason recorded, and the next attempt's prompt is told what the last one did — which is why this is worth automating at all: measured on a real unattended run, a plan that failed a check by five characters passed on the very next attempt, with no new instruction from anybody.
Three things bound it, and all three matter:
- It bounds exit
5and nothing else. A usage error (1), a money refusal (2) and an awaiting-human park (4) are attempted once however largenis. Each is a decision you own — a phase ceiling especially, which means a human decides about money, and a retry would turn that sentence into a delay. - Only consecutive failures count. A stage that succeeds puts the count back to zero, so a long run with one recoverable failure per phase never exhausts a small bound. What is being bounded is "this run is stuck", not "this run has ever failed".
- A retry spends. It is a fresh metered stage under the same phase ceiling and the same
--max-usd. When the bound is spent the loop stops on the failure's own exit5, and the last line says the count —3 consecutive stage failures at 03-plan/plan …— so therun.failedpayload on your phone says the loop tried, rather than a bare5.
The maximum is 3; anything higher is refused by name with exit 1.
Exit 4 is not a failure. It is "awaiting a person", and with the hook declared the person has already been told; what is left is your outer relaunch loop, which is yours to write.
The two things you gain over host mode are worth naming, because they are what the notification buys back:
- Parallelism.
--parallel <n>sets how many stories of one build wave run at once.waves.ymlalready guarantees a dependency is in an earlier wave, so a wave's stories are independent by construction. The shipped build stage declaresparallel: 2, so two at a time is what a workspace overriding nothing gets. - The meter. Every spawned turn is measured per stage and per attempt.
tldrx costprints what the work actually cost,tldrx cost --storiesputs each story beside the ceiling its spawn was given, and a ceiling that is getting close arrives as abudget.warnednotification with both numbers in it.
A first-run checklist
Declare
install:in.tldrx/workspace.yml, for every repo whose test command needs installed dependencies. A Build story runs in a freshgit worktree: it has your tracked files and nothing else — nonode_modules, no virtualenv, no restored packages, and none of your checkout's. tldrx runs the declared installer there before the developer and records it as its own check with an exit code and a duration. Without it, the story's DoD exits127and blocks, having already paid for a turn; the message names the absent binary and this slot, and the framework guesses no installer for you.yamlrepos: - name: app commands: install: "npm ci" # pnpm install --frozen-lockfile, uv sync, dotnet restore, … test: "npm run test"Declare
test_fastin.tldrx/workspace.yml— the fast subset the Build developer iterates on. It is not a Definition of Done command; the DoD re-runstest:.Write the adapter and declare it under
notify:. Start with every kind — narrowevents:later, once you know which ones you actually want waking you.Dry-run the adapter by hand, before any run depends on it:
bashecho '{"version":1,"kind":"status","at":"2026-01-01T00:00:00Z","run":"demo","root":"'"$PWD"'","stage":null,"summary":"hello","command":null,"detail":{"status_text":"hello","waiting_on":[]}}' | bin/notify-ownerIf that does not reach you, nothing will.
Launch with a short interval first —
--notify-every 60sfor one stage — so you find out the hook works while you are still at the keyboard. Then raise it.Watch
tldrx run statusfor the run's own view, andtldrx replay <run>for the event log as a narrative,notify.sent/notify.failedincluded.
Troubleshooting
A story blocked on exit 127, "command not found". The tests never ran: the story's worktree did not have the binary. Declare install: (item 1 above) and tldrx installs the dependencies there before the developer. The check that failed carries tree: "worktree", so it can be told apart from the Build-entry pre-flight, which runs the same command in your own checkout — where the dependencies already are, which is why it can be green minutes earlier.
The developer says "This command requires approval to run". Fixed in gh #209: every declared command is now granted both exactly and with trailing arguments, so a developer can run npm run test -- one/file.test.ts while it works instead of only the bare command.
The notifier is never called. Three usual causes, in the order they cost the least to check. The events: list does not name the kind you were expecting — remove the key entirely to subscribe to everything. The command is not executable, or is not on the path the run resolves it from — bin/notify-owner is relative to the workspace root, and it needs its executable bit. Or the command line contains a shell metacharacter and was refused rather than shelled: move the pipeline into a script and declare the script.
A malformed notify: block reads as no block at all. The reader will not produce a hook the validator would reject, so a broken block notifies nothing rather than spawning something unchecked. tldrx doctor is where a bad block is reported.
notify.failed in events.jsonl. The event carries the reason in words — needs a shell, could not be started, no such executable, timed out after N ms and was killed, or exit <n> with a tail of the child's output. Read it with tldrx replay <run>. Whatever it says, the run's own outcome is unchanged.
--wait-answers lapsed. You get one question.timeout carrying waited_ms and the same questions[], and then exit 4. Answer the question and start the loop again; nothing was lost and nothing was spent while it waited.
--wait-gates lapsed. You get one gate.timeout carrying waited_ms, the approve and reject lines and the gate's policy, and then exit 4. Sign or reject the gate and start the loop again. If the gate is on gates_policy: agent and you expected the run to carry on by itself: it will not — the loop signs nothing, and an agent gate says who MAY sign, not that anything has.
A dirty checkout no longer stops the run. The Build entry classifies every uncommitted path instead of counting it. Anything under tldrx-work/, .tldrx/ or .agent/ is the framework's own state and is ignored; a dirty path a pending story declares in its touches:, or a submodule, still refuses with exit 2; everything else is set aside in a pathspec-limited git stash push before the epic branch is cut, and popped back when the stage ends. Both moments are on the log (worktree.foreign_work_aside, worktree.foreign_work_restored) and nothing is ever deleted or force-popped. If git refuses the pop — because the tree changed that path meanwhile — the stage's LAST line, the handoff's ## Unknowns and the stage.done / run.finished notification all say foreign work NOT restored, with the stash and the literal command to take it back. The run's exit code does not move for it.
A repo in the middle of a merge or rebase refuses (exit 2). That state has no clean undo, so nothing is stashed into it. Finish or abort the operation and start the loop again.
The run is refused with exit 1. run auto will not run on a run marked attended_by: host — a lock and an engine are alternatives, never layers. Hand the run back to the framework with tldrx run attend --none <run>, or drive it from a session instead. (tldrx run attend host <run> is the other direction.)
A budget refusal (exit 2). --max-usd is checked between stages, so the loop can overshoot by at most one stage's share. tldrx budget show says what is left; raising a ceiling is a person's decision, and there is no flag that makes it not one.
The conceptual half — why attend and auto are opposites, and the mandate for driving a run from a session — is in Attended or unattended. The full chapter, including the four ways an agent gate falls through to a person, is 10 — Unattended mode.