Factory Beyond the Ticket Loop

Since the first Factory build log, the ticket-to-PR loop has grown a human handoff, an external review lane, and a way to inspect what happened without trusting the agent's account of it.

· 4 min · agents / developer-tools / build-logs

Last time I wrote about Factory, it was a week old. The useful trick was straightforward: put a coding agent behind a GitHub issue, make it run the repository's checks, review its PR, and let a guarded merge stage decide whether the result could land. Fourteen of Factory's own tickets had made it through. The sharp edges were already showing: a hung gate, a branch refresh that erased a PR, and the question of what happens when the machine cannot finish the job.

Three weeks later, the loop still exists. Most of the work since then has been around it: deciding what not to run, making evidence survive a run, and getting a useful question to a person when automation stops.

A plan is not a ticket

Factory now distinguishes an initiative from an executable issue. An initiative describes an outcome and its boundaries; the smaller tickets carry the work. The initiative guard refuses to send the planning issue to triage, dispatch, management, or merge. That sounds fussy until you imagine a worker treating “make the product better” as a task it can complete in one PR.

A linked ticket can carry a snapshot of the initiative's accepted plan. On retries, Factory checks that binding rather than quietly substituting whatever the plan says today. If the source changes or disappears, it stops and asks for a decision. The roadmap view shows the accepted plan alongside the live one and makes drift visible. It does not turn a plan into permission to execute.

The handoff is part of the work

The first post described ready-for-human as the end of Factory's rope. It was a label and an escalation packet, but somebody still had to work out who should answer, what they needed to decide, and whether an earlier attempt to ask had actually reached GitHub.

The routed handoff now posts a bounded question with evidence links and a proposed next step when automatic recovery is exhausted. It records the intent and the resulting comment so a crash between the two does not produce an endless stream of duplicate pings. Routing finds a decision owner; it does not assign a worker or pretend the human has agreed. That distinction matters more to me than another autonomous retry.

The other uncomfortable boundary was proof of work. A worker's worktree is temporary; once cleaned up, its account of what it did can be hard to inspect. Accepted handoffs now survive worktree cleanup, tied to the exact PR head with size and retention limits. Missing or incomplete evidence stays missing, not reconstructed from a confident summary.

Factory can review somebody else's PR

The original pipeline owned the branch it reviewed. The external review lane adds a narrower job: discover a contributor PR that opts in, publish findings against the head that was actually read, re-review when that head changes, and show CI readiness. It does not edit the contributor's branch or merge the PR. The distinction between “reviewed” and “ready to merge” is not cosmetic; a green review of an old head is no review of the new one.

There is now a read-only Factory Manager console and a dashboard that can show runtime state, PR feedback, roadmap questions, and retained results. The console gets bounded evidence reads, not a shell or permission to dispatch. It is useful for asking “why is this ticket here?” without granting the answerer the power to change the answer.

The current experiment

The recent work is less about adding stages than checking whether our judgments hold up. Factory has a reviewer-calibration corpus and adjudication rows for comparisons. There is also a bounded observation experiment over recorded ticket outcomes. These are experiments and recorded evidence, not proof that the reviewer is calibrated or that the system has learned to manage a project. The roadmap can show an experiment alongside a proposed next action without making the experiment a gate.

The 0.3.5 release is the last tagged release I am describing here; the observation and calibration work is on main after that tag. A launch prefix for workers is not a sandbox. The manager still runs a configured executable that Factory does not sandbox. And a dashboard showing a plan does not mean somebody has accepted the outcome.

The change since September 9 is not that the agents became more trustworthy. We have made it harder to mistake a plan for a ticket, a stale review for approval, or an escalation label for a completed human handoff. That is less flashy than “tickets in, PRs out.” It is what lets the loop keep running without pretending the edges aren't there.