The handoff doc is the real product
I had a version of this workflow sitting in a folder since May doing nothing. In July I tore it down and rebuilt all three skills. Here is everything I learned that applies to anyone running multi-phase work with agents.
The handoff doc is the real product of each phase.
Research does not hand the next phase 40 files of notes. It hands over roughly 200 lines. Plan reads only that document. A fresh thread reading one 200-line doc starts at 10 to 15 percent context usage instead of dragging 60,000 tokens of research chatter forward.
That is where the cost savings live. But it's a quality play before it's a cost play, because the failure mode you're avoiding has a name: context pollution. A thread that's been arguing with itself for two hours has opinions it can't justify and details it half remembers. Starting clean with a curated summary produces better output than continuing with everything.
Checkpoint edits are the highest-leverage thing a human does.
Reviewing 400 lines of specs beats reviewing 2,000 lines of generated output, every time. The research handoff deserves the most attention, the plan second, and you can skim the build reports. Most people do this exactly backwards: rubber-stamp the plan, then agonize over the output, which is the expensive end to be picky at.
Every phase needs a "done when" the agent can check itself.
Otherwise "looks done" is the only available signal, and "looks done" is not a signal. It doesn't have to be a test suite. "The sheet reconciles to the export" counts. "The doc answers these four questions" counts. Then a fresh subagent reviews the result adversarially, because fresh eyes have no attachment to the work. Attachment to your own output is not a human-only problem.
Subagents do the noisy reading.
The main thread never reads 40 files. It reads summaries of 40 files. This one change is most of why the whole thing stays coherent past the first hour.
A state file makes any thread resumable.
One small file per project holding where things stand. It sounds trivial until you come back three weeks later and the alternative is reconstructing your own reasoning from a transcript.
How I tested it
Three live runs, including a planted veto buried inside a handoff document and a planted contradiction inside a plan, to find out whether the mechanics would catch them or cheerfully carry them forward. They caught both. Six ambiguities surfaced along the way and got patched before I packaged anything.
The question I still can't answer: is 200 lines the right cap for bigger projects, or does it need a per-project override? I don't know. It has been right for everything I've run so far, which is not the same as being right.
Adapted from HumanLayer's frequent intentional compaction and Anthropic's Explore, Plan, Code guidance, generalized for people who aren't shipping code.