Define what ready means

An agent can produce a plausible implementation faster than a team can inspect every line. That makes the acceptance process a central part of factory design. Before increasing the amount of work in flight, decide what evidence will establish that a result meets its requirements. A successful run needs a result that can be evaluated independently of the worker's confidence.

Different products need different evidence. A protocol adapter can be checked against request and response fixtures. A desktop component needs interaction checks and visual references. A media export needs valid frames, correct timing, and a person who can judge the edit. A mathematical claim may admit a machine-checked proof. Each verifier answers a particular question; together they support an acceptance decision.

Make the check close and decisive

Give workers a fast way to learn that an attempt is wrong. A focused test, a schema validator, or a reproducible visual comparison can turn a vague problem into a specific correction. Good feedback identifies the violated requirement and provides enough context to act. The next attempt should start with new information.

Build a small harness around the actual boundary you care about. For an agent runtime, that may mean scripted model responses driving real file tools in a temporary workspace. For an email service, it may mean a request sequence that checks message threading and event delivery. Exercising the boundary helps catch integration mistakes that isolated assertions can miss.

Test the wrong thing too

Positive examples establish that a system accepts valid work. Negative examples establish where it should stop or refuse an input. Both matter. An importer should handle a valid record and identify a malformed one. A permission boundary should allow the intended participant and reject another. A quality detector should recognize a defect without flagging the known-clean case.

Include failures in the machinery around the task. Interrupt a write, repeat an event, lose a connection, or resume from an old checkpoint. The useful question is what state remains and whether work can continue correctly. A factory repeatedly encounters these conditions as it gets larger. Recovery behavior deserves an explicit place in its checks.

Attach evidence to the thing it describes

A passing check describes a particular artifact under particular conditions. Record the code snapshot, command, configuration, and result needed to understand that claim. When the relevant implementation changes, the evidence may need to be refreshed. This is especially important when one agent implements, another reviews, and a later session prepares the release.

Keep the scope of each claim legible. A unit test can establish a function's behavior on its cases. A successful production request can establish that a deployed path worked at that time. Neither automatically answers every other question about the system. An evidence record is useful when a reviewer can see both what it supports and what still needs examination.

Make room for judgment and repair

People still decide whether a feature solves the right problem, whether an interaction feels understandable, and whether an important tradeoff is acceptable. Build review surfaces that make those judgments concrete. Show the running preview, the changed behavior, the remaining uncertainty, and the checks already performed. Record the decision alongside the work.

Finally, use defects to improve the factory. A repeated failure is a candidate for a reusable fixture, a stronger contract, or a changed production step. Challenge the checks with known bad results and confirm they catch them. The system becomes more useful when each run improves its ability to recognize the next good result.

More notes from the workshop