A busy agent is not a measure of progress. Collect the events that explain a run: tool calls, changes, errors, time, and resource use. Connect those events to the code and artifacts they produced. Local session history, replay, and repository inspection make it possible to diagnose failures, compare approaches, and recover useful work from an abandoned attempt.
What this makes possible
- Session ingestion across different coding-agent providers
- Cost, token, timing, and outcome inspection
- Step-by-step change history and filesystem playback
- Repository, branch, and worktree inventory
On the production floor
Compare two runs that reached the same patch. Find the repeated tool failures that made one take longer, then improve the factory's setup.
How you know it works
Use fixture logs with known events and totals. Preserve source references so a displayed number can be traced back to the run that produced it.