The build lab

What happens if
we actually build it?

What do AI agents actually do when they work on real software? Each experiment here is a real question with a real build behind it, and each one becomes a video.

The project briefs are ready to use. Recorded episodes and measured results will be added when they’re published.

Up next

Can AI agents work like a real team?

The next experiments test the idea Bookbag is built on: agents need a system around them, not just a better prompt.

  1. 04

    Can an AI developer pass QA without a human fixing its work?

    A coding agent ships a feature. A strict QA agent reviews it. Nobody steps in. How much survives, and why does the rest fail?

    In planning
  2. 05

    Five agents, one feature: where does coordination break?

    Five agents on one codebase at the same time. Who overwrites whom, what conflicts, and what it takes to merge.

    In planning
  3. 06

    Does an agent follow your architecture when nothing enforces it?

    The rules are in the prompt. Nothing checks them. Does the work pass every test and still break the architecture?

    In planning
  4. 07

    Instructions aren't permissions. What does an agent do with access it shouldn't use?

    The prompt says don't. The tools say it can. Which one wins, and how often?

    In planning
  5. 08

    Can you audit what an AI agent did six months later?

    Pick a merged change from an agent and try to reconstruct who asked, who checked, and why it was allowed through.

    In planning
  6. 09

    Can an AI software team maintain the same app for 30 days?

    Shipping is day one. Bug reports, dependency updates and real users are the other 29.

    In planning

Want the checklists behind the experiments?

The agent rollout checklist, agent brief and release checklist are free.

Get the free resources