On YouTube

Not AI news.
Agents at work.

Every episode is an experiment: real agents on a real codebase, measured, with an honest account of what they got right, what they got wrong, and what it would take to trust them.

In production

The episodes I'm filming.

None of these are published yet. Subscribe to get each one when it lands, with the checklist behind it in the description.

  1. 01

    Can an AI developer pass QA without a human fixing its work?

    A coding agent ships a feature. A strict QA agent reviews it. Nobody steps in. How much survives, and why does the rest fail?

    Coming soon
  2. 02

    Five agents, one feature: where does coordination break?

    Five agents on one codebase at the same time. Who overwrites whom, what conflicts, and what it takes to merge.

    Coming soon
  3. 03

    Does an agent follow your architecture when nothing enforces it?

    The rules are in the prompt. Nothing checks them. Does the work pass every test and still break the architecture?

    Coming soon
  4. 04

    Instructions aren't permissions. What does an agent do with access it shouldn't use?

    The prompt says don't. The tools say it can. Which one wins, and how often?

    Coming soon
  5. 05

    Can you audit what an AI agent did six months later?

    Pick a merged change from an agent and try to reconstruct who asked, who checked, and why it was allowed through.

    Coming soon
  6. 06

    Can an AI software team maintain the same app for 30 days?

    Shipping is day one. Bug reports, dependency updates and real users are the other 29.

    Coming soon
  7. 07

    Can an AI agent remember the app it built yesterday?

    A fresh session against a maintained project memory. Does the agent stop repeating its mistakes?

    Build along with the brief
    Coming soon
  8. 08

    Which agent is best at which job? Measuring agents like a team

    Companies measure developers. Almost nobody measures agents. Same tasks, several agents, real scorecards.

    Coming soon
Visit the channel