Everyone wants end-to-end tests, nobody wants to maintain them — the real cost isn't typing the test, it's knowing what to assert and keeping it alive as the UI shifts every sprint. Martin hands both jobs to an AI agent live on stage: exploring the app, generating Playwright tests in a real browser, and repairing them after he deliberately breaks the UI. The second half drops the demo and measures instead — four setups, one real product, and an analysis plan written before the first run.