Static AI Review vs. Live Agentic Browsing: What 41 Seeded Bugs Actually Show
I built a deliberately buggy e-commerce app with a hidden 41-fault answer key, then scored two different automated exploratory-testing architectures against it: a static deterministic-scan-then-AI-review pipeline, and a live agentic browser driven by Claude through the real Playwright MCP tools. The two methods disagree on 16 of the 41 bugs - and only 4 were missed by both.