AI Hacking Agent and Human in the Loop with Diego Jurado | Critical Thinking
Inside XBOW’s Autonomous AI Pentesting Platform—and the Human Ingenuity Still Driving Elite Exploit Chains
How close is AI to performing the work of an elite penetration tester? In this episode of Critical Thinking, hosts Justin Gardner and Joseph Thacker sit down with Diego Jurado, security researcher at XBOW and one of the world's top bug bounty hunters, for an in-depth discussion on the future of autonomous offensive security.
The conversation begins with a fascinating walkthrough of a sophisticated account takeover vulnerability discovered during HackerOne's Ambassador World Cup. Diego breaks down how multiple seemingly unrelated weaknesses, including API downgrades, JSONP behavior, referer validation, and an Adobe Experience Manager XSS, were chained together into a complete account takeover. The discussion highlights why complex exploit chains remain one of the biggest challenges for both human and AI hackers, while also showcasing the power of collaboration among elite researchers.
The episode then shifts to XBOW's autonomous AI pentesting platform, where Diego provides an unusually transparent look inside its architecture. He explains how autonomous "solvers" pursue individual attack objectives under the guidance of a coordinating agent, using reasoning, custom Python tooling, browser automation, validation systems, and out-of-band interaction servers to discover and confirm real vulnerabilities. The hosts dive into the practical realities of AI hacking, including hallucinations, prompt engineering, context management, validation pipelines, false positives, and how XBOW balances autonomous reasoning with strict verification before reporting findings.
Beyond the technical deep dive, the discussion explores broader industry implications. The hosts debate whether AI will replace bug bounty hunters, the continued value of human creativity in discovering complex exploit chains, and why validation, not vulnerability discovery, may become the defining challenge for autonomous security systems. Diego also shares XBOW's philosophy of full autonomy, discusses how the company uses HackerOne as a large-scale training ground, and explains why lowering false positives remains a higher priority than maximizing finding volume.
The episode concludes with a candid conversation about HackerOne's Ambassador World Cup, covering team collaboration, competition incentives, duplicate findings, and why Spain has become a dominant force through disciplined teamwork and knowledge sharing.
For anyone interested in AI-driven offensive security, bug bounty hunting, or the future of autonomous penetration testing, this episode offers one of the most detailed and technically grounded conversations available on how frontier AI systems are beginning to think, reason, and exploit software like human hackers.
Speakers
Security Researcher | XBOW @ XBOW