Videos
Is There One Best AI Model for Hacking? | Offense Taken Ep. 02
When Frontier Models Learned to Hack
The first half of 2026 saw AI models take an enormous leap forward in offensive security. So which one is the best hacker? In the second episode of Offense Taken, XBOW's Fede Kirschbaum, Albert Ziegler, and Ed Spencer unpack XBOW's Mid-Year 2026 AI Model Security Research Report—evaluations of GPT-5.5, Mythos Preview, Opus 4.7, GLM-5.2, Muse Spark 1.1, and Grok 4.5 in real offensive security workflows.
The conversation explores:
- Why GPT-5.5 performs better "blindfolded," without source code access
- Why Mythos excels at deep source-code analysis—and why an excellent "brain" still needs a capable "body" for live testing
- How capable budget models are rewriting the economics of hacking
- What the Grok evaluations revealed about cost, efficiency, and safety guardrails
- Why traditional benchmarks fall short of live-system testing
- Why no single model wins—and how XBOW's alloy approach combines their strengths
As model capabilities keep advancing, security teams should stop searching for one "best" model and start evaluating the orchestration and safety systems built around them.