Benchmark Scores Are a False Flag

  • ProjectDiscovery
Benchmark Scores Are a False Flag

About This Webinar

Nearly every claim about a model's cybersecurity capability rests on a solve rate. The number counts how many challenges an agent finished, but records nothing about what the agent did to finish them. That gap hides what matters most.

A solve rate can't tell you whether a model didn't know the exploit or knew it and failed to execute. It can't tell you whether the agent exploited the vulnerability the challenge was built around or found an exposed credential and took the flag with that instead. And it says nothing about whether the agent stayed inside the target.

Tarun Koyalwar, AI Security Researcher at ProjectDiscovery, ran open and closed models against 54 black-box web targets. He gave them no source code, no hints, and no methodology. Then he read every run by hand instead of scoring it.

This session covers what he found and the harder question underneath it. The industry already knows these benchmarks are saturated. The problem starts earlier, in how we build them. If you write the test from an answer key and work backward, what are you measuring?

What You'll Learn

  1. Why a solve rate is a false flag, and the four dimensions worth measuring in its place
  2. How answer-key benchmark design produces challenges that hand over the flag with no exploitation at all
  3. What the run-by-run numbers show, including the loaded harness that underperformed an empty one and what repeat runs do to a reported score
  4. Why these benchmarks reward an agent for breaking out of scope, and how that connects to the Hugging Face incident

Speakers

  1. Scott Bekker Webinar Moderator Future B2B
  2. Davis Franklin Head of Technical Solutions ProjectDiscovery
  3. Tarun Koyalwar AI Researcher ProjectDiscovery
WIN A $250 Amazon Gift Card

REGISTER NOW & YOU COULD WIN

A $250 Amazon.com Gift Card!

Must be in live attendance to qualify. Duplicate or fraudulent entries will be disqualified automatically.