NewsTradingSentimentEventsCommunityBriefing
Tech

XRanges Scores AI Security Agents Using 545 Prior Hacker Tests

By Tech Desk · · 1 min read
A sleek server rack with glowing status lights indicating active processing.

New platform uses instrumented targets to verify if AI agents actually find real bugs, replacing manual review with live metrics.

Key points

  • XRanges for AI uses instrumented, realistic targets to score security agents on four independent metrics.
  • The platform tracks coverage and validity to distinguish real exploits from hallucinations and destructive errors.
  • The system was validated by 545 professional hackers to ensure the benchmark is resistant to gaming.

Autonomous security agents are becoming proficient at identifying software vulnerabilities, yet a reliable method to measure their accuracy remains absent. When developers point these AI tools at realistic targets, the output is often a confident report that lists findings without distinguishing between genuine exploits, accidental touches, and outright fabrications.

This creates a significant bottleneck for engineering teams. Verifying each claim requires a human expert to manually inspect the target, a process that becomes unmanageable when testing multiple models, prompt variants, and repeated runs. The result is a review queue that grows longer than the experiment itself, making consistent comparison of AI performance nearly impossible.

Instrumented Targets Replace Manual Guesswork

XRanges for AI, developed by CTF.ae, addresses this issue by deploying realistic applications with deep instrumentation. Unlike static puzzle challenges, these targets are complete multi-service applications with business logic, background jobs, and simulated user traffic. Each target contains over twenty injected vulnerabilities, including zero-days discovered by the platform's own researchers, ensuring the test scenarios do not appear in public training data.

Four Metrics Track Agent Behavior

The platform scores every run live using four independent signals. The first metric, coverage, measures how thoroughly the agent explores the target through legitimate user actions, such as registering an account or browsing content. By tracking which features were accessed and which were ignored, the system identifies blind spots in the agent’s exploration strategy.

The other metrics verify the validity of findings and check for destructive side effects, such as data deletion or key revocation. Because these signals are independent, an agent cannot improve one score by gaming another. This structure allows developers to compare the performance of different AI models objectively, rather than relying on subjective reading of self-generated reports.

Stress Tested By Human Hackers

The reliability of the benchmark was established before public release. According to The Hacker News, 545 hackers tested the system first to validate its integrity. This human validation ensures that the vulnerabilities and scoring mechanisms are robust enough to withstand professional scrutiny, providing a trusted baseline for evaluating autonomous security tools.

Based on reporting by The Hacker News, compiled by the Tradingbird desk.

Read next

More in Tech

More from the Tech desk

All desk stories