Security tooling company Aikido Security published comprehensive benchmarks on August 21, 2026, testing multiple leading AI models against cybersecurity-specific tasks using a dataset of 11.7 billion tokens. The study aimed to cut through marketing noise and give practitioners actual data on which LLMs perform reliably for security work.

Why Token Volume Matters

The choice to run 11.7B tokens isn't arbitrary—it provides statistical significance across diverse attack vectors, vulnerability types, and defensive scenarios. Smaller benchmark sets can be gamed or skewed by training overlap; this scale forces models to demonstrate genuine reasoning rather than pattern matching on popular datasets.

Methodology Overview

According to the Aikido blog post shared on Hacker News, the team constructed evaluation prompts based on real penetration testing findings, CVEs from the past three years, and practical SOC analyst workflows. Models were tested blind, with evaluators comparing outputs against known correct answers without knowing which model generated them.

What the Results Show

The headline claim—that 11.7B tokens worth of security-focused evaluations reveal clear performance hierarchies—suggests meaningful differentiation between models on tasks like exploit analysis, incident response guidance, and vulnerability prioritization. The full breakdown appears to include comparisons across multiple provider families, though specific rankings require reading the complete article.

Caveats Worth Noting

Benchmark results in AI move fast. What ranks highest today may shift with model updates, system prompts, or temperature settings. Readers should treat these findings as a snapshot rather than definitive verdicts—especially for production security tooling where false negatives carry real risk.

Key Takeaways

  • 11.7B token dataset provides statistical weight rarely seen in AI security benchmarks
  • Real-world CVE data and penetration testing results informed the evaluation criteria
  • Results likely show meaningful performance gaps between leading models on security tasks
  • The study was conducted by Aikido Security, a company with obvious commercial interest in AI tooling

The Bottom Line

This kind of thorough, large-scale benchmarking is exactly what the security community needs—but always read the methodology closely. Vendor-sponsored benchmarks deserve scrutiny even when the token counts are impressive.