Security researcher cekuu35 uncovered a staggering volume of exposed credentials in public GitHub repositories, finding over 60 live Supabase service_role keys in just three days. These keys, which bypass Row Level Security and grant full database access, were found in everything from config.php files to Docker Compose templates. The discovery sparked a critical question: if developers are increasingly relying on AI assistants to write this code, can the models themselves identify the security risks they help create?
The Benchmark Setup
To answer this, the researcher constructed a Kaggle Benchmarking Challenge task titled committed-supabase-key-detection. The test evaluated nine models, including Anthropic's Claude Sonnet 5, Google's Gemini 3 series, OpenAI's GPT-5.4-nano, and DeepSeek-R1. Each model was presented with ten realistic file snippets containing various credential leak patterns, ranging from obvious hardcoded keys to subtle NEXT_PUBLIC_ prefixes that bundle secrets into client-side JavaScript. The models were required to output a strict JSON verdict, classifying the presence of service_role keys, anon keys, and database passwords, with a score based on full correctness.
Frontier Models Ace the Obvious, Fail the Subtle
The results highlighted a clear divide between raw detection capability and nuanced security triage. Frontier models like Claude Sonnet 5 and Gemini 3.7 Flash achieved perfect 10/10 scores, correctly identifying even the most obfuscated threats. Notably, Gemini 3 Flash Preview demonstrated advanced reasoning by decoding a base64-obfuscated blob mid-response to identify the role claim. However, the budget-friendly GPT-5.4-nano struggled with edge cases, missing a commented-out key and failing to decode the base64 string, proving that cheaper models still lack the depth required for complex security audits.
The Danger of Alarm Fatigue
Perhaps the most telling failure came from DeepSeek-R1, a dedicated reasoning model that scored 0/10. Rather than missing keys, it flagged every single file as critical, including those containing only placeholders. In a real-world security workflow, a reviewer that cries wolf on every commit creates alarm fatigue, leading teams to ignore genuine threats. This underscores a critical lesson for AI integration in DevSecOps: high recall without precision is effectively useless. A model must distinguish between a live secret and a documentation example to be a viable security tool.
Key Takeaways
- Precision Matters: DeepSeek-R1's 0/10 score due to over-flagging demonstrates that false positives are as damaging as missed secrets in security triage.
- Frontier Dominance: Claude Sonnet 5 and Gemini 3 Flash tiers achieved perfect scores, successfully handling base64 obfuscation and complex context.
- Edge Cases Exposed: Cheaper models like GPT-5.4-nano failed on subtle leaks like commented-out keys and encoded secrets, highlighting the gap between cost and capability.
- Human Error Persists: The benchmark confirms that while models *can* spot leaks, the primary failure point remains the lack of automated review steps in the developer workflow.
The Bottom Line
AI assistants are technically capable of catching committed secrets, but their utility is nullified by poor precision in reasoning models and the human failure to actually ask them. Until developers integrate mandatory, high-precision AI review steps into their CI/CD pipelines, the 15-minute manual key rotation remains the most reliable security control.