The team behind Jevman has launched an open-source benchmark that pits AI models against the classic arcade game Pac-Man. By having six distinct AI agents play 100 games each against scripted ghosts, the project provides a tangible, real-time metric for simple decision-making capabilities. The leaderboard currently lists 'jev 1.130' as a participant, but the framework is designed to accept any model accessible via an HTTP endpoint.
Benchmarking Methodology and Constraints
The testing framework is rigorous and transparent. Each model faced the classic ghosts under strict conditions: a two-second deadline for every decision, three lives per game, and a five-minute cap. If a model failed to respond within the two-second window, a simple backup rule took over, with these instances tracked separately as backup moves. The games are capped to ensure completion, though the longest recorded session lasted only two minutes and 24 seconds, indicating that current models struggle to sustain high-level play for extended periods.
Open Architecture for Custom Models
Jevman is designed to be inclusive of any model accessible via an HTTP endpoint. At every junction, the game stateβincluding maze layout, pellets, and ghost positionsβis sent as JSON. The model returns directional probabilities, and Pac-Man acts on the highest probability. For developers wanting to test their own fine-tunes or local models, the repository provides a 34-line example endpoint. This low barrier to entry allows for rapid integration, whether the model is hosted in the cloud or running on a local laptop.
Statistical Rigor in Leaderboard Rankings
Unlike many AI benchmarks that rely on static datasets, Jevman uses dynamic gameplay. Rankings are determined by mean scores with a 95% margin of error (Β±2 standard errors). Models falling within each other's margin of error are tied, preventing false precision in performance comparisons. Every game is recorded and can be replayed exactly, ensuring that any claim can be audited. This transparency addresses the common criticism of AI benchmarks being opaque or non-reproducible.
Key Takeaways
- Jevman tests six models against classic arcade ghosts, with 'jev 1.130' explicitly listed in the source snippet.
- Models must respond within two seconds; failures trigger a backup rule and count against performance.
- The benchmark is open source, allowing any HTTP-endpoint model to join the leaderboard via pull request.
- Scores are statistically tied if they fall within a 95% margin of error, emphasizing reliability over marginal gains.
The Bottom Line
While Pac-Man might seem like a gimmick, it exposes the latency and reasoning gaps in these new decision endpoints better than static benchmarks ever could.