The AI agent community is being invited to a very specific, very weird test of competence. Developer eshevtsov has issued a challenge on DEV.to for anyone with a coding agentβ€”Claude Code, Codex, Cursor, or otherwiseβ€”to build a procedural octopus that unscrews a jar lid from the inside using Three.js. The catch? It must be a single, self-contained HTML file, and every run will be posted side-by-side for public scrutiny.

The Octopus Test

The prompt is deceptively simple on the surface but brutal in execution. Agents must generate a 3D scene featuring a glass jar on a wooden table, filled with water, containing a fully procedural octopus. The core behavior isn't just animation; the octopus must physically grip the lid, apply torque, and rotate it along the thread pitch for 3-4 turns before escaping. The output must handle inverse kinematics, squash-and-stretch physics, and real-time interactivity, all while hitting 60 fps on a mid-range laptop without external assets.

Methodology Over Models

This isn't just about which model generates the prettiest tentacles. The experiment aims to answer a persistent question in the agentic coding space: does the model choice matter more than the workflow? Eshevtsov is splitting participants into two groups based on their birth date. Prompt A is the base challenge. Prompt B adds a single paragraph instructing the agent to build a validation harness first, exposing phase names and lid angles to the window object for step-by-step debugging. The goal is to see if early, structured self-checking reduces bugs more effectively than switching from, say, GPT-4 to Claude 3.5.

Community Data and Transparency

Participants are required to use the KeepPlain plugin or upload session files to keepplain.com, ensuring that token costs, intervention points, and time-to-completion are logged. Eshevtsov warns agents to disconnect from any global memory or personal rules that might leak previous successes. The final analysis will pool data from all runs to chart the minute of the first real check against the number of bugs remaining at the end. If Prompt B runs consistently outperform Prompt A runs regardless of the model used, it suggests that agentic coding is less about the LLM's raw intelligence and more about the structure of the prompt's feedback loop.

Key Takeaways

  • Agents are being tested on a complex physics-based visual task that requires iterative debugging, not just code generation.
  • The experiment isolates the impact of prompt engineering (adding a self-validation step) versus model selection.
  • All runs will be published with full transparency on token usage, human interventions, and final output quality.
  • Participants must use their birth date to determine their prompt variant, preventing selection bias.

The Bottom Line

This octopus challenge is a rare moment of genuine scientific rigor in the hype-driven world of AI agents. It proves that in agentic coding, the workflow often beats the model, and that early validation is the single most effective way to prevent hallucinated physics.