A recent Hacker News post details a fascinating, low-stakes experiment where a developer attempted to use Anthropic's Claude models to play and win two obscure 1990s strategy games: Panzer General II (PzGII) and Imperialism II (ImpII). The project, documented on replicated.live, highlights the current limitations of Large Language Models in handling complex, probabilistic, and spatially dependent decision-making tasks. While Claude successfully reimplemented the game engines in JavaScript and C, its actual gameplay performance remained mediocre, ultimately failing to defeat the final boss in PzGII and struggling with long-term planning in ImpII.

The Setup: Obscure Games to Avoid Training Data Contamination

The author deliberately chose PzGII and ImpII over mainstream titles like Civilization II to ensure the models could not rely on memorized strategies from their training data. The hypothesis was that these games, with their dense rulebooks and complex interactions, would present a verifiable challenge: winning. The budget was tight, utilizing weekly token leftovers on claude.ai, a free quota reset, and a modest €20 Hetzner cloud account. The author spent the entire budget plus half a week's worth of the $200 Max plan to run these experiments.

Panzer General II: Bureaucratic Caution Dooms the Campaign

For PzGII, the author instructed Claude to build a "LLM wiki" of game rules and reimplement the DOS binaries. While the code generation was flawless, the AI's gameplay was hindered by a lack of spatial logic and long-term planning. The author employed six iterative improvements, including accumulating lessons in Markdown field manuals and applying historical military concepts like Heinz Guderian’s *Auftragstaktik*. Despite reaching the "final boss" stage (the invasion of the US), the AI failed to secure a "Brilliant Victory" at Dunkirk. The analysis revealed that Claude’s accumulated expertise led to an overly cautious, bureaucratic playstyle that avoided immediate disasters but allowed long-term objectives to slip, a flaw inherent to the method of recording "hard-earned lessons" in text.

Imperialism II: Evolutionary Loops and Economic Solvers

In Imperialism II, the author attempted to use Claude to implement solvers for different game planes—logistics, resources, and diplomacy—inspired by Leontief’s input-output analysis. Claude excelled at writing Linear Programming (LP) code and discussing economic models but failed to integrate these solvers into a coherent strategic decision-making process. A more successful approach involved an evolutionary loop where five Claude agents competed in tournaments against a base version. Each tournament, running on a 16-core Hetzner box, took about an hour to average results across 720 possible player seatings. This method produced steady, albeit slow, improvements in heuristic AI performance, paralleling techniques seen in Google’s AlphaEvolve project.

Key Takeaways

  • Spatial Reasoning Gap: Claude struggles with the high-dimensional combinatorics of strategy games, lacking the intuitive spatial logic required for effective unit movement and terrain exploitation.
  • Markdown Limitations: Relying on text-based "field manuals" for learning leads to risk-averse, bureaucratic behavior rather than creative or aggressive strategic adaptation.
  • Evolutionary Improvement Works: Running multiple agents in competitive tournaments with automated scoring provides a scalable path for incremental AI improvement, even on a tiny budget.
  • Code Generation vs. Gameplay: LLMs are exceptionally good at reimplementing game engines and writing solvers but currently fail to translate that code into winning human-like strategies in complex simulations.

The Bottom Line

This experiment proves that while LLMs can write the game, they cannot yet play it at a competitive level. The reliance on text-based memory creates a safety-first paralysis, and true strategic dominance requires spatial intuition that current token-based models simply do not possess.