Michael J. Klaiber published a technical critique on September 27, 2026, challenging the growing narrative that Large Language Models (LLMs) will replace AI compilers. The article, titled "LLMs Will Not Replace AI Compilers. They Will Call Them," argues that while LLMs are powerful, using them to perform compilation tasks directly within their weights is computationally inefficient and prone to verification issues. Instead, Klaiber posits that the future lies in LLMs orchestrating existing compiler infrastructure, much like they currently invoke external tools such as ffmpeg.

The Efficiency and Verification Problem

Klaiber highlights a fundamental disparity in computational cost. A traditional compiler’s memory planner or tiling search executes in milliseconds or seconds using deterministic algorithms. In contrast, an LLM attempting to generate the same optimized code token-by-token consumes billions of floating-point operations per token. The author notes the irony that LLMs themselves are among the most heavily compiled programs, relying on the very fusion and quantization techniques they are purported to replace. Furthermore, compilation correctness exists on a spectrum from bit-exact to numerically close. Introducing a non-deterministic LLM into this chain makes debugging accuracy regressions significantly harder, as the model cannot reliably explain its decisions or reproduce results consistently.

Orchestration Over Emulation

The proposed solution is a shift from emulation to orchestration. Klaiber suggests that LLMs should function as intelligent agents that drive compiler passes, autotuners, and profilers. This approach mirrors how LLMs handle video encoding today: they do not emulate the encoder but instead construct command-line arguments for tools like ffmpeg. By treating compiler passes as tools, LLMs can leverage decades of optimized algorithmic work. The article also points to the value of persistence in this model. Just as autotuners save optimal configurations to logs to speed up future builds, an LLM orchestrator can record successful pass pipelines and hints, amortizing the cost of reasoning over multiple compilations.

Impact on Low-Level Code Generation

Where LLMs are expected to have the most significant impact is in the authorship of low-level code. Tasks such as writing Triton or CUDA kernels, creating MLIR lowering patterns, and porting operators to new hardware are well-scoped and benefit from LLM assistance. However, Klaiber warns that this does not diminish the need for robust compiler infrastructure. Instead, it increases the demand for rigorous verification harnesses, clear intermediate representations (IR), and fast search mechanisms. The role of the compiler engineer shifts from writing every line of code to defining the interfaces and correctness criteria within which LLM-generated code operates.

Key Takeaways

  • LLMs acting as direct compilers are computationally inefficient compared to purpose-built algorithms.
  • Non-deterministic LLM outputs complicate the verification and debugging of AI compilation.
  • The optimal architecture involves LLMs orchestrating existing compiler tools and autotuners.
  • LLMs will significantly impact the generation of low-level kernels and backend code.
  • Compiler infrastructure must become more rigorous to support and verify LLM-generated artifacts.

The Bottom Line

The hype around LLMs replacing compilers ignores the massive efficiency gains of deterministic algorithms; the real revolution is in LLMs acting as smart drivers for existing, battle-tested toolchains.