The gap between raw database documentation and what an AI agent actually needs to write correct SQL is wider than most teams realize. A new compiler, part of the OKF (Open Knowledge Format) series, aims to bridge that divide by transforming schema dumps and column dictionaries into structured, linked bundles that agents can navigate like a wiki. Written by Lakhan Malviya and published on DEV.to, this third installment in the series introduces okf_compiler, a Python package that reads raw documentation from the LiveSQLBench benchmark and outputs concept files for each table and business rule.

The Problem With Raw Context

Most SQL agents are fed the entire schema and all business rules in one giant prompt, which is inefficient and often leads to hallucinations when the model gets lost in the noise. The author argues that agents need a way to find specific knowledge without loading everything into context. The solution is an OKF bundle: a directory of markdown files, each representing a single concept (like a table or a rule), linked together by an index. This allows an agent to read the index first, then fetch only the specific files relevant to a user's question, drastically reducing token usage and improving accuracy.

Building the Compiler

The okf_compiler package is structured to handle the messy reality of real-world documentation. It uses a readers.py module to parse three different file formatsβ€”a SQL text file for the schema, a JSON object for column meanings, and a JSON Lines file for business rulesβ€”into unified Python dataclasses. A key design choice highlighted in the post is that table descriptions are generated programmatically from column names and join relationships, rather than by an LLM. This ensures the bundle remains faithful to the source documentation, avoiding the introduction of 'hallucinated' context that could skew benchmark results. The compiler also flags columns that lack descriptions in the source, such as patients.clinleadref in the 'mental' database, ensuring gaps are visible rather than hidden.

Executable Knowledge

The compiler targets the disaster database from LiveSQLBench, which contains 10 tables and 54 business rules about disaster-response operations. By running python -m okf_compiler.cli disaster, the tool generates a bundle in bundles/compiled/disaster. Each table concept file includes YAML frontmatter with metadata like generated.by and sources, followed by a markdown body listing columns, types, meanings, and join paths. This structure allows the subsequent agent component to treat the database documentation as a navigable knowledge graph, rather than a static text blob. The next part of the series will build the actual agent that consumes these bundles and compares its performance against a baseline where all documentation is pasted directly into the prompt.

Key Takeaways

  • The okf_compiler transforms raw schema and rule files into linked markdown bundles for AI agents.
  • Descriptions are generated deterministically from source data to prevent LLM-induced context pollution.
  • The system explicitly reports missing documentation gaps instead of silently ignoring them.
  • This approach enables agents to fetch only relevant context, optimizing token usage and accuracy.

The Bottom Line

If your agent is drowning in tokens, stop pasting the whole schema into the prompt. Build a compiler that turns docs into a navigable knowledge base. This is the future of agentic RAG.