The era of naive context management in AI agents is officially over. A new TypeScript implementation by developer Bobby Hall demonstrates that rule-based context compaction can reduce token usage by up to 96 percent while preserving critical failure signals. The project, released October 5, 2026, moves beyond simple truncation by integrating error-aware previews and aggressive elision of old tool results, directly addressing the 'quiet habit' of agent loops that re-read entire conversation histories on every turn.
The Cost of Forgetting Nothing
Standard agent architectures pay for every token, every time. As Hall notes, a 20,000-token test log is not a one-time cost; it is billed on every subsequent turn after it lands in the context window. Recent industry developments confirm this is the primary bottleneck for efficiency. On September 17, 2026, researchers published 'An Empirical Study of Harness Design for Coding Agents,' finding that context management benefits are driven primarily by preventing overflow failures. Similarly, the Strands Agents team reported 28 percent lower token costs in late September by truncating tool results over 1,500 tokens and triggering compaction above 85 percent of the window capacity.
Rule-Based Elision Beats Summarization
The new compactor rejects LLM-based summarization in favor of deterministic rules, aligning with findings from the DeepSeek Harness v0.2.1-alpha.1 release on October 3, 2026. The implementation uses four core policies: capping big tool results with head-and-tail previews, pinning failure lines (matching 'FAIL' or 'ERROR') to ensure they survive truncation, eliding tool results older than two turns into one-line stubs, and dropping entire messages when the context exceeds 85 percent of the window. This approach ensures that the full log remains immutable in the background, while the model only reads a curated 'view' of the conversation.
Benchmarking the Compactor
In a scripted agent run involving nine tool calls and a 32,000-token window, the demo revealed stark differences in efficiency. The 'naive' policy, which sends the full history, billed 290,208 tokens and failed to fit within the window at call 7. The 'cap' policy reduced billing to 22,905 tokens but lost the critical failure line. The 'cap+errors' policy cost only 152 more tokens than 'cap' but successfully preserved the failure signal. The full compactor, combining all rules, billed just 9,367 tokens with a peak of 1,637, demonstrating that cheap, rule-based machinery can outperform expensive summarization.
Key Takeaways
- Rule-based elision before LLM summarization offers the best efficiency-to-accuracy ratio, as confirmed by both academic studies and industry harnesses.
- Error-aware capping is essential; standard head-and-tail truncation often discards the specific line explaining a test failure, rendering the context useless for debugging.
- Immutable logs with dynamic views prevent 'compaction of compaction' errors, ensuring data integrity while allowing aggressive context trimming.
- Context windows are budgets, not containers; spending tokens on irrelevant historical data leads to 'wrong at a discount' outcomes where agents miss critical signals.
The Bottom Line
If your agent is still dumping raw tool outputs into the context window, you are burning money to pay for noise. The future of agent harnesses is not bigger windows, but smarter views.