Developer Ryan Alberts has published Best-of-Agent-Harnesses, a curated GitHub repository ranking 167 AI agent harnesses, orchestration frameworks, and related techniques. The list rescoring weekly and includes machine-readable outputs for agents to query directly via an MCP server, llms.txt, and JSON endpoints. Categories span coding agents, personal runtimes, multi-agent orchestration, memory layers, sandboxing, and evaluation harnesses.

Harness Choice Outweighs Model Upgrades

The repository's central thesis is backed by SWE-bench Pro data cited from AINews (August 8, 2026) and analysis by @joelniklaus: swapping the agent harness changed pass@1 more than many model upgrades did. On GLM-5.2, different harnesses produced pass@1 scores ranging from 23% to 52%. On Gemma 4 26B, the spread was 15% to 36%. Critically, harness rankings barely transfer across models, with a rank correlation of -0.05, meaning a small model paired with the right harness can approach a larger model trapped in the wrong one. The gap widens on long-horizon tasks: a bare frontier model verified at ~30% on ARC-AGI-3, while Prime Agent's harness pushed Opus 5 to 95.5%.

Built for Agents, Not Just Humans

Unlike static awesome-lists, Best-of-Agent-Harnesses ships with an MCP server (io.github.RyanAlberts/agent-harnesses) that lets coding agents call recommend, pick_harness, compare, and pick_infrastructure to choose a harness matched to their model and task. The server includes graveyard warnings for abandoned projects and flags repos suspected of star manipulation. Three agent skeletons โ€” harness-scout, stack-auditor, and harness-radar โ€” ship in the repo, each running on existing AI subscriptions and delivering results to Slack or Notion. The stack-auditor can trace agent session logs to show how the harness steered technical decisions.

Landscape and Decision Framework

The 167 projects are organized into 12 categories including Progressive Disclosure Harnesses (8 projects), Coding Agent Products (24), Frameworks (26), and Multi-Agent Orchestration (12). Each entry is scored on four dimensions: GitHub stars (captured 2026-09-27), simplicity-to-capability tiers, headless-readiness (โ˜…), and durability (โœฑ). The repo includes head-to-head comparison pages for major matchups โ€” OpenClaw vs Hermes, opencode vs Codex vs Gemini CLI, Mem0 vs Letta vs claude-mem, and E2B vs Daytona vs Modal. Templates cover universal AGENTS.md files, safe Claude Code permission settings, and a minimal 180-line Python harness for learning.

Key Takeaways

  • Harness choice moves benchmark scores more than model upgrades; GLM-5.2 swung 29 points on SWE-bench Pro from harness alone
  • Rank correlation between harnesses across different models is -0.05, meaning you must re-evaluate harness pairing every time you change models
  • The MCP server and machine-readable formats let agents autonomously select harnesses, not just humans browsing GitHub
  • Prime Agent's harness took Opus 5 from ~30% to 95.5% on ARC-AGI-3, demonstrating the ceiling for well-designed runtime scaffolding
  • Weekly rescoring and graveyard warnings address the rapid churn in the agent tools space

The Bottom Line

The framework era is over and the harness era is here โ€” this repo is the first serious attempt to catalog that shift with empirical data. If your agent is failing, the problem is probably your runtime, not your model.