Agent developers have long complained that wiring up multiple Model Context Protocol (MCP) servers leads to unpredictable behavior, not because of transport issues, but because of poor tool design. A new open-source project called mcp-lint, released by developer haoli, aims to solve this by statically analyzing tool definitions and assigning a design score out of 100. The tool identifies common pitfalls like vague naming conventions and missing descriptions that cause large language models to select incorrect functions.
Diagnosing the Design Flaws
The core problem, according to the project documentation, is that tools with generic names like handle_data or those lacking input schemas confuse agents. mcp-lint analyzes the tools/list output and flags specific issues, such as missing-description, vague-name, and empty-schema. Each violation deducts points from a baseline of 100, providing a granular view of which tools are likely to cause hallucinations or misrouted calls in an agent workflow.
Enforcing Quality in CI Pipelines
Beyond simple auditing, mcp-lint integrates into development workflows via a --fail-under flag. If the average design score of a server's tools falls below a specified threshold, such as 80, the command exits with a status code of 1, allowing teams to gate their builds. The tool also distinguishes between design quality and token cost, positioning itself as a sibling to mcp-tax, which audits context consumption. While mcp-tax ensures you aren't burning tokens unnecessarily, mcp-lint ensures the tools you do use are structured correctly for agent interpretation.
Key Takeaways
- mcp-lint scores MCP tools on eight design rules, including pagination hints and confirmation requirements for destructive actions.
- The tool is written in Python 3.9+ with no external dependencies, using only the standard library.
- It supports both human-readable table outputs and machine-readable JSON for automated integration.
- A score below the threshold can fail a CI build, enforcing design standards before deployment.
The Bottom Line
Tool description hygiene is the new unit test. If your agent is picking the wrong function, itβs not a model failureβitβs a linting failure. Run the audit, fix the schema, and stop wasting context on ambiguous names.