The standard wisdom for choosing AI tooling is simple: stop arguing, start measuring. Run the same tasks through each option, track time-to-merge, defect rates, rework cycles, and review effort, then let the data settle the debate. It's solid advice — and exactly what developer Javier Aguilar attempted in a recent experiment that went sideways.
Setting Up the Experiment
Aguilar structured their evaluation like any good infrastructure benchmark: identical representative tasks, consistent metrics tracked across both tools, and a commitment to following the evidence wherever it led. The goal was straightforward — turn an opinion into a decision backed by numbers. They even acknowledge this is advice they regularly give to others facing similar tooling decisions.
When Measurement Backfires
The catch? Their own data told a different story than their gut instinct. Rather than confirming which tool they'd preferred going in, the metrics pointed somewhere unexpected. The title of their post says it all: 'The Instrument Fails in Your Favour.' It's not that the measurement framework broke — it's that the results didn't align with what they wanted to find.
Why This Matters for Dev Teams
This is exactly why tooling debates drag on for months in engineering orgs. Even when teams agree to measure, there's often an unspoken assumption that measurement will validate existing biases. When it doesn't, the instinct is to question the instrument rather than reconsider the premise. Aguilar's experience shows how hard it is to stay truly neutral — even with a structured approach.
Key Takeaways
- Measurement frameworks are only as good as your willingness to accept uncomfortable results
- Time-to-merge and defect rates are useful metrics, but they don't capture everything about developer experience
- Starting an evaluation with a preferred outcome makes objective analysis difficult
- The real skill isn't running benchmarks — it's knowing when to trust the data over your gut
The Bottom Line
The lesson here isn't that measurement fails — it's that you need to commit fully or not at all. If you're going to run an experiment, accept that your tooling preference might be wrong. That's the whole point of bringing evidence into the conversation in the first place.