A new benchmark titled "We blind-tested ChatGPT, Claude, and Gemini on 20 everyday tasks (Open Dataset)" was posted to Hacker News on September 12, 2026. The post links to dailyskill.ai/ai-benchmark-2026/, which appears to be an independent evaluation comparing the three leading frontier models across a suite of practical, non-academic tasks. The dataset is described as open, suggesting the prompts and evaluation criteria are available for replication.
Methodology Remains Opaque
The source material is sparse. The Hacker News post contains only a title, URL, and a single comment, with a score of four points. The raw content of the linked article is largely unreadable in the provided source, appearing as compressed or corrupted data. This lack of accessible detail is a red flag for any benchmark claiming to provide rigorous analysis. Without clear visibility into the scoring rubric, the specific models tested (e.g., GPT-4o vs. GPT-4 Turbo, or Claude 3.5 Sonnet vs. Opus), and the task definitions, the results are difficult to trust.
The 'Everyday Tasks' Angle
Most academic benchmarks (MMLU, HumanEval, etc.) focus on specialized reasoning or coding. A shift toward "everyday tasks" is welcome in principle, as it better reflects how users actually interact with LLMs. However, the term "everyday tasks" is notoriously subjective. Does it mean writing emails, summarizing news, planning trips, or debugging a home router? The specific nature of these 20 tasks is critical to interpreting the results, and the current source material does not provide that clarity.
Key Takeaways
- A new open-dataset benchmark from dailyskill.ai claims to blind-test ChatGPT, Claude, and Gemini on 20 everyday tasks.
- The benchmark was posted to Hacker News on September 12, 2026, but received very low engagement (4 points, 1 comment).
- The source article content is largely unreadable in the provided data, making it impossible to verify specific scores, model versions, or task definitions.
- The focus on "everyday tasks" is a promising alternative to standard academic benchmarks, but requires transparent methodology to be credible.
The Bottom Line
Without clear access to the actual benchmark data, model versions, and scoring criteria, this "open dataset" is just noise. Until dailyskill.ai publishes a readable, detailed report, treat these rankings with extreme skepticism.