The Blind Taste Test: Why AGI is Still a Mirage
Testing the latest frontier models reveals a gap between intelligence and utility.
The race for Artificial General Intelligence (AGI) is often described as a steady climb, but for those actually using these tools to build products, it feels more like a series of erratic jumps and sudden falls. When Anthropic and OpenAI released their latest models—Claude Opus 5.5 and GPT-6 Sol—on the same morning, the industry expected a clear winner. Instead, a blind testing process revealed something far more complicated. The results didn't point to a single dominant intelligence, but rather a collection of specialised tools that excel in some tasks while failing spectacularly in others. This isn't just a matter of preference; it is a matter of reliability in professional workflows.
The Performance Gap
In a rigorous blind evaluation, the models were pushed through real-world stressors: writing PRDs, generating frontend prototypes, and managing long-running agentic tasks. The findings were split. GPT-6 Astra captured the emotional resonance of the user, providing a sense of 'personality' that felt more natural. Meanwhile, Claude Opus 5.5 emerged as the workhorse, particularly for B2B frontend development and complex, multi-step agent tasks. Yet, despite these strengths, the 'Barbie Bench'—a test involving 3D fashion games—exposed the cracks. The models struggled with basic spatial and aesthetic consistency, producing 'tragic' hands and broken geometry. It is a stark reminder that while a model might write a perfect email, it still cannot grasp the fundamental physics of a digital object.
AGI has not arrived; we are still playing with very clever, very broken toys.
The discrepancy between human judgment and LLM-based judges is another layer of the problem. When an AI was used to score the outputs, it often rewarded different qualities than a human expert. The AI judge tended to favour structure and adherence to explicit instructions, whereas the human user valued nuance, tone, and the 'vibe' of a prototype. This suggests that as we build more automated evaluation systems, we risk training our models to satisfy a mathematical metric rather than a human need. We are building for the judge, not the user.
- GPT-6 Astra: High personality and intuitive interaction
- Claude Opus 5.5: Superior for B2B frontend and long-running agents
- GPT-6 Sol: Strongest for clear writing and readable documentation
- The Failure Point: 3D rendering and spatial reasoning
The takeaway for agency owners and product builders is clear: do not bet your entire stack on a single model. The current state of the art is a fragmented landscape of specialists. You need a model for your code, a model for your copy, and perhaps a different model entirely to act as the supervisor. The era of the 'one model to rule them all' is a marketing myth that ignores the technical reality of how these systems actually function.
Intelligence is currently fragmented; success requires a multi-model strategy rather than a single-vendor dependency.