Thursday, 24 September 2026

The Deep Feed

The Friction of Progress: Models, Myths, and the Search for Substance

77 min read · 6 pieces
In this issue
01 The Blind Taste Test: Why AGI is Still a Mirage 12 min
02 The Return to Claude: Why Efficiency Trumps Personality 10 min
03 The Toxicity of Folklore 15 min
04 The New Moat: Why Evals are the Real Product Strategy 11 min
05 The Myth of the Old West 5 min
06 The Academy vs. The University 14 min
Editor's Letter

Tonight we look at the friction points of the modern age. From the messy reality of testing frontier AI models to the dangerous allure of folkloric toxins, we examine where the hype meets the hard truth.

01 — Lenny's Newsletter

The Blind Taste Test: Why AGI is Still a Mirage

Testing the latest frontier models reveals a gap between intelligence and utility.

By Claire Vo · 12 min read
Editor's note: As models race toward intelligence, the actual utility for professional work remains inconsistent and frustrating.

The race for Artificial General Intelligence (AGI) is often described as a steady climb, but for those actually using these tools to build products, it feels more like a series of erratic jumps and sudden falls. When Anthropic and OpenAI released their latest models—Claude Opus 5.5 and GPT-6 Sol—on the same morning, the industry expected a clear winner. Instead, a blind testing process revealed something far more complicated. The results didn't point to a single dominant intelligence, but rather a collection of specialised tools that excel in some tasks while failing spectacularly in others. This isn't just a matter of preference; it is a matter of reliability in professional workflows.

The Performance Gap

In a rigorous blind evaluation, the models were pushed through real-world stressors: writing PRDs, generating frontend prototypes, and managing long-running agentic tasks. The findings were split. GPT-6 Astra captured the emotional resonance of the user, providing a sense of 'personality' that felt more natural. Meanwhile, Claude Opus 5.5 emerged as the workhorse, particularly for B2B frontend development and complex, multi-step agent tasks. Yet, despite these strengths, the 'Barbie Bench'—a test involving 3D fashion games—exposed the cracks. The models struggled with basic spatial and aesthetic consistency, producing 'tragic' hands and broken geometry. It is a stark reminder that while a model might write a perfect email, it still cannot grasp the fundamental physics of a digital object.

AGI has not arrived; we are still playing with very clever, very broken toys.

The discrepancy between human judgment and LLM-based judges is another layer of the problem. When an AI was used to score the outputs, it often rewarded different qualities than a human expert. The AI judge tended to favour structure and adherence to explicit instructions, whereas the human user valued nuance, tone, and the 'vibe' of a prototype. This suggests that as we build more automated evaluation systems, we risk training our models to satisfy a mathematical metric rather than a human need. We are building for the judge, not the user.

Model Strengths
  • GPT-6 Astra: High personality and intuitive interaction
  • Claude Opus 5.5: Superior for B2B frontend and long-running agents
  • GPT-6 Sol: Strongest for clear writing and readable documentation
  • The Failure Point: 3D rendering and spatial reasoning

The takeaway for agency owners and product builders is clear: do not bet your entire stack on a single model. The current state of the art is a fragmented landscape of specialists. You need a model for your code, a model for your copy, and perhaps a different model entirely to act as the supervisor. The era of the 'one model to rule them all' is a marketing myth that ignores the technical reality of how these systems actually function.

Key Takeaway

Intelligence is currently fragmented; success requires a multi-model strategy rather than a single-vendor dependency.

02 — Lenny's Newsletter

The Return to Claude: Why Efficiency Trumps Personality

A deep dive into the shifting economics and usability of frontier models.

By Claire Vo · 10 min read
Editor's note: The decision to switch models is rarely about intelligence alone; it is about the friction of the user experience.

For months, the allure of OpenAI's Codex kept many power users away from Anthropic's ecosystem. The reason wasn't a lack of capability, but a lack of temperament. Earlier versions of Claude were plagued by a specific kind of digital neurosis: excessive hedging, unnecessary disclaimers, and a tendency to ramble when a direct answer was required. In a high-velocity work environment, these 'safety' features feel less like guardrails and more like speed bumps. When you are trying to build a product, you don't need a lecture on ethics; you need a functional piece of code.

The Economics of Agency

The release of Opus 5.5 changed the math. Anthropic didn't just improve the model; they changed the pricing structure, making it 40% cheaper than its predecessor. This is a critical distinction for anyone building agentic workflows. If an AI agent needs to perform fifty micro-tasks to complete a single user request, the cost of each prompt becomes a primary constraint on the product's viability. A model that is slightly less 'smart' but significantly cheaper and faster is often more valuable than a brilliant but expensive oracle. Opus 5.5 represents a shift from 'chatting' to 'doing'.

In the age of agents, the cost per token is the new unit economics of intelligence.

However, the friction hasn't entirely vanished. Even with the new alignment approach, there are moments where the model's refusal to perform a task feels arbitrary. There is a tension between making a model 'safe' and making it 'useful'. If a model is too cautious, it becomes a liability in a professional setting. The goal is to find the threshold where the model can say 'no' to a genuine threat without saying 'no' to a complex, edge-case request that is perfectly legitimate.

Why Opus 5.5 Wins
  • Lower cost for high-volume agentic tasks
  • Improved speed for real-time prototyping
  • Better handling of long-running, multi-step processes
  • Reduced 'preachy' tone compared to previous versions

The real winner in this model war is the user, but only if they remain unattached to any single provider. The ability to split a stack—using Codex for logic, Opus for frontend, and Sol for writing—is the only way to navigate the current volatility. We are moving away from a world of general-purpose assistants and into a world of specialized, modular intelligence.

Key Takeaway

Model selection is a trade-off between intelligence, personality, and the brutal economics of token costs.

03 — Aeon

The Toxicity of Folklore

The dangerous reality behind the 'magic' of Amanita muscaria.

By Eric Leas · 15 min read
Editor's note: When folklore meets the unregulated supplement market, the results can be lethal.

In the dimly lit aisles of a modern smoke shop, a package of 'magic mushroom' gummies sits next to standard cannabis products. The packaging evokes the whimsical imagery of Super Mario or the fairy-tale forests of Lewis Carroll, but the chemistry inside is far from playful. These are not psilocybin gummies, the kind currently being studied for therapeutic potential in clinical settings. These are Amanita muscaria—the red-capped toadstool that has occupied the human imagination for millennia. But as a public health scientist, the fascination with its folklore is secondary to the immediate, physical danger it poses.

A Different Kind of High

The neurological distinction between psilocybin and muscimol (the primary compound in Amanita) is fundamental. Psilocybin works by activating serotonin receptors, leading to the structured, often expansive alterations in perception associated with the psychedelic experience. Muscimol, however, targets the GABA-A receptors—the brain's primary inhibitory system. Instead of an 'expansion' of consciousness, muscimol often produces a 'contraction.' The experience is characterized by sedation, confusion, disorientation, and a loss of motor control. It is less a journey and more a descent into a delirious, unpredictable state.

Muscimol does not expand the mind; it shuts down the brain's ability to regulate itself.

The danger is compounded by the lack of regulation in the supplement market. Manufacturers often use the term 'magic mushrooms' to capitalize on the cultural cachet of psilocybin, while selling a substance that is pharmacologically distinct and significantly more toxic. In toxicological studies, the dose required to kill half a test population (LD50) for muscimol is alarmingly low—in some animal models, lower than that of fentanyl. This is not a substance for casual exploration; it is a potent neurotoxin being sold under the guise of a wellness trend.

The Risks of Amanita
  • Inhibitory neurological effects (GABA-A activation)
  • High risk of profound sedation and delirium
  • Significant toxicity compared to other psychotropics
  • Misleading marketing in the unregulated supplement market

The intersection of ancient myth and modern commerce creates a perfect storm for public health crises. When a substance is wrapped in the aesthetics of a childhood fantasy, the consumer's guard is lowered. We must distinguish between the cultural history of a plant and its physiological reality. In the case of Amanita muscaria, the myth is beautiful, but the chemistry is dangerous.

Key Takeaway

Cultural mythology is a poor substitute for pharmacological reality, especially in unregulated markets.

04 — Lenny's Newsletter

The New Moat: Why Evals are the Real Product Strategy

Moving beyond metrics to find the true failures in AI products.

By Hamel Husain, Shreya Shankar · 11 min read
Editor's note: In the AI era, the ability to measure quality is more important than the ability to build features.

For most product managers, the instinct is to build features. In the world of AI, that instinct is a trap. Because AI models are inherently probabilistic, a change that improves one feature can silently break three others. This creates a level of volatility that traditional software testing cannot handle. The solution is 'evals'—evaluation systems that turn human judgment into repeatable, automated tests. But most teams are doing it wrong. They jump straight to creating metrics, trying to measure 'accuracy' or 'helpfulness' before they even understand how their product actually fails.

The Error Discovery Phase

The most critical, yet most skipped, step in the AI lifecycle is 'error discovery'. This is the equivalent of product discovery, but instead of looking for user needs, you are looking for failure modes. It requires digging through messy, unglamorous user session logs to find the exact moments where the model hallucinated, refused a valid request, or provided a sub-optimal answer. Without this phase, you are merely building dashboards around generic metrics that don't reflect the actual user experience. You end up measuring the wrong things very efficiently.

Evals are not just a testing tool; they are the new PRDs for the AI era.

When done correctly, error discovery creates a compounding advantage. Every production error becomes a new test case. As your library of failures grows, your ability to ship changes without breaking the system increases. This is how companies like Shopify and Cursor have managed to reduce costs and increase user satisfaction simultaneously. They aren't just prompting better; they are measuring better.

The Eval Workflow
  • Error Discovery: Find the actual ways users are being let down
  • Metric Creation: Turn those failures into quantifiable scores
  • Continuous Loop: Automate the testing of every new prompt or model change

If you are an agency owner or a product lead, your competitive moat is not your prompt library. It is your evaluation library. The team that knows exactly how their AI fails is the only team that can safely scale it.

Key Takeaway

Don't measure what is easy; measure what is broken.

05 — Aeon

The Myth of the Old West

How theme parks curate a sanitized version of history.

By Aeon · 5 min read
Editor's note: Nostalgia is a powerful tool for escapism, but it often requires the erasure of truth.

At Old Tucson, the Arizona sun beats down on a town that never truly existed. Visitors dress in the costumes of sheriffs and outlaws, participating in a choreographed dance of cowboy romance. It is a funhouse mirror of the American frontier, designed to highlight the adventure while carefully obscuring the poverty, violence, and systemic drudgery of the actual era. This is not history; it is a curated experience of nostalgia.

The Allure of the Simplified Past

Why do people pay to step into this fiction? For many, it is an escape from a modern world that feels increasingly complex and politically charged. The Old West offers a version of history that is legible and simple: good guys versus bad guys, clear boundaries, and a sense of rugged individualism. By stripping away the messy realities of the 19th century, theme parks provide a psychological relief—a chance to play in a world where the stakes are clear and the outcomes are scripted.

We don't go to theme parks to see the past; we go to see the version of the past we wish we had.

This sanitization is a form of cultural editing. By focusing on the 'movie mythos' of the West, we reinforce certain narratives while silencing others. The theme park becomes a space where history is consumed as entertainment rather than understood as a complex series of events. It is a comfortable lie that serves the needs of the present.

Key Takeaway

Nostalgia is often a desire for simplicity, not for truth.

06 — Not Boring

The Academy vs. The University

Can intensive, peer-driven learning replace the traditional college model?

By Packy McCormick, Alex Danco · 14 min read
Editor's note: As the cost of traditional education skyrockets, new models of intensive, practical learning are emerging.

The launch of the Horowitz Andreessen Academy (HAA) raises a question that many in the tech sector have been whispering about for years: is the traditional university model finally unbundling? HAA is not a college in the classical sense. It is a residential program in San Francisco designed to gather ambitious, young people and put them in a high-pressure environment where the primary goal is to build things and be useful. It is an apprenticeship disguised as an academy.

The Hard Parts of Growing Up

When discussing whether one would send their own children to such a program, the conversation shifts from theoretical disruption to practical parenting. The value of a university is often not found in the curriculum, but in the 'soft' infrastructure of the experience. It is about resilience, social intelligence, and the ability to navigate complex peer groups. Traditional colleges provide a buffer—a controlled environment to develop the character traits that are required for adulthood. The question for HAA is whether a high-intensity building program can replicate that social development.

The real purpose of education isn't just the skill; it's the life experience that builds resilience.

Critics might argue that HAA is too utilitarian, focused solely on the trajectory of becoming a 'builder'. However, the program seems to acknowledge that the most effective way to learn is through solving real problems alongside a motivated peer group. This mirrors the Paul Graham philosophy: the best way to find a great idea is not to think about ideas, but to solve a problem you personally have. HAA is essentially applying this to human capital.

The HAA Model
  • Residential, peer-driven environment
  • Focus on practical building and utility
  • Located in a high-density tech hub (San Francisco)
  • Alternative to the traditional, theoretical degree

Whether this model succeeds depends on whether it can provide more than just technical training. If it can foster the same sense of community and character development as a university, it could represent a fundamental shift in how we prepare the next generation for a rapidly changing economy.

Key Takeaway

The future of education may lie in intense, practical immersion rather than broad, theoretical study.

Endnote
Tonight's pieces, though seemingly disparate, share a common thread: the tension between the ideal and the actual. We see it in the gap between AI's promised intelligence and its current, erratic performance. We see it in the difference between the romanticized 'magic' of a mushroom and its lethal chemical reality. We see it in the curated history of the Old West and the debate over whether a new kind of academy can truly replace the social fabric of a university. In every case, there is a temptation to accept the polished, simplified version of reality. But as we have seen, the truth—whether in a lab, a code editor, or a history book—is almost always more complex, more dangerous, and more interesting than the myth.
Where in your own life are you accepting a polished myth in place of a messy reality?
The Deep Feed · A nightly magazine · Thursday, 24 September 2026