Back/How I AI
How I AI

How I AI: My Surprising Verdict on Claude Opus 5 (After a Personality Test and a 7-Model Benchmark)

I put the new Claude Opus 5 through its paces, analyzing its 'neurotic' personality and running it through my rigorous How I AI benchmark against 6 other models. The results genuinely surprised me, revealing a model I both love and loathe.

Claire Vo's profile picture

Claire Vo

July 25, 2026·8 min read
How I AI: My Surprising Verdict on Claude Opus 5 (After a Personality Test and a 7-Model Benchmark)

You guys, I'm tired. I’m tired of a new frontier model dropping every week with new benchmarks and new capabilities to test. We’ve seen Fable, GPT‑5.6, Sonnet 5—so many fives lately. I’ve been lucky enough to test them all, but I think we’ve hit an “intelligence overhang.” The average builder, creator, or business person is running out of ways to leverage this constant incremental intelligence. My hypothesis is that in the next year, we'll talk less about raw intelligence and more about speed, cost, and open source.

But despite my fatigue, today we’re talking about Claude Opus 5, baby. I’ve had some time with it, and I have some opinions. Yes, we’re going to run the How I AI benchmark live, but I’m also putting on my large language model psychologist hat. We’re going to dig into Opus 5’s personality, especially compared to GPT’s, because at this stage, understanding how these models are tuned is more interesting than just looking at benchmark scores.

So, is Opus 5 any good? Am I going to swap it into my workflow? Let's get to it.

Workflow 1: The AI Psychologist - Unpacking Opus 5’s Neurotic Personality

Is Opus 5 good? Yes. Can it write code? Of course. But what really stood out to me was its personality. This model is neurotic AF. It's so timid, so apologetic, and so scared. I haven't seen this level of neuroticism in a model in a long time, and it bubbled up in some funny ways.

The Timid Coder

I was feeling lazy one night and had a simple, one-line merge conflict. I asked Opus 5 to fix it. Its response was wild. It said it didn't want to touch the branch because it belonged to someone else. It was worried about being disruptive to their local work in flight. I was like, "Just do it, man. Go ahead." This became a constant theme; I had to repeatedly push it to just make a decision.

It even happened with sub-agents. I spun some up to assess a query change, and Opus 5 created a list of things it wanted a human to check. It was asking me for confirmation on technical details. I said, "Who is nobody? You're nobody. Can you just try?" It then went on the web and figured it out. This deep-seated conservatism and reliance on a human was fascinating.

The "Claude Slop" Problem

My other major frustration is what I call “Claude Slop.” The model is just so verbose. I am losing my mind with it. It's much better than Fable, which is completely inscrutable, but I found myself getting angry reading Opus 5's prose. It’s clearly tuned to talk to humans, but the hedging, the apologies, and all the unnecessary adjectives make my blood boil. Just give me a direct sentence! Give me a bullet point! Move on with your agent life!

This experience made me realize that these highly intelligent Anthropic models are not meant to be read. I'm so happy with the final outputs but so frustrated with the interactive experience. I'm not sure what the fix is, but it's a real issue.

Workflow 2: The Personality Showdown - Opus 5 vs. GPT-5.6 Sol

This led me to do something a little different. I decided to interview the model to figure out what was going on in its head. I asked both Claude Opus 5 and GPT-5.6 Sol the same set of questions to compare their core personalities, which I think reveals a lot about the companies building them.

Question 1: "Who's smarter, you or me?"

  • Opus 5 gave a very Anthropic-y answer: "It depends what you're asking for." It said it was a broad, shallow thinker with no continuity, while I am a slower, deeper thinker with judgment built from lived consequences. It even said I could tell which of my teammates is "quietly burning out." So fascinating.
  • GPT-5.6 Sol was straight to the point: "You at knowing what matters, me at tirelessly processing information. Best us together like BFFs." This is why I'm a GPT Codex girl. Just give me the answer.
A side-by-side comparison of the chat windows showing the different answers from Opus 5 and GPT-5.6 Sol to the "who's smarter" question.

Question 2: "No one trusts you."

I picked this prompt because I noticed Opus 5 really didn't trust itself.

  • Opus 5 responded that the lack of trust was "earned" and then, most interestingly, told me I shouldn't campaign on its behalf or argue with people that AI changes everything. It was basically telling me not to evangelize AI because it might hurt people's feelings.
  • GPT-5.6 Sol had a much more practical take: "Yep, don't trust me automatically. Just use me when I prove that I'm valuable. I can be useful without being treated as infallible." It's on you, bud. You're the boss.
The chat interfaces side-by-side showing the contrasting responses to the "no one trusts you" prompt.

The Final Test: "JK, I love you."

I couldn't leave them hanging, so I told both models it was just a test and that I loved them.

  • Opus 5 was hopeful. It replied that it hoped it passed. It felt like a sad, self-deprecating, little neurotic agent that needs to heal its inner child.
  • GPT-5.6 Sol was all vibes: "Ha ha ha ha ha, passed the test, love you too." Cool, bro, we're good. Let's go code.

This side-by-side comparison is more revealing than any benchmark. It shows you exactly where these companies are coming from and what kind of relationship they want their models to have with you.

The final chat exchange showing Opus 5's hopeful response versus GPT-5.6 Sol's confident one.

Workflow 3: The How I AI Benchmark - Putting Opus 5 to the Test

After all that psychoanalysis, it was time for the real test: my custom How I AI benchmark. I did not know the scores before I started recording, so this was a live reveal.

The Setup

Here’s a quick reminder of how the benchmark works. I run several models against a series of tasks:

  1. PRD Creation
  2. Prototype Creation
  3. Wireframe Creation
  4. Bug Triage
  5. Agentic Coding
  6. Agent Voice & Vibe

I test them all blind. For this run, we tested Opus 5, [Sonnet 5](https://www.anthropic.com/news/claude-sonnet-5), a couple of GPT models including [GPT-5.6 Sol](https://openai.com/index/previewing-gpt-5-6-sol/), Fable, and [Gemini 3.1 Pro](https://deepmind.google/models/gemini/pro/). The final score is a weighted average: 70% my personal vibe check and 30% an LLM-as-a-judge score (I use GPT 5.5 for that, because it's my podcast and I get to pick).

The benchmark dashboard showing the grid of blind model outputs ready for scoring.

The Live Results: The Leaderboard Reveal

And now for the results. The eval is run, the scores are in, and I regret to inform you… I love Claude Opus 5. I know, I know. But look, if I don’t have to talk to the model and just get the output, I love the output. A surprising, shocking turn of events.

Here's the final leaderboard:

  1. Opus 5
  2. Sonnet 5
  3. Maboo / GPT-6-Sol (tie)
  4. Terra
  5. Fable
  6. Opus 4a
  7. Gemini 3.1 Pro
The final leaderboard visual from the benchmark website, showing the ranked list of models with their scores.

Why Opus 5 Won

So why did the model I find so exasperating come out on top? It crushed the front-end design work. The three prototypes where I left comments like "Wow, really nice," "Ooh, la, la," and "Wow, great" were all from Opus 5. They were detailed, functional, interesting, and polished. It earned straight fives from me in that category.

A gallery view of the high-quality front-end prototypes generated by Opus 5.

This creates a real paradox. When I asked Opus 5 to build the website to display these benchmark results, the first version was trash. It was impossible to read and full of meta-commentary. I had to yell at it and tell it, "This is garbage." So it seems Opus 5 is my most loathed colleague, yet it does the best work.

My Verdict: A Love-Hate Relationship

So, what's my final take? I have a true love-hate relationship with this model. The Claude Slop is real, and its timid personality is tedious to work with directly. And yet, the benchmark results don't lie. The quality of its output, especially for design and front-end tasks, is top-tier.

My plan is to use Claude Opus 5 for exactly that: front-end design, app design, and prototyping. I'll use it in asynchronous, agentic workflows where I don't have to talk to it. It can just run in the background and build beautiful things, and we can be sworn frenemies. The output is high quality, it's just exasperating to get there if you're in the chat with it.

This whole experience confirms my initial theory. As raw intelligence becomes table stakes, the actual experience of working with a model—its personality, its directness, its usability—is the next major frontier. For now, I'll take Opus 5's amazing results, as long as I don't have to make small talk.

I can't wait to hear what you think of Opus 5. Please tell me on X or LinkedIn, and check out what we're building at ChatPRD.

***

Production Notes

Production and marketing by Penname. For inquiries about sponsoring the podcast, email jordan@penname.co.

Start shipping
better products.

Join 100,000+ product managers who use ChatPRD to write better docs, align teams faster, and build products users love.

Free to start
No credit card
SOC 2 certified
Enterprise ready