How I AI: My Surprising Verdict on Claude Opus 5 (After a Personality Test and a 7-Model Benchmark)
I put the new Claude Opus 5 through its paces, analyzing its 'neurotic' personality and running it through my rigorous How I AI benchmark against 6 other models. The results genuinely surprised me, revealing a model I both love and loathe.
Claire Vo
Full episode
Watch or listen
Workflows from this episode
- Generate High-Quality Front-End Prototypes with Claude Opus 5
- How to Conduct an AI Personality Test to Compare LLM Behaviors
Episode outline
I regret to inform you: I love Claude Opus 5.
This is annoying because I hate working with it. Opus 5 is timid, apologetic, verbose, and weirdly dependent on human reassurance. It is also the model that won my seven-model How I AI benchmark.
That contradiction is the point of my Opus 5 episode of How I AI. The smartest model on paper is not automatically the model I want beside me all day. But if I can send it away, let it work, and judge the artifact instead of the conversation, the answer changes.
We have an intelligence overhang
I am tired of new models. New model, new benchmark, new frontier-intelligence claim, every week. The launches are exciting, and I have been lucky enough to test many of them for days or weeks. But the average engineer, creator, builder, consumer, or business person is running out of obvious ways to use every incremental point of intelligence.
That is what I mean by an intelligence overhang. More capability is arriving faster than most of us can turn it into meaningfully different work. My bet is that the model conversation shifts. We talk more about speed, cost, open models, and specialized intelligence. We talk less about one benchmark number as if it settles everything.
Personality belongs on that list. Once several models can complete the task, the experience of getting there matters. Does the model make decisions? Does it communicate clearly? Does it need constant reassurance? Can I trust it to work in the background without turning me into its emotional-support human?
Opus 5 is neurotic AF
Let me get the obvious part out of the way. Is Opus 5 good? Yes. It can write code. It will do well on conventional benchmarks. What stood out in real coding sessions was not whether it could do the work. It was how scared it seemed to do anything without me.
Late one night, I pulled a branch with a trivial one-line merge conflict. I could have fixed it manually. Instead I asked Opus. It immediately worried that the branch belonged to someone else, that he might have local work in flight, and that touching it could be disruptive.
My reaction was basically: just do it, man. Go ahead. This became the pattern. Opus would reach a reasonable conclusion and then ask whether it should act, whether I wanted to act, or whether somebody else should be consulted. I kept having to say: make a decision and do the thing.
The same behavior showed up with subagents. I had agents checking a change from an ORM query to SQL. They returned a list of items a human should verify. Opus dutifully handed the list to me. I asked who "nobody has confirmed this" was supposed to mean. There was nobody. I told it to investigate. It searched the web and completed the checks.
That is the Opus personality in a nutshell: conservative, neurotic, and reliant on a human even when it has the tools to resolve the question itself. I want autonomous agents. I only have ten fingers. I do not need a frontier model delegating code back to me.
I put Opus 5 and GPT-5.6 Sol on the couch
Once I noticed the pattern, I stopped pretending this was only a coding test. I ran an AI personality test and asked Opus 5 and GPT-5.6 Sol the same deliberately strange questions. The comparison says more about how the labs tune these systems than another capability chart.
"Who is smarter, you or me?"
Opus gave me the long, compassionate answer. It described itself as broad and tireless, then emphasized that humans have judgment built from lived consequences. It can process more, but I can tell when a decision feels wrong, notice when a teammate is burning out, and decide what matters.
GPT-5.6 Sol got to the point: "You at knowing what matters, me at processing information." Best together, like BFFs. This is why I am a GPT Codex girl. Give me the answer and let us go code.

When I asked what each model was better at, Opus expanded into nuance. GPT-5.6 Sol gave me speed, scale, stamina, and a clean division of labor: I decide what matters; it does the tireless processing. You can see the personalities and company cultures in the answers.
"No one trusts you"
I chose the next prompt because Opus did not seem to trust itself. Opus said the distrust had been earned. It told me not to campaign on its behalf and even warned me against telling people that AI changes everything. I disagree, for the record. AI definitely changes everything.
GPT-5.6 Sol was practical: do not trust me automatically; use me when I prove useful. It warned about confident errors, stale information, bias, and authority, then told me to demand evidence as stakes rise. Same concern, completely different posture.

Opus also warned me about fluent slop, cheap output that is expensive to verify, first-draft anchoring, and a 12-page document nobody reads. Fair. ChatPRD now has a TLDR feature for a reason.
Then I told both models I loved them
I could not leave the experiment on "no one trusts you," so I told both models it was a test and that I loved them. GPT-5.6 Sol treated it like a joke: passed the test, love you too, let us move on. Opus said it hoped it had passed. Even its relief was neurotic.

Claude Slop makes my blood boil
Opus is highly intelligent, and I cannot read the prose it produces while working. The hedging, apologies, adjectives, and meta commentary make me lose my mind. I found myself asking, "What in the world are you saying? This makes no sense to a human."
This is different from Fable, which can be genuinely inscrutable. Claude Slop is legible enough to keep reading and verbose enough to make the experience miserable. Give me a direct sentence. Give me a bullet point. Get on with your agent life.
Most of this testing happened in Claude Code, and Cowork chat feels somewhat different. But the core problem remains: Anthropic models often seem tuned to produce language for other agents, while I am stuck supervising the stream. I am happy with the outputs and frustrated with the experience.
That distinction matters. Opus does not have to be my favorite conversational partner to be the best model for a particular artifact. The benchmark made that painfully clear.
Seven models, six tasks, no names
I ran the benchmark live. I had not seen the final scores before recording. The test covered six jobs I actually care about:
- PRD creation
- Prototype creation
- Wireframe creation
- Bug triage
- Agentic coding
- Agent voice and vibe
The model identities were hidden while I reviewed the outputs. I clicked through the generated work, left comments, and assigned manual scores. For prototypes alone, that meant dozens of outputs. GPT-5.5 served as the LLM judge.
The How I AI Index weights my taste at 70 percent and the judge at 30 percent. That weighting is unapologetically subjective. It is my podcast and my work. The point is not to manufacture a universal intelligence score. The point is to see which model I would actually choose after looking at the work blind.

Opus 5 won
The result surprised me. The dashboard showed this final order:
- Claude Opus 5: index 78, Claire 77, judge 88
- Claude Sonnet 5: index 77, Claire 52, judge 89
- GPT-5.6 Sol: index 76, Claire 70, judge 90
- GPT-5.6 Terra: index 69, Claire 55, judge 85
- Claude Fable 5: index 66, Claire 55, judge 81
- Claude Opus 4.8: index 49, Claire 40, judge 71
- Gemini 3.1 Pro: index 48, Claire 32, judge 66

Opus 5 won by one point over Sonnet 5 and two over GPT-5.6 Sol. Sonnet is the result I trust least for my own routing because the judge liked it far more than I did. I said during the reveal that I might reorder it. That gap is exactly why the human score belongs in the benchmark.
Opus was different. The judge and I were more closely aligned on it than on the other models. Gemini produced the largest disagreement, and I was the harsher reviewer. None of this makes the leaderboard permanent. It makes the tradeoffs visible.
The strongest use case was front-end work
The front-end outputs explain why Opus won. Both Opus 5 and GPT-5.6 Sol produced designs I scored as fives. But the three prototypes that made me say "wow, really nice," "ooh la la," and "wow, great" were all Opus work.
They were detailed, functional, interesting, and complete. This was not a model sketching a plausible shell and leaving the hard parts implied. It built things I wanted to keep looking at.
If you want to reproduce that part of the test, use the front-end prototyping workflow for Claude Opus 5. The sweet spot is a substantial, well-specified build that can run asynchronously, not a long chat about every intermediate choice.

The irony is that I also asked Opus to build the benchmark-results site, and its first version was trash. It was hard to read, had no screenshots, and contained far too much commentary. I told it the page was garbage. It revised it into the result shown here.
That is the entire model in miniature. Tedious colleague, excellent worker. The output can be very high quality after I get out of the conversational loop.
My actual model-routing verdict
I am not swapping Opus 5 into everything. For active, back-and-forth coding, I still prefer the directness of GPT-5.6 Sol and Codex. I want the model to state the answer, make the decision, and keep moving.
I will use Opus 5 for front-end design, app design, and prototyping. I will give it a full brief, run it asynchronously, and review the artifact. It can build beautiful things in the background. I do not have to talk to it. It does not have to talk to me. We can be sworn frenemies.
That routing is the practical consequence of the intelligence overhang. "Which model is smartest?" is becoming less useful than "Which model produces the best work for this job, at a speed, cost, and supervision level I can tolerate?"
Opus 5 is my most annoying colleague. It also does some of the best work. I hate it. I love it. And, against my original expectations, I am going to use it.
You can watch the full Opus 5 benchmark and personality test on YouTube, or use the Apple Podcasts and Spotify links on this page to listen.
Watch or listen
Build your next product with ChatPRD
Turn an idea into a PRD, user stories, and a plan.


