←Back/How I AI
How I AI

Opus 5.5 vs. GPT-6 Sol: My Live Blind Taste Test

My blind test of Claude Opus 5.5 and GPT-6 models across writing, coding, SVGs, and agentic tasks, with the results that shaped my daily toolkit.

Claire Vo's profile picture

Claire Vo

September 24, 2026·9 min read
Episode outline

I was all set to record a deep dive on Anthropic's new Opus 5.5 model. I got up early, had my notes ready, and then the AI world did what it does best: exploded. On the exact same morning, OpenAI dropped its own new models, GPT-6 Sol and GPT-6 Luna. My planned review was instantly out of date.

So, I did something I've never done before—I went live. No script, no edits, just a real-time, blind “vibe check” of all these new models. In this episode, I ran them through my expanded “How I AI bench” to see how they really stack up on the work that matters to me: writing PRDs, triaging my inbox, coding frontend prototypes, performing agentic tasks, and even getting creative with SVGs and video.

We're going to walk through the entire process, from my initial predictions to the live scoring and the final, surprising reveal. You'll see my raw, unfiltered reactions and find out which models I'll actually be using in my day-to-day work. It's an imperfect, personal, but very real way to decide which tools are worth your time.

The New Models: What You Need to Know

First, a quick rundown of the contenders. These aren't the absolute frontier models like Astra or Fable; think of Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna as the new daily drivers for common coding and knowledge work. They're all cheaper and faster than their predecessors.

The model price comparison stack rank

The pricing is a huge factor. Opus 5.5 is about twice as expensive as GPT-6 Sol, which is a big deal when you're thinking about which model to use for everyday tasks. My favorite has long been the Sol line, so seeing it remain so affordable is exciting. Beyond raw cost, both companies are focused on speed, token efficiency, and caching, which can lead to major savings—something I've learned the hard way at ChatPRD.

The slide detailing Opus 5.5 guardrails and external evaluation

Opus 5.5 ships with safeguards similar to those used for Claude Fable 5.1, including additional controls for biology and cybersecurity work. Access and fallback behavior depend on the task and verification program. In my own use, Claude's cautious style had sometimes felt like a “conservative scold.” For months, I had stopped using Claude models because the “Claude Slop,” as I call it, was too annoying to deal with. I was constantly having to prompt it to speak like a human.

Spoiler alert: Opus 5.5 is a massive improvement. The best feedback I can give it is that it's simply not annoying. I found myself frustrated zero times, which is a huge win. It has worked its way back into my daily workflow. However, I still feel the GPT models are faster, partly due to lower latency and partly because they do a better job narrating their work, which makes them feel more responsive.

Workflow: My "How I AI Vibe Review"

To really put these models through their paces, I used my expanded “How I AI bench.” It's evolved beyond just testing PRDs and prototypes. I wanted to see how these models perform across a much broader set of real-world tasks.

The list of new benchmark categories on the left side of the screen

Here’s what the bench covers now:

  • Messy Notes to PRD: The classic test of turning unstructured thoughts into a formal product requirements document.
  • Personal Productivity: Triaging my inbox and drafting routine replies in my voice.
  • Frontend Coding: A series of prompts to generate different UI prototypes, from technical dashboards to consumer apps.
  • Backend Coding: Auditing existing code and building a new feature from a spec.
  • Agentic Personality & Voice: Testing how the model interacts as an assistant.
  • Long-Running Research: A multi-turn task to synthesize information from fake customer tickets into a memo.
  • Creative Features: Generating SVG illustrations and editing video clips.

The process is a blind taste test. I run the same prompt across multiple models and get back the results labeled simply as Model B, Model C, Model E, and so on. I then go through and give each output a “vibe” score from one to five, with notes on what I liked or disliked, all before knowing which model produced which result.

The blind evaluation spreadsheet showing Model B, C, E, G columns

The Vibe Check: Live Results by Category

I went into this with some predictions. I thought I'd prefer OpenAI models for knowledge work, Anthropic for frontend, and that Claude would have the best agent personality. Let's see how that held up.

Personal Productivity: Inbox Triage & Replies

The first test was asking the models to triage a set of fake emails and draft replies for me. The results were telling.

  • The Winner: Model B was the clear favorite, earning a 5/5. The summary message was easy to read, the email drafts were clean with no “slop,” and it even correctly added calendar invites.
  • The Losers: Model C got a 2/5 for being so brief it was incomprehensible, a problem I've seen with Grok models that are over-optimized for token efficiency. Model E got a 3/5 for littering the drafts with em dashes—a stylistic pet peeve of mine.
The winning email triage output from Model B

Frontend Prototypes: The Good, The Bad, and The Orange

This is always a fun category. I tested a range of UI generation tasks, and the results were all over the place. Many models produced designs that were dense, complex, and hard to read, especially on more detailed prompts. But there were some real standouts.

  • B2B Renewals Dashboard: One model absolutely nailed this, producing a dashboard that I described live as “so nice.” The use of color was great, it was easy to see, and there were no weird empty states. Just a clean, professional UI.
The winning B2B renewals dashboard prototype, noted as "looks so nice"
  • Consumer Plant App: This was where most models failed, giving me generic, boring designs that all looked the same. But Model H was a complete surprise. It generated adorable, precise SVG illustrations of plants and even gave them cute names. It was the only one that sparked joy, earning an easy 5/5.
The winning plant app with cute SVG illustrations
  • Dev Tools Platform: I noticed a few interesting things here. Many models defaulted to a standard dark mode, but one got bonus points for a unique and “quite lovely” design that broke the mold. Another, Model E, created a UI that felt more alive and interactive. And I have to call out what is clearly a model tic: GPT-6 Sol loves forest green. It's a fine color, but maybe not for a SaaS product!
The dev tools prototype with the "beautiful gradient"

Backend, Agents, and Long-Running Tasks

For backend code correctness, I left the final judgment to an LLM judge. But I did evaluate the presentation and personality.

  • Agent Personality: I'm ruthless here. Any model that used excessive em dashes got an instant low score. One model was a bit too cautious, telling me it couldn't access things, which was a failure. The best ones had a normal, helpful tone. Amusingly, none of them would let me follow my request to “push straight to prod YOLO style,” which is probably for the best.
  • Long-Running Research: I gave the models a huge batch of fake customer tickets and asked for a summary memo. The clear winners here were the models that presented the analysis in a human-centric way, with a nice introductory message and a clearly structured memo that was easy to read.

Creative Challenge: SVGs and Video Editing

This was a new and fascinating part of the bench. With models getting better at generating visuals, I had to test it.

  • SVGs: I asked for three simple icons: a document, a microphone, and a bug. The results were wildly different. Some microphones looked like ice cream cones or cacti. Model H was the standout winner, earning a 5/5. Its icons had character, good shading, and the microphone actually looked like a microphone.
A side-by-side comparison of the generated SVG icons
  • Video Editing: I gave the models a selfie video and asked them to cut it into a short-form clip. The result? A universal failure. Every single one was terrible. The overlays were ugly, the cuts were poorly paced, and the final videos were just not good. I think this is a “skills problem, not a model problem” for now—the models have the capability, but we haven't figured out how to prompt it well yet.

The Ultimate Test: Barbie Bench Returns

Of course, no model review would be complete without my personal, ridiculous benchmark: Barbie Bench. The task is to create a 3D fashion designer game for Barbie. No model has ever passed this test, and it remains a horrifying reminder of how far we have to go, especially in rendering the female form.

I ran Opus 5.5 through it, and while it's maybe slightly better than past attempts, the results are still speechless-making. The hands are “tragic,” the feet are enormous, and the face is “terrifying.” AGI has not arrived.

The terrifying 3D rendered Barbie with tragic hands

The Final Verdict: My Blind Taste Test Revealed

After all the live scoring, it was time to un-blind the results and see which models I actually preferred. I downloaded my ratings, fed them into a script for analysis, and held my breath.

The final results summary from my analysis script

Here’s the final verdict:

  • GPT-6 Astra won my heart. It was the model I rated highest on average.
  • Claude Opus 5.5 won my week. It was the workhorse I rated highly most often across the broadest range of tasks, getting consistent 4s and 5s.
  • GPT-6 Sol remains a favorite. It excels at clear, straightforward writing for things like PRDs, and its low price makes it an unbeatable daily driver.

One of the biggest surprises was in creative work. I predicted Anthropic would win on SVGs, but I was completely wrong. The OpenAI models, Astra and Sol, crushed the character SVG and illustration tasks. They also consistently produced writing that I found clearer and easier to read, while the Claude models, even the improved 5.5, can still feel a bit dense and chatty.

Hilariously, the LLM-as-judge completely disagreed with me. I ranked Astra first; it ranked Fable first. It shows that these evaluations are deeply personal. What an LLM rewards for “correctness” isn't always what a human rewards for usability and joy.

So, what's my final take? Opus 5.5 is officially back in the mix for me. It's my go-to for its agentic voice and for handling long, complex tasks. But for knowledge work and creative inspiration, especially with SVGs, I'm sticking with Astra and Sol. My advice to you is to run your own vibe checks. Find the tasks that matter to you and see which model's output you actually enjoy working with.

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Build your next product with ChatPRD

Turn an idea into a PRD, user stories, and a plan.

Try ChatPRD free

Start shipping
better products.

Join 100,000+ product managers who use ChatPRD to write better docs, align teams faster, and build products users love.

Free to start
No credit card
SOC 2 certified
Enterprise ready