Back/How I AI
How I AI

How I AI: My GPT-5.6 Sol Benchmark & 4 Game-Changing Workflows (vs. Fable)

GPT-5.6 Sol is back, and I ran it through my full How I AI vibe benchmark against Fable and others. I'm sharing the results, plus four workflows for building apps, video editing, and browser automation that show exactly where Sol wins.

Claire Vo's profile picture

Claire Vo

July 9, 2026·9 min read
Episode outline

I kept coming back to GPT-5.6 Sol for the same reason: it reliably got me to a usable result faster. In this episode of How I AI, I compare Sol, Terra, and Luna against Anthropic’s Fable 5 and Sonnet 5, then walk through the real workflows where Sol earned a permanent spot in my stack: functional prototyping, debugging stuck coding systems, video editing, and browser automation.

The comparison uses my "How I AI vibe review benchmark," a deliberately opinionated test built around the work I actually do every day. Instead of abstract benchmarks alone, I grade outputs myself: reading PRDs, clicking through prototypes, reviewing debugging results, and judging whether the models are tolerable collaborators over long sessions.

Sol is the flagship model in the GPT-5.6 family, Terra is aimed at balanced everyday work, and Luna is the lower-cost high-volume option. My conclusion is not that one model wins universally. Different models were better at different jobs, and the benchmark reflects my own preferences for direct writing, functional interfaces, and systems that adapt when an implementation gets stuck.

GPT-5.6 Sol, Terra, and Luna model images from the OpenAI blog post

The How I AI vibe review benchmark

My benchmark focuses on four categories that map closely to my day-to-day product and coding workflow:

  • PRD generation: whether the model produces concise, useful product requirements without drifting into generic AI business writing.
  • Wireframes and prototypes: whether it can create distinctive interfaces with enough functionality to meaningfully test an idea.
  • Coding and debugging: whether it can complete a multi-step technical investigation instead of stopping at surface-level fixes.
  • Agent voice: whether the model sounds readable and natural enough to collaborate with for hours at a time.

The harness runs the same tasks across GPT-5.6 Sol, Terra, and Luna, plus Fable 5 and Sonnet 5. GPT-5.5 acts as the model judge for the automated scoring. I then review the outputs myself, read the generated documents, click through the interfaces, test the interactions, and leave written notes alongside my scores.

How the final score is weighted

I weighted the final index 70 percent toward my own review and 30 percent toward the model judge. Across dozens of outputs, Sol earned the strongest "taste score" by a wide margin. That result reflects the kinds of things I consistently rewarded: clear writing, prototypes that actually worked when clicked through, and models willing to revise an approach instead of rigidly defending a failing implementation.

The How I AI benchmark dashboard showing model scores

Fable 5 still produced strong work in cases where I did not need to collaborate conversationally with it. Terra and Luna held up well for practical use, while Sonnet 5 finished lower overall in this benchmark. Because the evaluation heavily emphasized front-end prototyping and app-building workflows, the final rankings naturally favored those capabilities.

The Claire Weighted Index results showing Sol at the top

Different models won different categories

The category-level breakdown mattered more to me than the single leaderboard score because each model developed recognizable strengths and weaknesses:

  • Prototypes: I consistently preferred GPT-5.6 Sol because the interfaces felt more functional, visually opinionated, and less like standard AI-generated dashboards.
  • PRD writing: GPT-5.6 Terra became my favorite for requirements documents because the writing was cleaner, shorter, and more straightforward.
  • Bug hunting: GPT-5.5 as judge favored Sonnet 5 for completeness and accuracy, though I said this benchmark still needed refinement before I fully trusted the result.
  • Agent voice: Sonnet 5 sounded the most human to me. It still had AI quirks, but it felt noticeably easier to collaborate with over long conversations.

A dense operations dashboard prototype made the differences especially obvious. Sol’s version used clearer hierarchy, semantic color choices, and controls that actually worked when tested. I repeatedly rewarded prototypes that carried the functionality all the way through instead of stopping at a pretty mockup.

Side-by-side comparison of the doc scheduler dashboard (Sol vs. Fable)

Fable’s dashboard looked competent but much more familiar: dark mode, monospace styling, and the standard developer-tool aesthetic I see repeatedly in AI-generated interfaces. My scoring consistently favored designs with a stronger point of view and enough implementation depth to evaluate the product itself, not just the visual shell.

Building a surprisingly complete homework app from a single PRD

I took an XP-system proposal my husband generated through OpenClaw, pasted the PRD into Codex, and asked GPT-5.6 Sol to build a fully gamified homework tracker for my children. The goal was not a toy wireframe. I wanted a functioning system with incentives, dashboards, rewards, and parent controls. See How to Build a Gamified Homework App with AI in a Single Shot.

The gamified homework app dashboard showing quests for both kids

The result still had obvious AI-generated design tells, but I was impressed by how much product surface area appeared in a single pass. Instead of producing a shallow mockup, Sol generated a multi-user app with distinct experiences for kids and parents, integrated progression systems, and enough functionality to immediately test the concept with my family.

  • Student dashboards assigned each child a different summer quest. The app included personalized tasks, focus timers, progress tracking, and rewards tied to the children’s actual interests, including basketball coaching, Minecraft-related rewards, and family activities.
  • The reward system layered in XP, companion characters, power auras, collaborative bonuses, and unlockable progression systems. Sol even built cooperative mechanics where siblings could earn additional rewards for completing activities together.
  • The Parent HQ included controls for reviewing progress, enabling or disabling quests, changing XP values, adding tasks, editing rewards, and viewing history. My main takeaway was not that the app was ready to launch, but that the model carried details from the PRD consistently across both the child-facing and administrative experiences.
The "Parent HQ" section of the homework app

I was explicit that I would not ship the result as a real consumer product without substantial refinement. What mattered was the speed and completeness of the first draft. In one attempt, I had enough functionality to evaluate the incentive structure, test interactions, and decide whether the idea was worth investing in further.

Using Sol to escape a debugging dead end

One of my clearest distinctions between the models was that Fable often felt theoretically brilliant while Sol felt practically effective. I repeatedly preferred models that could loosen constraints and rethink assumptions when a system stopped working instead of doubling down on technical purity.

While building an integrated prototyping tool inside ChatPRD, I hit a wall with the tool-calling architecture. Only GPT-5.5 would reliably complete the loop. Sonnet, Opus, and several other models repeatedly failed despite extensive evaluation runs.

In my testing, Fable kept framing the issue as a limitation of the other models instead of questioning whether the harness itself had become too rigid and over-engineered.

"No, bro, that's, it's, it's totally these models, model's fault."

I switched to Codex with GPT-5.6 Sol and asked it to reevaluate the setup from first principles rather than preserve the existing architecture.

"Look, I'm just not convinced we can't get Sonnet 5 to work. This is ridiculous. Just do what you think is correct."

Sol altered the implementation enough to get Sonnet 5 running almost immediately. I did not describe the result as elegant or fully solved, but it broke a debugging stalemate that multiple models had reinforced. That willingness to revise constraints instead of protecting them became one of my strongest arguments in Sol’s favor.

Turning a conference talk into publishable social clips

After speaking at a Cursor event about the future of product management, I received the full session recording and wanted short promotional clips for social media. Normally that process meant manually searching through a long video, finding strong moments, trimming them, and tightening pacing by hand. You can follow the full implementation in How to Quickly Create Social Media Video Clips Using AI. See How to Quickly Create Social Media Video Clips Using AI.

Instead, I dropped the recording directly into Codex and asked GPT-5.6 Sol to generate five clips optimized for social platforms.

Can you cut this video into five clips for social?

I then iterated with more specific editing direction, asking for horizontal formats, clips pulled from different sections of the talk, and faster pacing with tighter cuts. I moved the outputs into CapCut, added music, and published them. The workflow did not replace final editing judgment, but it eliminated most of the tedious searching, clipping, and rough-cut assembly work.

The Codex interface showing the video file being dropped in for editing

Using browser automation inside a logged-in Chrome session

I also tested GPT-5.6 Sol with Codex browser control connected to an authenticated Chrome session. The setup allowed the model to operate directly inside visible websites and complete multi-step tasks while preserving my existing login state. See How to Automate LinkedIn Messaging with AI Browser Control.

The Codex interface showing the @Chrome prompt for LinkedIn automation

One experiment focused on LinkedIn inbox management. I instructed the system to review incoming messages, maintain a very high threshold for valuable ChatPRD or How I AI conversations, and only accept connection requests from executives at selected companies.

"Can you use Chrome to reply to messages that are a very high value to ChatPRD or the How I AI podcast? Keep the bar very high... only accept them if they're executives of tier one companies."

The automation processed roughly 500 messages, replied to selected contacts, thanked listeners who had sent positive podcast feedback, and filtered connection requests according to my rules. I also used browser control for testing web apps and filling out repetitive forms. I repeatedly emphasized that workflows capable of sending messages or changing account state require careful review, narrow permissions, and compliance with platform policies.

Where Sol fit best

For me, Sol’s biggest advantage was the combination of direct communication, highly functional prototypes, flexibility during debugging, strong multimedia handling, and browser automation capabilities. I still preferred Terra for some PRD writing and Sonnet 5 for conversational agent voice, especially inside OpenClaw-style workflows.

The benchmark also surfaced recurring stylistic habits. Sol repeatedly defaulted to forest-green interface palettes and recognizable visual patterns, which I treated as a reminder that even strong models still need explicit creative direction and human taste.

The most useful takeaway from this episode of How I AI is not the leaderboard itself. It is the evidence produced by running real work across multiple models and reviewing the outputs closely. Sol looked strongest in workflows where fast iteration, practical implementation, and functional prototypes mattered more than architectural perfection. Terra worked well for concise business writing. Sonnet still handled conversational tone best. The part that remains stubbornly human is judgment: deciding when a prototype is actually useful, when a debugging system has become too rigid, and when automation should stop before it creates new problems.

Build your next product with ChatPRD

Turn an idea into a PRD, user stories, and a plan.

Try ChatPRD free

Start shipping
better products.

Join 100,000+ product managers who use ChatPRD to write better docs, align teams faster, and build products users love.

Free to start
No credit card
SOC 2 certified
Enterprise ready