Back/How I AI
How I AI

How I AI: My Hands-On Review of GrokBot, Cursor Origin, and the Grok 4.6 Model

I'm testing the entire Grok ecosystem, from the multi-agent workspace in GrokBot to the agent-native code hosting of Cursor Origin, and putting the Grok 4.6 model through my personal design and coding benchmark.

Claire Vo's profile picture

Claire Vo

August 18, 2026·10 min read

Full episode

Watch or listen

Episode outline

I don't know if you've noticed, but it seems like everyone I know is secretly becoming a Groq boy. The whole OpenAI versus Anthropic rivalry feels like old news. Ever since the (entirely fictional) acquisition of Cursor by SpaceX, the buzz around the products from the Elon cinematic universe of code builders has been impossible to ignore.

I’ve been hearing more and more positive things about the Groq models, so I decided it was time to put the whole ecosystem to the test. In this special episode breakdown, I’m going through the latest releases from the Cursor/xAI team to give you my honest take on what works, what doesn’t, and where the real potential lies.

We’re going to look at three key things: GrokBot, their new chat-style agent; Origin, their ambitious GitHub competitor; and of course, the Groq 4.6 model itself, which I ran through my personal benchmark gauntlet. Let's get to it.

Workflow 1: Using GrokBot for Multi-Account Knowledge Work

First up is GrokBot, a chat-style agent available as a desktop and mobile app. Think of it as a simpler, hosted alternative to something like OpenClaw, designed for knowledge work. The UI is clean and streamlined, pulling from that familiar iMessage-style chat experience. But what I really appreciate is its commitment to a multi-agent philosophy. Instead of one agent to rule them all, GrokBot encourages you to create specialized agents for specific jobs, each with its own name and purpose.

The list of my custom GrokBots, including Prody McProd, commitment tracker, and money maker bot

The Killer Feature: Multiple Accounts Per Connector

If I had to point to one piece of magic in GrokBot, it's the plugins—specifically, how it handles them. I've always thought Cursor had the best client for connecting to different services in my stack, and they’ve brought that expertise to GrokBot with a game-changing improvement: you can connect multiple accounts for a single service.

This is huge. I have a dozen different email addresses and Slack accounts, and no other tool seems to understand this reality. With GrokBot, I can connect all four of my Gmail accounts and all seven of my Slacks, and my agents can work across all of them. This feature alone would have product-market fit with me.

The GrokBot connector settings showing multiple Gmail accounts configured and active

How I'm Using GrokBot

Setting up a new bot is incredibly simple. You just create a bot and give it instructions. For example, I created a new one on the fly:

I want you to be a family manager bot that keeps track of my personal email and personal calendar.

It immediately understands, identifies the necessary connectors, and gets to work. I’ve already set up several specialized agents:

  • Prody McProd: My product manager bot, hooked up to ChatPRD.
  • Commitment Tracker: An out-of-the-box bot that keeps me honest about my promises.
  • Money Maker: Chases invoices, payments, and sales deals.
  • Case Study Buddy: Helps me with a big batch of case studies I'm writing.
  • Data Monitor: Watches my ChatPRD data for important trends.

Each GrokBot also comes with its own virtual computer, complete with Chrome, a terminal, and a file system. It’s like a lightweight version of OpenClaw, giving the agent the ability to browse the web and execute commands.

The GrokBot virtual machine interface showing the Chrome browser, Terminal, and Files icons

My Verdict: Simple, Powerful, but Not for Hackers

The biggest pro of GrokBot is that it just works. The onboarding is smooth, the connectors are top-tier, and the multi-agent approach is intuitive. However, that simplicity is also its biggest drawback for me. I love my chaotic, high-maintenance OpenClaw agents because I can tune and hack them endlessly.

The joy is in the challenge. And I will tell you, GrokBot has not challenged me because it works simply and it works out of the box.

With GrokBot, you have less control. You can't choose the underlying model, and you can't get into the weeds of its configuration. The personality of the model it uses also has what I call "Claude Slop"—just no vibes. It's not a bot I'd want to hang out with. If you're looking for a highly tunable, transparent agent setup, you're better off with something like OpenClaw or Hermes. But if you want a powerful, multi-account agent that's easy to set up, GrokBot is seriously impressive.

Workflow 2: Evaluating Cursor Origin, the Agent-Native GitHub Replacement

Next up is Origin, the new GitHub replacement from the Cursor team. The core idea is to build a code hosting platform that is agent-native from the ground up. It has all the Git primitives we know—repos, diffs, pull requests—but the user experience is redesigned for collaboration with AI agents, particularly Cursor's own cloud agents and CLI.

Cursor's bet is that as we rely more on agents to write and review code, we'll need a platform built for that new reality, and that GitHub won't get there fast enough.

Getting Started and First Impressions

You access Origin (currently called Code Base in the app) through the Cursor web app. The first step is to import your existing repositories from GitHub. You just authorize your account, pick the repos you want to sync, and they appear in Origin. I had a bit of trouble with this, as I was testing it on a day GitHub was having a major outage—either terrible timing or a stroke of marketing genius by the Cursor team.

The Cursor Origin UI showing the 'Sync from GitHub' button and a list of imported repositories

Once imported, the experience feels a lot like GitHub, but with a Cursor-centric redesign. The PR view, for example, has some nice affordances for showing feedback from agents like bugbot. You can also @Cursor in a comment to ask it to perform tasks, like fixing a failing Vercel preview branch.

A pull request view in Origin highlighting a comment from an AI agent reviewer

My Verdict: A Foundation for the Future, But Not a Must-Have Today

Right now, Origin is in a very early beta, and it shows. For anyone deeply invested in the GitHub ecosystem with custom actions, code owners, and complex CI/CD pipelines, Origin feels more like a limited, slightly slower wrapper on the GitHub API. It's a nice redesign, but there isn't a compelling, must-have feature that would make me switch my team over today.

I can see the vision. They're laying the groundwork for a world where agents are first-class citizens in the development lifecycle. Having a tightly integrated experience between your code host and your AI coding assistant makes perfect sense. But for now, it's one to watch, not one to migrate to. It will be fascinating to see if they can build enough value to pull developers away from GitHub's massive gravity.

Workflow 3: The Benchmark - Putting Grok 4.6 to the Vibe Test

Finally, the moment you've been waiting for. How good is the actual Groq 4.6 model? To find out, I ran it through my "How I AI Vibe review," a custom benchmark I use to test new models. I don't just look at leaderboards; I test them on real-world tasks I do every day: writing PRDs, building prototypes, creating designs and wireframes, and making technical changes. And, of course, I evaluate whether I actually like talking to the model.

I run all these evaluations blind, grading the outputs myself before I know which model produced what. For this round, I even added two new tests: a redesign task where the model has complete creative freedom, and a wireframe for a highly complex claims adjudication UI.

The "Claire Weighted Index" Results

After grading everything, I put it all into my Claire Weighted Index, which is 70% my personal taste and 30% the taste of an LLM-as-a-judge. The results were genuinely surprising.

Grok 4.6 came in second place, just behind my favorite, GPT-5.6 Sol, and ahead of both Sonnet 5 and Opus 5. I did not expect it to rank so highly.

The final 'Claire Weighted Index' chart showing the ranking of models, with Grok 4.6 in the number two spot

Breaking it down by task, my preferences were clear:

  • PRDs & Prototypes: GPT-5.6 Sol remains my favorite for its clean, comprehensive, and matter-of-fact style.
  • Technical Implementation: The LLM-as-a-judge rated Opus 5 as the best for a technical bug triage problem.
  • Agent Chit-Chat: Sonnet 5 continues to be my undefeated champion for pithy, enjoyable back-and-forth in an OpenClaw agent.
The 'Overall Recommendation by Task' slide from the benchmark presentation

Where Grok 4.6 Really Shines: Design Freedom

So where did Grok 4.6 excel? In design tasks where it was given broad creative freedom. My theory is that I can spot GPT and Claude "slop" a mile away—GPT-5.6 loves forest green, while Claude models gravitate toward a brown, tan, and orange palette. Grok 4.6 felt like a breath of fresh air.

My favorite output from it was an adorable and functional interactive ordering system for a coffee shop. It was creative and well-designed. In contrast, when following very specific art direction or designing highly complex UIs (like a container ship docking app), I still prefer GPT-5.6 Sol. It's simply the best at managing density and technical completeness without being overwhelming.

The side-by-side comparison of different model designs, highlighting Grok 4.6's creative coffee shop UI

Interestingly, while I was impressed with Grok 4.6, the LLM-as-a-judge (I use GPT-5.5 because it's the harshest grader) absolutely hated it and loved the Claude models. It just goes to show how much personal taste matters in these evaluations.

Conclusion: It's Time to Pay Attention to Groq

After spending a week with these new products, my main conclusion is that the Groq model is a legitimate competitor. GrokBot is a fantastic tool for anyone who wants a simple but powerful multi-account agent and doesn't need deep hackability. Cursor Origin is an interesting bet on the future, though it's too early to call. And the Groq 4.6 model itself is genuinely good, especially for creative design tasks.

I'm still spending most of my time in Codex and with the GPT-5.6 models, but this experience has convinced me to spin up Cursor more often for coding. It seems it might be time for all of us to become, at least to some degree, a Groq boy (or girl, or person).

I'd love to hear from you. Have you tried GrokBot? Do you think Origin can replace GitHub? Have you switched to Groq for any of your coding tasks? Let me know in the comments.

---

A word from our sponsors

A big thank you to our sponsors for making this episode possible.

  • bolt.new: This episode is brought to you by bolt.new, the AI app builder for people who have ideas and want to ship them. You describe what you wanna build, and Bolt generates production-ready code in minutes. Connect Stripe, hook up your domain, and deploy it live. You just need an idea and a weekend. Check it out at bolt.new/howiai.
  • Jira by Atlassian: This episode is brought to you by Jira by Atlassian. The teamwork graph in Jira delivers forty-four percent more accurate agent results with forty-eight percent less token usage. The teamwork graph pulls context from across your entire stack—from Jira and Confluence to GitHub—and feeds it directly to your agents before they write a single line. Same team, smarter agents. Try them free at jira.dev. That's J-I-R-A dot D-E-V.

---

Build your next product with ChatPRD

Turn an idea into a PRD, user stories, and a plan.

Try ChatPRD free

Start shipping
better products.

Join 100,000+ product managers who use ChatPRD to write better docs, align teams faster, and build products users love.

Free to start
No credit card
SOC 2 certified
Enterprise ready