My Jev Playbook for Data Analysis, Product Insights, and Real-Time Apps
How I use Jev to analyze GitHub PRs, organize product feedback, find YouTube audience insights, and build a real-time voice-to-color app.
Claire Vo
Full episode
Watch or listen
Workflows from this episode
- How to Analyze YouTube Comments for Audience Insights with AI
- How to Build a Multi-Model AI Product Insights Engine
- How to Analyze GitHub PRs to Understand Engineering Effort with AI
Episode outline
Jev, Jev, Jev. Welcome to Jev week on How I AI. In the past week alone, we’ve seen a flood of new models hit the scene—Opus 5.5, GPT-6 Sol, GPT-6 Luna, Muse is all over the timeline. And yet, the one model I can’t stop talking about is this incredibly fast, cheap, “System 1” decision model from TypeSafe AI called Jev.
As soon as I saw Jev launch, I started testing it, and I have to say, it has unlocked more personal productivity, coding, and product use cases for me than any other model lately. I've used it to accomplish work that I believe is worth hundreds of thousands of dollars, and I've spent less than $10 in tokens to do it. Some of that usage was subsidized through Vercel AI Gateway during my testing; check current pricing before planning your own runs.
Today, I’m giving you a tour of what Jev is, what it does (and what it doesn’t do), and several workflows I’ve built with it in the last week. From analyzing datasets to organizing product feedback, Jev is finding its way into much of what I build.
What Exactly is Jev?
So, what makes Jev different from the standard LLMs we’re all used to? The best way to understand it is to look at the comparison table on the TypeSafe blog. While both standard LLMs and Jev take in unstructured text, their outputs are fundamentally different.
- Standard LLMs: Text in, text out. You give them a prompt, and they generate strings of text.
- Jev: Text in, TypeSafe values out. You give Jev text, and it returns a predefined, structured value. It doesn't write prose; it makes a decision.

Jev essentially returns one of three things, which you can read about in their docs on primitives:
- Choice: You provide a list of options, and Jev picks one. For example, if I ask, "What should I wear on this date?" and give it the choices
[dress, jeans, workout clothes], it will pick one. - Score: It ranks something on a scale. This is perfect for severity rankings, like triaging a bug as
cosmetic,broken, orblocking. - Noul: This returns a probability between 0 and 1 that a statement is true. Ask it "Is Claire a podcaster?" and it can return a value such as 0.99.
This might sound simple, but when you're a software engineer, you realize that 90% of what you build involves making choices, scoring things, and routing based on yes/no decisions. The other incredible thing about Jev is that it’s dirt cheap—just four cents per million input tokens—and blazing fast. This opens up classification and decision-making at a scale that was simply too expensive or slow before.
Workflow 1: Large-Scale Data Analysis for Pennies
The number one thing I've been using Jev for is classifying vast sets of unstructured data. This is work that would have been incredibly annoying to do manually but provides immense value.
Analyzing GitHub PRs to Understand Engineering Effort
As a CPO or CTO, you’re constantly asked by the board, “What percentage of our work is going to tech debt versus new features?” Answering that used to be a nightmare of spreadsheets and guesswork. Now, it costs nine cents.
Here’s the workflow I built using Codex to orchestrate it:
- Pull the Data: First, I connected to GitHub and pulled every single pull request from a repository.
- Cluster with Jev: This is where Jev shines. Instead of asking it to categorize each PR from scratch, I had it compare candidate pairs. For each pair, I asked whether the PRs addressed the same thematic area. Jev's answers helped group related items without needing predefined categories.
- Label with a Cheap LLM: Once Jev created the clusters, I used a cheap and fast model (in this case, Gemini Flash Lite) to look at each cluster and assign it a descriptive label.
The Results:
I first ran this on our marketing site, which had 112 PRs. It cost me 1.1 cents and gave me a clear breakdown of work across maintenance, content, tools, and more.
Then, I ran it on my main app, ChatPRD, which had about 1,700 PRs this year. It analyzed all of them, found 17,000 potential pairs, and clustered them. The cost? Nine cents. The whole process took about two minutes.

The analysis revealed that almost 30% of our effort goes to platform, security, and infrastructure—which is great to see! This kind of insight is invaluable for strategic planning and reporting.
Analyzing My Local Claude Code and Codex Sessions
You can apply this same technique to your own data. I ran a similar analysis on the Claude Code and Codex session records available locally on my machine.
I had Jev categorize my activity based on user turns. The results showed a fascinating shift in my work. In January, I was almost exclusively doing engineering tasks. By September, engineering was less than 40% of my work, with more time spent on agents, publishing, media, and my new business.

Quick Triage for Personal Email
I even ran this on my personal Gmail. I had Jev look at the subject line and snippet of each email and simply asked it to score whether or not I can delete this email. It quickly tagged a huge volume of messages that could be safely deleted, allowing me to focus on what actually mattered. This leads me to my next big point...
Workflow 2: Building a Product Insights Graph with Jev's "Buddy System"
Jev alone is okay. Jev with an LLM buddy is super powerful.
This is the key takeaway. I like to use Jev to take a massive corpus of information, and then tag, categorize, cluster, and filter it. Once the data is organized, I can apply more powerful AI actions to the right clusters. The best example of this is the product-insights graph I’m building for ChatPRD.
The goal is to suck in all our business data—support tickets, GitHub PRs, user conversations, Linear tickets—and generate insights. I had spent tens of thousands of dollars trying to brute-force this with frontier models, but it was too expensive and the nuance was too high. Jev cracked it for me.
Here’s the multi-model architecture:

- Classification (Jev): I pull in about 1,100 individual signals. Jev’s job is to do all the heavy lifting of classification, clustering, and mapping. It performs over 200,000 classification and pairwise groupings to organize this mess of data.
- Analysis (Astra): Once Jev creates structured groups, I pass them to a powerful reasoning model like Astra and ask the big-brain questions:
"What the heck's going on across all of this?" - Generation (Sol/Luna): Finally, for generating user-facing summaries or text, I use a creative model like Sol or Luna.
This system allows me to find the gap between what customers are telling us and what we’re actually working on. The cost for Jev’s part in processing all this data? About four dollars. This architecture has made the entire feature more margin-accretive and unlocked a product that was previously too difficult and expensive to build.
Workflow 3: Creating Interactive Dashboards and Real-Time Apps
Jev's speed makes it perfect for use cases that require real-time feedback, which was always a challenge with slower, more expensive models.
Analyzing How I AI YouTube Comments
I wanted to understand what you, our wonderful audience, are saying in the YouTube comments. There are about 4,500 of them, so I hooked up the YouTube v3 API and used Jev to analyze them.
- Data: Pulled all 4,500 comments.
- Classification: Used Jev for two tasks:
-
Categorize them into positive, negative, or neutral. -
Identify whether or not the comment includes an idea for a future episode.
- Dashboard: The categorized data fed a live dashboard.

The results were awesome. I discovered that about half the comments are positive, and I found 58 with fantastic episode ideas! It also allowed me to build a live search over the comments. When I search for a term like "slop," it scans the comment collection and returns the relevant ones.
A Real-Time Voice-to-Emotion App
To really show off the real-time capabilities, I built a fun little app in an afternoon. It takes my voice as input, uses Jev to determine the emotion, maps that emotion to a color, and displays a relevant quote.
The stack is OpenAI's real-time voice API, Jev, and a quote API. Here's how it works:
- The app ingests a phrase from the real-time voice API.
- Jev is given the text and a list of hex color values. It
scoresthe colors to find the best match for the emotion in the text. - Jev also helps filter and score quotes from the quote API based on the detected sentiment.
- The app displays the quote with the corresponding colored background.

Watching it react as I spoke was magical. You can see the slight latency from the quote API, but Jev’s decision is instantaneous. It’s a great example of the kinds of interactive, magical experiences you can now build.
The Future is Fast and Decisive
These are just a few of the ways I've started using Jev, and it's already fundamentally changing how I approach building with AI. By using it as a smart, fast, and cheap filter, you can better orchestrate more powerful models to do what they do best: complex reasoning and generation.
I hope this has unlocked some ideas for how you can use Jev on your own data, whether it's for business analytics, personal productivity, or building fun new applications. I truly believe these small, decisive models are a huge piece of the puzzle.
This is just part one of Jev week! I'll be back midweek with one of our most popular guests to dive even deeper. If you're excited, subscribe and let me know what questions you have about Jev in the comments.
Thanks for joining How I AI!
Sponsors
This episode is brought to you by Open Art Arena, the global leaderboard for creative intelligence. Every week, new AI models launch, and everyone claims to be the best. But best at what? Open Art Arena is built to answer the question that actually matters: which model is best for your specific job? Instead of one overall winner, Open Art Arena ranks models across real creative tasks, from advertising and film to animation, product, graphic design, editing, and lip sync, covering both image and video. And these rankings aren't based on hype. They're judged by professionals, industry leaders, and working creators through blind evaluations, so judges never know which model produced which output. That means you can see how models actually perform when it comes to the creative work you're doing. So stop guessing which model to use. Explore rankings based on real creative work and find the right model for your project and save time and cost. See the rankings at Open Art Arena.
About Claire Vo
Claire Vo is the host of How I AI, CEO of ChatPRD, and an expert in AI product strategy. You can find her on X and LinkedIn.
Watch or listen
Build your next product with ChatPRD
Turn an idea into a PRD, user stories, and a plan.


