Define Your Benchmark Tasks
Identify 3-7 core tasks you perform regularly, such as drafting documents, coding, summarizing notes, or triaging your inbox.
Systematically evaluate AI models like GPT-6 and Opus 5.5 to find the best fit for your tasks. This blind test method helps you choose tools based on real-world performance and personal preference, not just industry benchmarks.

OpenAI's cloud-based AI software engineering agent that can execute code, run tests, and handle complex multi-file tasks autonomously.
Anthropic AI assistant
Step by step
Follow the sequence once, then adapt the prompts, checks, and handoffs to your own setup.
6 steps
Identify 3-7 core tasks you perform regularly, such as drafting documents, coding, summarizing notes, or triaging your inbox.
Use a tool or ask a colleague to run your prompts on different models and label the outputs anonymously (e.g., Model A, Model B) so you can't tell which is which.
Submit the exact same prompt for each of your benchmark tasks to every model you are testing. This ensures a fair and direct comparison of their outputs.
Review each output without knowing the source model. Rate them on a personal 'vibe' scale (e.g., 1-5) and add notes on what you liked or disliked about the style, structure, and quality.
Un-blind the model names and aggregate your scores in a spreadsheet. Calculate average scores and identify which models performed best on average and on specific tasks.
Based on your analysis, assign specific models to specific tasks in your daily workflow, optimizing for both quality and cost to build your personal AI toolkit.
Turn an idea into a PRD, user stories, and a plan.
Keep building

Quickly analyze thousands of YouTube comments for sentiment and new content ideas using the Jev AI model. This workflow helps creators turn audience feedback into actionable insights for their content strategy.

Create a powerful product insights engine by combining Jev with other LLMs. Use Jev for large-scale classification of customer feedback and development data, then apply reasoning models to analyze the structured output for deep insights.

Use the Jev AI model to cheaply and quickly cluster thousands of GitHub pull requests into thematic areas. This helps engineering and product leaders understand the breakdown of work between new features, tech debt, and maintenance.
Join 100,000+ product managers who use ChatPRD to write better docs, align teams faster, and build products users love.