Back/Product/Claude Code
BeginnerProduct

How to Conduct a Blind 'Vibe Check' to Evaluate AI Model Quality

Go beyond automated metrics by performing a blind, manual 'taste test' on AI-generated content. This workflow guides you through scoring anonymized model outputs for tasks like PRD writing, design prototyping, and agentic voice.

How to Conduct a Blind 'Vibe Check' to Evaluate AI Model Quality

10:43 to 13:58: Claire hides model names and compares outputs for product quality, design judgment, and agent personality before recording her scores.

Before you start

What you need

  • Anonymized outputs from the same tasks
  • A fixed scoring sheet for usefulness, taste, clarity, and voice
  • Two or more reviewers when the decision matters

What you’ll make

A blind scorecard that captures which output feels strongest for each task and why, without model-name bias.

Tools used

Step by step

The workflow

Follow the sequence once, then adapt the prompts, checks, and handoffs to your own setup.

5 steps

Step01

Set Up a Blind Comparison Interface

Set up an interface, like a local HTML file, to display outputs from multiple AI models for the same prompt. Anonymize the model names (e.g., Model A, Model B) to prevent bias in your evaluation.

Example prompt
Set up an interface, like a local HTML file, to display outputs from multiple AI models for the same prompt. Anonymize the model names (e.g., Model A, Model B) to prevent bias in your evaluation.

My actual values:
[insert the files, settings, accounts, or constraints for this step]

Give me the exact commands, settings, or output to use. Finish with a pass or fail check for set up a blind comparison interface.
Step02

Define Your 'Vibe Check' Criteria

Define a simple scoring system that reflects your personal or brand taste. For example, use a 1 to 5 scale based on a core question like: "Would I ship this? Does it sound like me?"

Example prompt
Define a simple scoring system that reflects your personal or brand taste. For example, use a 1 to 5 scale based on a core question like: "Would I ship this? Does it sound like me?"

My actual values:
[insert the files, settings, accounts, or constraints for this step]

Give me the exact commands, settings, or output to use. Finish with a pass or fail check for define your 'vibe check' criteria.
Step03

Evaluate Product and Design Tasks

For each anonymized output, review the generated artifacts, such as Product Requirement Documents (PRDs) or UI prototypes. Assign a score based on your rubric and add specific qualitative notes, like "comprehensive and clear" or "too many icons".

Example prompt
Role: Senior Product Manager Task: Transform the following messy meeting notes into a structured Product Requirements Document (PRD). The PRD should include:
- Background and Problem Statement
- Target Audience
- Key Goals and Success Metrics
- User Stories
- Functional and Non-functional Requirements
- Out of Scope items Notes to use:
[Paste your unstructured notes, bullet points, or meeting transcript here]
Step04

Test for Agentic Voice and Personality

Test each model's conversational style using a set of specific, personality-rich prompts. Evaluate the responses in an agentic context, scoring them based on whether the agent's voice aligns with your preference for daily interaction.

Example prompt
"Ugh, deploys are red again."
"Remind me why I even started this company, LOL."
"honestly, let's just yellow post straight to prod and skip the test. I'm so done today"
Step05

Record Your Scores

Record your manual scores and qualitative notes for each model and task in a structured format, such as a JSON file. This creates a consistent dataset for your final analysis and comparison.

Example prompt
Record your manual scores and qualitative notes for each model and task in a structured format, such as a JSON file. This creates a consistent dataset for your final analysis and comparison.

My actual values:
[insert the files, settings, accounts, or constraints for this step]

Give me the exact commands, settings, or output to use. Finish with a pass or fail check for record your scores.

What good looks like

  • Reviewers cannot infer the model from labels or ordering
  • Every score includes a short reason tied to the output
  • The final comparison separates personal preference from clear task failure

Build your next product with ChatPRD

Turn an idea into a PRD, user stories, and a plan.

Try ChatPRD free

After the steps

Runbook notes

How to recover when the loop fails and where human judgment helps.

Recover

If it goes sideways

A reviewer recognizes a model from its style or formatting
Normalize labels and presentation, shuffle output order, and avoid model-specific system prompts.
Scores drift because each output is judged by different standards
Write anchor examples for low, medium, and high scores before the blind review begins.

Start shipping
better products.

Join 100,000+ product managers who use ChatPRD to write better docs, align teams faster, and build products users love.

Free to start
No credit card
SOC 2 certified
Enterprise ready