Back/Design/Braintrust
IntermediateDesign

How to Scale Expert Judgment in AI Systems with a Human Feedback Loop

Implement a feedback loop to translate subjective expert feedback into quantitative evaluation criteria. This workflow allows you to scale the 'taste' of key stakeholders, like a lead designer, across your entire AI system, continuously raising the quality bar.

How to Scale Expert Judgment in AI Systems with a Human Feedback Loop

From 23:06 to 33:12, Ankur Goyal demonstrates an evaluation loop that turns expert feedback and designer taste into reusable scoring criteria. Clip range: 23:06 to 33:12.

Before you start

What you need

  • Evaluation dataset for the target AI task
  • Quantitative scoring criteria such as accuracy or concision
  • Human domain expert or design reviewer
  • Examples of reviewed AI outputs
  • Rubric storage location for updated evaluation criteria

What you’ll make

An expanded evaluation rubric that converts subjective expert feedback into measurable scoring criteria.

Tools used

Step by step

The workflow

Follow the sequence once, then adapt the prompts, checks, and handoffs to your own setup.

3 steps

Step01

Run Quantitative Evals to Establish a Baseline

First, define a set of measurable criteria for your AI's output, such as accuracy, concision, or code quality. Use a quantitative evaluation system to test the model against a dataset of inputs until it consistently reaches a high baseline score, for example, 90% or higher.

Step02

Conduct a Human 'Vibe Check'

Once the quantitative scores are high, present a representative set of AI-generated outputs to your team's designated 'taste maker,' such as a lead designer or domain expert. The goal is a final qualitative review or 'vibe check' to catch issues that the numbers miss.

Step03

Capture and Encode Feedback into New Evals

When the expert provides subjective feedback, such as 'The tone is helpful but feels patronizing,' translate that qualitative insight into a new, measurable evaluation criterion. Add this new criterion to your evaluation suite to systematically test for this quality in all future runs.

Example prompt
You are an expert in creating AI evaluation criteria. Your task is to translate subjective human feedback into a new, measurable scoring rubric that another AI can use for evaluation.

Here is the context:
- Original Task: [Describe the AI's task, e.g., "Answer user questions about our documentation."]
- Example AI Output: [Paste the AI-generated output that was reviewed.]
- Expert Feedback: [Paste the expert's qualitative feedback, e.g., "The tone is helpful but feels patronizing."]

Based on this, generate a detailed scoring rubric from 1 to 5 for the quality of "[describe the quality, e.g., 'professional and helpful tone']". For each score, provide a clear definition and an example of what that score would look like. The rubric should be clear enough for an LLM to apply it consistently.

What good looks like

  • Baseline evaluation scores are recorded before qualitative review
  • Human feedback is tied to specific example outputs
  • New rubric defines scoring levels with concrete examples
  • Updated evaluation suite can score future outputs against the new criterion

Build your next product with ChatPRD

Turn an idea into a PRD, user stories, and a plan.

Try ChatPRD free

After the steps

Runbook notes

How to recover when the loop fails and where human judgment helps.

Recover

If it goes sideways

Subjective feedback is too vague to operationalize
Request examples of acceptable and unacceptable responses from the reviewer
New rubric conflicts with existing evaluation goals
Test the rubric against historical outputs and adjust scoring weights
Evaluation data overrepresents one user scenario
Expand the dataset with additional task variations before benchmarking

Start shipping
better products.

Join 100,000+ product managers who use ChatPRD to write better docs, align teams faster, and build products users love.

Free to start
No credit card
SOC 2 certified
Enterprise ready