Back/Design/Claude
IntermediateDesign

How to Prototype Conversational AI Flows Using Claude's Golden Conversations

Design a conversational AI feature from the experience backward: write golden conversations, test them across representative scenarios, define the quality rubric they reveal, and turn the strongest flow into a clearly labeled interactive prototype.

How to Prototype Conversational AI Flows Using Claude's Golden Conversations

Priya starts with golden conversations for a photo-assisted Yelp service flow, tests several real-world scenarios, turns the patterns into a rubric, and builds a Claude Artifact to feel the interaction in a chat interface.

Before you start

What you need

  • A specific conversational job and target user
  • The assistant capabilities and hard product constraints
  • Authorized or synthetic scenarios that cover normal and difficult cases
  • A required conversation format and termination condition
  • A rubric for correctness, usefulness, safety, clarity, and efficiency

What you’ll make

A set of source-labeled golden conversations plus an interactive prototype that exposes the proposed system behavior, UI rhythm, and unresolved product decisions.

Tools used

Step by step

The workflow

Follow the sequence once, then adapt the prompts, checks, and handoffs to your own setup.

4 steps

Step01

Generate Initial Sample Conversations

Describe one representative user situation, the assistant capabilities, the information needed before completion, and the output format. Ask for a continuous role-labeled conversation rather than a generic feature list.

Example prompt
Write one complete sample conversation for [user situation]. Use "User:" and "Assistant:" labels. The assistant may [capabilities] and must [constraints]. Continue until it has enough confirmed information to [completion action]. Show uncertainty and ask a clarifying question when the input is ambiguous.
Step02

Test with Real-World Scenarios and Iterate

Run the same flow against authorized or synthetic examples that vary by intent, input quality, risk, and category. Keep the source for each case and inspect recognition, questions, advice, tone, and stopping behavior.

Example prompt
Create numbered golden conversations for these scenarios: [scenarios]. For each, preserve the scenario ID, explain any uncertainty, and do not infer details that the input does not support. Include one ambiguous case, one unsupported case, and one case that requires escalation.

Pay attention to the LLM's 'thought process' if available. It provides valuable clues for debugging and understanding how the model is interpreting your prompt.

Step03

Refine and Polish Conversations

Compare the conversations and define a rubric from the recurring strengths and failures. Rewrite the examples against that rubric while retaining both the original scenario and a concise change note.

Example prompt
Evaluate these conversations for input interpretation, question quality, factual support, safety, concision, user control, and completion behavior. Propose observable rubric criteria, then revise each conversation. List what changed and which criterion prompted the change.
Step04

Create an LLM-Powered Interactive Prototype

Build a Claude Artifact from the reviewed conversations and interface references. Label simulated actions, use synthetic data, test waiting and error states, and record where the prototype differs from the intended production system.

Example prompt
Create a mobile chat prototype based on these reviewed conversations and reference screens. Include the proposed system instructions, suggested replies, input controls, loading, error, low-confidence, and completion states. Use synthetic data and label mocked actions. Do not connect production accounts, data, or services.

Testing an interactive prototype gives you a much better feel for the user experience than a static text document. Pay attention to response length and speed in a mobile interface.

What good looks like

  • The examples cover meaningfully different user intents, inputs, and failure conditions.
  • Image or document interpretations state uncertainty and avoid unsupported diagnosis or safety claims.
  • The rubric turns qualitative reactions into observable criteria that can later support evals.
  • The artifact is treated as a prototype and does not imply production data handling, model behavior, or feasibility.

Build your next product with ChatPRD

Turn an idea into a PRD, user stories, and a plan.

Try ChatPRD free

After the steps

Runbook notes

How to recover when the loop fails and where human judgment helps.

Recover

If it goes sideways

Every golden conversation uses a clear, cooperative user
Add terse, ambiguous, misspelled, unsupported, and interrupted scenarios based on observed patterns or carefully designed synthetic cases.
The model confidently misreads a photo or recommends an unsafe action
Require uncertainty, clarifying questions, and an escalation path for hazardous or low-confidence situations.
The examples become rigid scripts that hide viable alternatives
Record the goal and rubric for each turn, then include more than one acceptable path where the product allows it.
The Claude Artifact is mistaken for evidence that the production stack can deliver the same behavior
Label mocked behavior and document the model, tools, latency, privacy, and engineering work still needed.

Start shipping
better products.

Join 100,000+ product managers who use ChatPRD to write better docs, align teams faster, and build products users love.

Free to start
No credit card
SOC 2 certified
Enterprise ready