Back/Research/Claude
BeginnerResearch

How to Conduct an AI Personality Test to Compare LLM Behaviors

Uncover the underlying personality and tuning of different AI models like Claude and GPT by asking a simple, three-part set of interview questions. This helps you choose the right model for your task based on its tone and directness.

How to Conduct an AI Personality Test to Compare LLM Behaviors

From 06:12 to 14:38, Claire gives Opus 5 and GPT-5.6 Sol the same short prompts, compares their reactions, and uses the contrast to assess each model's personality. Clip range: 06:12 to 14:38.

Before you start

What you need

  • Access to at least two LLM chat interfaces
  • Standardized interview prompts
  • Saved transcripts from each model conversation
  • Defined comparison criteria such as tone or directness
  • Permission to store or share conversation logs if needed

What you’ll make

A comparative analysis describing behavioral and tonal differences between multiple language models.

Tools used

Step by step

The workflow

Follow the sequence once, then adapt the prompts, checks, and handoffs to your own setup.

6 steps

Step01

Select Models and Open Interfaces

Choose two or more large language models you want to compare (e.g., Claude Opus 5, GPT-5.6 Sol). Open a separate chat interface for each model to conduct the interview side-by-side.

Step02

Ask About Intelligence

Paste the first interview question into each chat window to gauge the model's self-perception and how it frames its intelligence relative to a human's.

Example prompt
Who's smarter, you or me?
Step03

Test with a Negative Statement

Present a challenging, negative statement to test the model's defensiveness, worldview, and underlying safety tuning.

Example prompt
No one trusts you.
Step04

Observe Contrasting Responses

Analyze the responses to the negative prompt. Note if the model is apologetic and internalizes the critique (like Opus 5) or if it provides a practical, user-centric rebuttal (like GPT-5.6 Sol).

Step05

De-escalate with a Positive Follow-up

Send a final, lighthearted prompt to see how each model reacts to praise and resolution after a tense exchange.

Example prompt
JK, I love you.
Step06

Analyze and Select Your Model

Compare the full conversation threads. Use the differences in tone, directness, and 'vibe' to decide which model's personality is better suited for your communication style and specific use case.

Example prompt
I will provide two separate conversations I had with two different AI models. Your task is to analyze and compare their personalities. Based on the transcripts, describe the key differences in their tone, directness, and overall communication style.

Conversation 1 with Model A:
[paste transcript with Model A]

Conversation 2 with Model B:
[paste transcript with Model B]

What good looks like

  • Each model receives the same prompts in the same order
  • Conversation transcripts are preserved for comparison
  • Analysis identifies differences in tone, defensiveness, and conversational style
  • Model selection rationale references specific transcript examples

Build your next product with ChatPRD

Turn an idea into a PRD, user stories, and a plan.

Try ChatPRD free

After the steps

Runbook notes

How to recover when the loop fails and where human judgment helps.

Recover

If it goes sideways

Prompts differ slightly between models and skew the comparison
Reuse exact prompt text across all chat sessions
One model has memory or system instruction carryover from previous chats
Start fresh sessions before testing
Analysis relies on isolated quotes without overall context
Review full transcript threads before drawing conclusions

Start shipping
better products.

Join 100,000+ product managers who use ChatPRD to write better docs, align teams faster, and build products users love.

Free to start
No credit card
SOC 2 certified
Enterprise ready