Back/Operations/ChatGPT
IntermediateOperations

How to Build a Custom GPT to Evaluate Team Work Like a Manager

Make a custom GPT that gives artifact feedback from a manager's explicit rubric and examples, while keeping the manager responsible for context, coaching, performance decisions, and any feedback that affects a person rather than the work itself.

How to Build a Custom GPT to Evaluate Team Work Like a Manager

Hilary collects before and after slide examples, asks ChatGPT to articulate the differences and become far more specific, turns the resulting criteria into instructions for a custom slide evaluator, scores each criterion from one to five, and beta tests the GPT with a teammate before wider sharing.

Before you start

What you need

  • A narrowly defined artifact such as slides, briefs, or interview plans
  • Representative before and after examples you may share with the model
  • An explicit rubric with observable criteria, examples, and exceptions
  • A privacy safe GPT workspace and rules for confidential company material
  • A beta tester, feedback channel, and manager owner for ongoing calibration

What you’ll make

A reusable feedback assistant that scores and explains artifact criteria, cites specific passages or elements, proposes actionable revisions, states uncertainty, and routes contextual or people related judgment back to the manager.

Tools used

Step by step

The workflow

Follow the sequence once, then adapt the prompts, checks, and handoffs to your own setup.

5 steps

Step01

Collect 'Good' and 'Bad' Examples

Collect permissioned examples of the same artifact type, including weak drafts and the manager's edits. Remove author identity and sensitive content, and label the context and why each change improved or harmed the work.

Step02

Reverse-Engineer Your Criteria with AI

Ask AI to compare the examples and propose observable criteria. Challenge vague phrases and separate general quality, team convention, audience need, and personal preference.

Example prompt
Compare these paired before and after artifacts. Identify observable differences, the likely purpose of each edit, exceptions, and candidate criteria. Separate general communication quality, team conventions, audience specific needs, and subjective preference. Do not infer anything about the authors.
Step03

Refine and Get Hyper-Specific

Turn the promising criteria into a specific rubric with definitions, positive and negative examples, evidence requirements, scoring anchors, and a not enough context outcome.

Example prompt
Make this rubric 100 times more specific. For each criterion define what it measures, why it matters, observable evidence, one through five anchors, examples, exceptions, and when to return NOT_ENOUGH_CONTEXT. Keep all judgment about the artifact, not the person.
Step04

Generate the GPT Instructions

Generate custom GPT instructions that ask for the artifact purpose and audience, score the rubric with citations, prioritize the most important revisions, and explain rather than simply rewrite.

Example prompt
Write instructions for a custom GPT that evaluates [artifact] using this rubric: [rubric]. It must request purpose and audience, cite exact evidence, score with anchors, explain why each issue matters, propose specific revisions, state uncertainty, and never assess employee performance, intent, or potential.
Step05

Build, Test, and Deploy the GPT

Test the GPT on held out examples with one teammate, compare its feedback with the manager's, and collect disagreements. Revise the rubric and examples before broader sharing, then review use and failure reports regularly.

What good looks like

  • The rubric describes observable work qualities rather than attempting to imitate a manager's personality or private intuitions.
  • Feedback cites the artifact, explains why the criterion matters, and suggests a concrete next attempt.
  • The GPT does not make performance ratings, promotion decisions, or claims about an employee's intent or capability.
  • Beta users find the feedback useful, consistent, and easy to challenge, and the owner updates the rubric when patterns recur.

Build your next product with ChatPRD

Turn an idea into a PRD, user stories, and a plan.

Try ChatPRD free

After the steps

Runbook notes

How to recover when the loop fails and where human judgment helps.

Recover

If it goes sideways

Examples encode one manager's preferences as universal quality or disadvantage a style or culture
Use diverse examples, name subjective criteria, invite challenge, and review the rubric with the team.
The GPT turns artifact feedback into judgment about the employee
Constrain evaluation to observable work, remove identity data, and keep coaching and employment decisions with the manager.
A one through five score creates false precision or becomes a hidden performance metric
Treat scores as navigation, require evidence, and prohibit aggregation or reuse for performance management.
Sensitive strategy, customer, personnel, or unreleased work is uploaded to an unapproved GPT
Use an approved environment, redact unnecessary context, and document what material may be submitted.

Start shipping
better products.

Join 100,000+ product managers who use ChatPRD to write better docs, align teams faster, and build products users love.

Free to start
No credit card
SOC 2 certified
Enterprise ready