How I AI: Hilary Gridley's Playbook for Scaling Yourself with Custom GPTs
Discover how Hilary Gridley, Head of Core Product at WHOOP, builds custom GPTs that think like her to evaluate slide decks and uses AI as a writing coach to make her team's ideas go viral. This episode is a masterclass in using AI to scale your expertise as a manager.
Claire Vo
Full episode
Watch or listen
Workflows from this episode
- How to Use AI as a Sparring Partner to Improve Your Business Writing
- How to Build a Custom GPT to Evaluate Team Work Like a Manager
Episode outline
Whoop Head of Core Product Hilary Gridley turns recurring management feedback into reusable custom GPTs. Her examples cover slide critique and writing support, two tasks where a rubric and a strong set of examples can make an early review more consistent.
Hilary has built many specialized GPTs. They make her criteria available before a formal review, while she keeps responsibility for context, coaching, evaluation, and the decisions that affect her team.
In this episode of How I AI, Hilary derives a slide rubric from before-and-after examples, packages it as a custom GPT, and uses ChatGPT to clarify and challenge a written argument without handing over authorship.
Build a custom GPT for first-pass slide feedback
The manager-feedback GPT workflow starts with Hilary's own before-and-after edits, turns the recurring choices into a draft rubric, and calibrates the result with the team before anyone relies on it.
Hilary begins by making her own standards explicit. The custom GPT can apply those standards repeatedly, but it does not reproduce all of her judgment or replace a conversation with the person doing the work.
Pair early drafts with edited versions
For slide-deck feedback, Hilary creates a two-column document with examples before and after her edits.
- Before: Early slides that show common problems or decisions she would change.
- After: The same slides revised to reflect her preferred hierarchy, narrative, and level of specificity.
She exports the pairs as a PDF and uploads it to ChatGPT. The examples should be approved for the selected AI workspace and stripped of confidential strategy, personal data, customer information, and employee performance details that are not necessary for the task.

Ask the model to infer candidate criteria
Hilary starts with a simple prompt so the model can surface patterns in the examples before she names every criterion herself.
"I kind of want the AI to start by interpreting this in ways that I might not even be able to predict. And then I'm gonna get an intune it and that's when I get super, super specific."
Her initial prompt asks the model to compare the pairs and describe what changed:
Here are some examples of slides. In one column the slides are not good. And in the other column I have edited them and made them good. Can you help me articulate the principles I am using to determine what is good and bad based on these examples? Use these examples.
The model identifies candidate principles such as succinct headlines, intentional visual hierarchy, and one idea per slide. These are observations from the supplied examples, not a trained model or an objective definition of quality.

Refine the rubric until it is specific
Hilary challenges vague criteria and adds the priorities the model could not infer. She also asks what may be missing.
Be 100 times more specific.
The resulting rubric combines model-detected patterns with Hilary's explicit choices. She decides which criteria matter, how they should be weighted, and which kinds of work need different standards.
Turn the rubric into custom GPT instructions
Hilary asks ChatGPT to draft the instruction set for a custom GPT, then edits that instruction set rather than starting from a blank page.
My job is to create a GPT that can evaluate slide decks for people on my team based on my specific criteria that I care about. It's really important to me that this GPT explains the criteria, why it matters, and gives the user specific feedback on how to improve. YOUR job is to write the prompt for it.
The generated instructions define a direct coaching voice and criteria for vagueness, clutter, and narrative structure. Persona language can make feedback consistent, but it should not encourage disrespectful, punitive, or falsely authoritative evaluations.

Calibrate the GPT with the team
Hilary calls the GPT Deck Doctor and tests it with sample slides. It assigns ratings and suggests revisions based on the rubric.

The team can use the GPT for an early pass before asking Hilary for review. Calibration should include representative decks, disagreements between the GPT and qualified reviewers, accessibility checks, and a way for people to ignore or challenge unhelpful advice. Its scores should not become employee-performance measures.
Use AI to pressure-test a written argument
The business-writing sparring workflow gives the model a narrower role than ghostwriter: identify the thesis, question the stakes, surface weak assumptions, and help the author see where the argument needs more work.
Hilary uses ChatGPT as a reader and critic while developing a point of view. The goal is a clearer argument, not a formula for virality or a substitute for the writer's judgment.
Test whether the thesis is legible
Hilary begins with an unfiltered draft, then asks ChatGPT to identify the thesis and supporting points. Sensitive plans, personnel matters, customer details, and unpublished company information should be removed or handled only in an approved environment.
Here is a written first draft of a newsletter I'm writing. Can you succinctly express my thesis back to me and my main supporting points?
If the summary matches her intended argument, the structure is carrying the idea. If it does not, she revises the thinking rather than merely asking for smoother prose.
Strengthen the stakes and support
Once the thesis is clear, she asks where the framing, evidence, or momentum can improve:
How can I make this more compelling?
The model acts as a sparring partner by suggesting stronger framing and pointing to sections that lose momentum. Hilary decides which suggestions serve the argument and which would flatten her voice.
Find weak assumptions and missing perspectives
Hilary asks for counterarguments, blind spots, and gaps in logic instead of asking the model to validate her position.
What blind spots might I have as I'm talking about this?
The critique may identify an overlooked audience, structural constraint, or missing qualification. It can also invent objections or reflect model bias, so the writer should verify factual claims and seek relevant human perspectives before publication.
"Try to get the AI to help you make it right as opposed to assuming that it's right and getting the AI to validate that."
Restructure, then rewrite in your own voice
Hilary may ask for alternate structures after testing the argument. She rewrites the final piece herself and checks every sourced claim, quotation, example, and implication.
Reusable feedback creates more room for management
A custom GPT can make a manager's recurring criteria available earlier and more consistently. It cannot read interpersonal context, notice when the rubric is wrong for the situation, deliver accountable performance feedback, or replace mentorship.
A sensible first use is one low-stakes artifact with a clear rubric and a small calibration set. Keep the manager available, invite disagreement, protect the source material, and revise the tool when the team finds advice that is wrong or unhelpful.
Hilary is trying to move feedback earlier, not remove herself from management. A teammate can ask the custom GPT for a first pass before a slide review, fix obvious issues, and arrive at the human conversation with sharper questions. Hilary still owns the context the model cannot see: the audience, the person's growth goals, the politics around the decision, and whether the rubric fits this particular assignment.
The calibration process is the product. Hilary and the team compare the GPT's advice with her real edits, identify where it is too generic or too rigid, and revise the instructions. The tool becomes useful because disagreement improves it; a manager should never present its output as an invisible authority.
Her writing workflow uses the same boundary. She asks AI whether the claim is legible and well supported, then rewrites in her own voice. The model creates distance from a draft she is too close to, while authorship and accountability stay with her.
This is especially useful for a manager whose feedback has become a bottleneck. The GPT can make recurring standards available when Hilary is in another meeting, but it also gives her evidence about which standards were never clear to the team. If everyone challenges the same recommendation, the next management task may be improving the rubric rather than retraining the people. The disagreement is useful management data, not a failure to suppress.
The best signal is not that the GPT agrees with Hilary. It is that teammates can predict which feedback is routine, which advice they should challenge, and when a draft needs the manager's judgment rather than another automated pass.
Watch or listen
Sponsors
Thanks for supporting How I AI
Build your next product with ChatPRD
Turn an idea into a PRD, user stories, and a plan.


