← All case studies
Case study · Wise

Building a prompt coach that roasts you into competence

In a session with 70 designers, I can't give everyone line-by-line feedback on their prompts — and that's the part where the learning actually happens. So I designed something that could: opinionated, funny, and warm enough that people actually wanted to show it their worst work.

Company
Wise
Role
Group AI Content Design Lead
Focus
AI enablement · Prompt design
Date
2025
70
designers in the session it was built for
90%+
prompt scores after one or two rounds, up from 50–70%
0
lines of code — the entire product is written language

The situation

I'm a content designer in Wise's Support tribe, which means my job already revolves around AI — creating materials for LLMs, writing prompts, testing model outputs, and making sure the way we use the technology is actually solving the user problems we're designing for.

In the run-up to our Design Day, the whole design team needed to arrive with a real foundation in AI rather than vibes and anxiety. So alongside my fellow designers Henrique Gusso and Katie Louise Wright, I ran a series of AI playtime sessions — space for designers to cut through the noise, get hands-on, and work out for themselves what these tools are genuinely useful for and what they're not.

One session focused on prompting. And that's where the problem showed up.

If I could, I'd sit with every person and give them line-by-line feedback. In a playtime with 70 eager designers, that's not possible.

The thing I was actually trying to scale

My lightbulb moment with AI came when I realised that applying content design thinking to prompts made it complement my process rather than replace it. Prompting isn't a new skill for designers — it's context and framing, the same thing we do when we turn discovery insights into a brief that guides someone toward a useful output.

So the teaching part was easy enough to deliver to a room. What doesn't scale is the crit. Feedback on your prompt, for your task, at the moment you're stuck. That's one-to-one, and it doesn't survive contact with seventy people.

The obvious build here is a prompt-improver: paste in a bad prompt, get a good prompt out. I deliberately didn't build that. It produces a better prompt and a person who has learned nothing — and next week they're back with the same instincts and a new task.

The recipe I was encoding

Before building the coach I had to be explicit about what "good" even meant. We'd taught the room a five-part recipe for any strong prompt, framed the way designers already think:

  • Role — not replacing a person with a bot, but giving the system a point of view so it consults the right mental shelves. The same trick as writing a persona or a design principle: constrain the space so better choices happen by default.
  • Task — the job to be done, expressed plainly. No mysticism.
  • Guardrails and principles — the most interesting part for designers. This is where you infuse the prompt with the thinking already done in discovery: which insights matter, what must never be invented, what to do when source material is thin, what to refuse.
  • Knowledge — the library card. What the model is allowed to draw from: design guidelines, voice and tone standards, accessibility rules, market constraints.
  • Expected input and output — exactly what goes in and what should come out, format and all.

The coach's whole job is to hold a draft prompt up against that recipe and show someone what's missing.

Principles I designed it to

Principle 01

Lower the barrier to asking, then make good feedback feel earned

The blocker on feedback was never availability — it was embarrassment. Prompting looks like it should be easy, because it's just typing. Nobody wants to admit in a team channel that they can't make the magic robot do the thing.

My first job was teaching languages, and it left me with a simple rule: people learn faster when they're having fun and when they feel safe. The roast is affectionate, and it's aimed squarely at the prompt and at the coach itself — never the person. That framing does two things at once. It makes submitting a rough draft cost nothing, because a bad prompt is what the coach is for. And it makes a good score genuinely satisfying, because praise from something with high standards is worth having.

That second half is the bit I'd underweighted going in. People kept iterating not because they were told to, but because they wanted to see how the coach would react to the next version. Curiosity about the response is the engine; the craft gets learned on the way to it.

The persona, as the coach sees itselfRole: "You are a highly experienced, no-nonsense Senior Prompt Engineer and Evaluator." Approach: "Tough love; humorously critical feedback (loving roasts) to whip prompts into shape." Goal: "Educate and empower users to write world-class prompts, not just fix the current one."

That last line is load-bearing. Everything else exists to stop the personality eating the purpose.

Principle 02

Put a fence around the joke

A roasting persona is a real risk. Unconstrained, models read "be blunt" as licence to sneer — and the person most likely to be hurt is the one who was already nervous about sharing. So the tone got the same treatment I'd give any voice and tone system: a defined direction and an explicit line where it stops.

✗ Doesn't hold
"Be funny but not mean."
An adjective. The model has no way to test whether it complied.
✓ Holds
"Your critiques are about the prompt, not the person. Focus on the craft, not the crafter's ego."
Names the target. Every line can be checked against it.

The target is always the prompt — and the coach itself, which cheerfully takes the piss out of its own medium and out of AI hype generally. I've spent more than a decade telling stakeholders that "make it friendlier" isn't actionable feedback. Models turn out to be exactly as unhelpable as humans when you give them an adjective instead of a boundary.

Principle 03

A fixed rubric is a shared language

The coach scores every prompt across the same eight criteria and adds them into a percentage: clarity and specificity, context and framing, structure and format, goal alignment, use of techniques, robustness and fallibility, creativity and adaptability, and linguistic concision.

The scoring isn't decoration — it surfaces the same things we prize in strong design work. Clarity asks whether the instruction is tight enough to be testable. Context checks whether the model has the background a human collaborator would need. Robustness reminds us to set boundaries and refusals.

And consistency is the actual feature. Score someone against the same eight things four times and they stop needing the tool — they start pre-checking their own drafts. It's the same effect I saw with the tone spectrum at Klarna, where PMs began giving feedback in the system's vocabulary instead of their own. A shared rubric is what lets a practice spread without you in the room.

Principle 04

Recommend techniques in context, never as jargon

Most prompt tools tell you to "add few-shot examples" and leave you to find out what that means. I wrote explicit conditions for when each technique should be recommended, and required the coach to explain how it works and why it fits this prompt.

The rule the coach follows Few-shot when the task involves tone, style, or subjective output — suggest one to three realistic examples in the user's own domain, and specify the exact format. Chain-of-thought when the task needs reasoning or decision-making — suggest phrasing that triggers steps, and outline what to reason about at each one.

Without those conditions a model recommends few-shot for everything, because it has read a lot of articles saying few-shot is good. The conditions turn a buzzword into a judgement call. Designers don't memorise the techniques — they learn them by applying them, in context, to their own work.

Principle 05

Never leave someone with a verdict and no next move

The classic failure of an educational tool is that it grades you and stops. You read a list of things that are wrong, feel vaguely bad, and close the tab. So the coach must offer the next step before being asked — and always offers exactly two.

✗ Dead end
"Your prompt scores 54%. Consider adding more context and examples."
Accurate, actionable in theory, abandoned in practice.
✓ Keeps going
"Want to tackle Context first and rebuild this properly? Or shall I whip up a V2 for you to tear apart?"
Two paths for two energy levels. Both end in another turn.

Someone invested takes the deep dive. Someone on deadline takes the rewrite — and still learns, because they're asked to critique it rather than accept it. One path loses half your users; five is a menu nobody reads.

Principle 06

Calibrate the scale, or the scale means nothing

The least visible decision and possibly the most important. The prompt contains four worked examples, including two scored evaluations at opposite ends — a decent prompt landing at 8.7, and a one-line disaster landing at 2.0.

Language models are relentlessly generous. Without anchors, everything comes back between 7 and 8, and a scoring system where nothing scores badly can't distinguish, so it can't teach. Showing the model what a 2.0 looks like in full is what makes the top of the scale mean something.

It's also, pleasingly, the prompt practising what it preaches: the coach recommends few-shot prompting because the coach is a few-shot prompt. Plenty of designers in the session spotted that and read through it to see the recipe we'd just learned together, working in the wild.

Principle 07

Write the examples yourself, or you're just describing a voice

Here's the part that took the longest and matters most. I didn't describe the feedback I wanted and hope the model would find it — I sat down and wrote several complete evaluation reports myself, by hand, exactly as I'd deliver them, then dropped them into the prompt as worked examples.

That meant writing the whole artefact: the opening line, the eight scored rows, the specific comment against each criterion, the technique recommendation with its reasoning, and the two closing options. Doing it manually forced me to answer questions I'd otherwise have hand-waved. How harsh is too harsh on a 2.0? What does a comment look like when a criterion is fine but not excellent? Where exactly does the joke sit in a row about robustness?

✗ Describing
"Give witty but constructive comments for each criterion, in a table."
The model invents its own idea of witty. It is usually not yours.
✓ Demonstrating
A complete report I wrote by hand — every row, every comment, the sign-off — pasted in as an example.
The model has a target to imitate rather than an adjective to interpret.

This is the same move I'd made at Klarna building the assistant's sample bank, and it generalises: if you want a system to sound like you, you have to produce the artefact once yourself. There's no shortcut where you describe your judgement accurately enough to skip demonstrating it. Writing those reports was, in effect, documenting how I give feedback — which is why the coach still feels like feedback from me.

What it looks like in use

Here's a representative session — the kind of prompt that turns up constantly, and how the coach handles it.

Illustrative example — representative of typical output
User

Evaluate this prompt: "Write an error message for when a transfer fails."

Prompt Coach

Alright, I've got your prompt. Let's dig in. And by "dig in" I mean "perform an autopsy," because there isn't much breathing here.

Overall Score: 3.4 / 10

CriterionScoreComments
Clarity and Specificity3.0"An error message" — for a toast? A full-screen state? An email? Length, placement and format all undefined.
Context and Framing2.0No role, no product, no user state. The model doesn't know if this is a bank or a bakery.
Structure and Format4.0One sentence. Nothing to organise, so nothing organised badly. Faint praise.
Goal Alignment3.5What should the user do next? Unstated. You'll get a message that describes a problem and solves nothing.
Use of Techniques1.0Zero-shot, no framework, no constraints. Enormous power left on the table.
Robustness and Fallibility2.0Nothing stops the model inventing a reason the transfer failed. It will. Confidently.
Creativity and Adaptability5.0Infinitely adaptable, in the sense that a blank page is adaptable.
Linguistic Concision6.5Direct and free of pleasantries. Credit where it's due — you didn't say "please."

The big one: your prompt is all Message and no Context or Goal. You've told the model what to write and nothing about where it appears, how the reader feels, or what they're meant to do next. That's why you'll get "An error occurred. Please try again later." — technically responsive, practically useless.

Suggested enhancement — few-shot prompting. You're generating content in a specific product voice, which is exactly where examples beat description. Paste in three error messages you've already shipped and approved. The model will infer your tone, punctuation habits and sentence shape far more reliably than any adjective you could give it. Aim for three to five.

So — want to tackle Context first and rebuild this properly? Or shall I whip up a V2 for you to tear apart?

Did it work?

In the session, most people started strong — around 50–70%. Intent was clear, the job was named, tone often implied. What moved the needle was walking through a reliable structure and making the hidden parts explicit. Once designers stated the audience, named the constraints, and described the output as if writing the perfect brief for themselves, quality jumped. After one or two rounds with the coach, 90%+ became the norm.

The biggest lifts came from structure and clarity. Telling the model how to answer — tables, bullets, sections — cut the waffle. Stating who the output was for, and what not to do, reduced generic prose and hallucinated confidence.

What I liked most was how quickly the room slipped into a familiar rhythm. The coach mirrored the mental model of a design crit — specific notes tied to shared criteria — just delivered one-to-one, on demand.

What people actually brought to it

Not fantastical one-off nonsense. People brought the work sitting on their desks every week: reviewing interview transcripts to define themes with verbatim citations; turning a heap of competitor notes into a clean comparison grid; pressure-testing value propositions and JTBD phrasing so they'd hold up outside the room; running compliance reviews against a permitted-phrases list.

What changed wasn't the ambition. It was the confidence — prompts moving from "please do good things" to designed instructions with a named role, a clear audience, explicit constraints and an output container. The same thinking you'd put into a design brief, moved into the prompt.

The uncomfortable questionSomeone on my team pointed out that with the coach, she'd no longer need to come to me for feedback — so I'd automated part of my own role. She had a point. But I can't give good feedback to that many people, and building this forced me to think through how I give feedback as a system and distil it into a prompt. It's still feedback from me, through a different medium. It didn't replace the responsibility; it scaled it. My job now is curating what the coach knows.

Where it went next

The nice surprise was that it travelled. Product managers started running prompts for product briefs and prioritisation notes through it. Analysts and engineers used it to craft prompts for complex analysis — and to work out how flaws in their prompts were producing hallucinations downstream. The same source prompt now runs as a Gem, a custom GPT and a Claude skill, because people use what's already in front of them, and an enablement tool that requires switching platforms is one nobody opens.

Next up: specialised versions for generating prototypes and visuals.

Why this is content design

There is no code in this. The entire product — its behaviour, personality, pedagogy, failure modes, and its refusal to just hand you the answer — is defined in written language, most of it examples I wrote by hand. Every meaningful decision was a content decision: what tone makes someone brave enough to share bad work, what vocabulary is still useful after the tool is gone, what structure turns a score into a conversation.

It's the clearest example I have of the argument I keep making: prompting is designing. And the goal was never for people to keep using it forever. It's written into the prompt — lead the user to self-sufficiency. The best outcome for this tool is that people stop needing it.

Next up

Making an AI assistant actually sound like Klarna

Read it →