Play to learn.
Playgrounds are small, focused tools that isolate one design lever at a time. Each one produces an artifact you can download.
Diff mode
Run one prompt through two configurations side-by-side. The fastest way to feel how prompts shape outputs.
Tone dial
Treat style as a design token. Move dials for warmth, verbosity, energy, directness — see the prompt compose itself.
Persona lab
Design a character — backstory, beliefs, blind spots — and watch the model embody them.
Refusal lab
Probe boundary design with a panel of edge cases. Tune the line between over- and under-refusal.
Eval lab
Rubric-based evaluation. Define what good looks like, score the model against it, watch the average move.
Conversation choreographer
Write the user's side of a conversation in advance, then run it. Same script, different system prompt — see how the model holds the thread.
Spread lab
Run one config many times. Outputs are a distribution, not a value — find out which clauses of your spec actually hold.
Race lab
One prompt, two models, at once. Watch what the better answer actually costs — in seconds and in dollars.
Portability lab
One spec, several models. Find out which clauses are real rules and which are incantations tuned to a single vendor.
Context lab
One question, several context sets. Your system prompt is a fraction of what the model reads — see the rest, and where the answer came from.
Tool bench
Now it does things, not just says things. Write the tools and the policy, then find out where it draws the line between asking and acting.
Judge lab
Hand the scoring to a model, then check it. Every comparison runs twice with the answers swapped — a judge reading position gives itself away.
Round table
Three seats, one decision, a planted dissenter. Edit the protocol — who speaks first, whether the first round is blind, when it stops — and watch the room decide, not the prompts.