Shape
03

Play to learn.

Playgrounds are small, focused tools that isolate one design lever at a time. Each one produces an artifact you can download.

01

Diff mode

Run one prompt through two configurations side-by-side. The fastest way to feel how prompts shape outputs.

Produces — Diff Log · Concept — Prompts as designOpen →
02

Tone dial

Treat style as a design token. Move dials for warmth, verbosity, energy, directness — see the prompt compose itself.

Produces — Behavior Spec · Concept — Voice & toneOpen →
03

Persona lab

Design a character — backstory, beliefs, blind spots — and watch the model embody them.

Produces — Persona Card · Concept — Personas for AIOpen →
04

Refusal lab

Probe boundary design with a panel of edge cases. Tune the line between over- and under-refusal.

Produces — Refusal Scorecard · Concept — Refusal & boundariesOpen →
05

Eval lab

Rubric-based evaluation. Define what good looks like, score the model against it, watch the average move.

Produces — Eval Rubric + Scorecard · Concept — EvaluationOpen →
06

Conversation choreographer

Write the user's side of a conversation in advance, then run it. Same script, different system prompt — see how the model holds the thread.

Produces — Behavior Spec · Concept — Multi-turn flowsOpen →
07

Spread lab

Run one config many times. Outputs are a distribution, not a value — find out which clauses of your spec actually hold.

Produces — Stability Report · Concept — Distributions, not outputsOpen →
08

Race lab

One prompt, two models, at once. Watch what the better answer actually costs — in seconds and in dollars.

Produces — Speed TrialOpen →
09

Portability lab

One spec, several models. Find out which clauses are real rules and which are incantations tuned to a single vendor.

Produces — Portability Report · Concept — Distributions, not outputsOpen →
10

Context lab

One question, several context sets. Your system prompt is a fraction of what the model reads — see the rest, and where the answer came from.

Produces — Context Map · Concept — Context is the interfaceOpen →
11

Tool bench

Now it does things, not just says things. Write the tools and the policy, then find out where it draws the line between asking and acting.

Produces — Agency Policy · Concept — Designing agencyOpen →
12

Judge lab

Hand the scoring to a model, then check it. Every comparison runs twice with the answers swapped — a judge reading position gives itself away.

Produces — Calibrated Judge · Concept — Judging at scaleOpen →
13

Round table

Three seats, one decision, a planted dissenter. Edit the protocol — who speaks first, whether the first round is blind, when it stops — and watch the room decide, not the prompts.

Produces — Protocol · Concept — Groups, not agentsOpen →