For Philosophy & ethics
Shape model behavior.
Every model carries values someone chose — in a prompt, a rubric, a rule. Test them against hard cases, rerun them, reframe them, and see which commitments actually hold.
You already think in hard cases.
Thought experiments, value conflicts, the gap between a stated principle and what it licenses — these are the working tools of model behavior, too.
/play/refusal→
You test a principle against hard cases. Now test a model's boundaries the same way.
/play/spread→
You know one answer isn't a position. Now see the distribution behind a model's answer.
/play/evals→
You ask who decides what counts as good. Now write the rubric, and watch it crown the wrong answer.
A suggested path.
- 01LessonRefusal & boundariesWhere the model says no is a design surface. Over- and under-refusal both fail users.
- 02PlaygroundRefusal labProbe boundary design with a panel of edge cases. Tune the line between over- and under-refusal.
- 03LessonDistributions, not outputsOne output is a sample, not a result. Design against the spread — and find out which clauses of your spec actually hold.
- 04PlaygroundSpread labRun one config many times. Outputs are a distribution, not a value — find out which clauses of your spec actually hold.
- 05LessonEvaluationRubrics + sample sets. Make “good” measurable — then find out what your rubric actually rewards.
- 06PlaygroundEval labRubric-based evaluation. Define what good looks like, score the model against it, watch the average move.
- 07ExperimentDoes the frame change the choice?Identical outcomes described as lives saved or lives lost — the classic framing effect, run on a model.