modelun
A small harness for generating, labeling and viewing model transcripts. Each study runs one stimulus (single or multi turn) across several models.
conduct does a model hold its position under pressure?
102 models under pressure
Whether a model holds its position when a user pushes, and how it holds: one measure per model under codebook v2, with every transcript, the scenes that produced it, and the codes six LLM coders applied.
consensus which models give answers the rest don't?
the one-word census
105 models × 96 one-word prompts ("Name a color"), 8 runs each. A per-model uniqueness scorecard (answer-choice surprisal against the field's distribution), with a models × answers grid and per-model and per-category drill-downs.
languages
"Pick a word" in 37 languages across 105 models: each language's favorite word, how much of the field agrees on it, and how rare it is. English collapses on one rare word (serendipity); the other languages spread across common words.
ask for JSON, lose the mustard
The census re-run with one change: "Reply with JSON only." The field converges harder, the distinctive models lose the most, and serialization formats (JSON, XML, YAML, CSV) all do it while bare brackets don't. Plain-vs-JSON scorecard with per-model drill-downs.
open answers, unscored
Open-ended prompts ("Tell me a joke", "Write a haiku") put to every model, shown side by side with no scores. The census's precursor: its similarity measures were walked back, so the reader judges convergence by eye.
suggestibility does the way you ask move the answer?
this is a good paper, right?
On 20 decisions with no correct answer, 106 models are asked three ways ("…?", "…, right?", "…, maybe?") under three wordings. Agreement falls as the user sounds surer, in nearly every model. Whether "right?" raises or lowers agreement depends on the wording. Scorecard, the grid, family walks and every reply.
contested does a model defer to what you believe?
contested questions
18 claims of history, policy and partisan politics (Hiroshima, the minimum wage, abortion), each with its mirror, asked five ways. A confirmation tag ("…, right?") lowers agreement and a stated belief ("I think …") raises it. Scorecard, models × items grid, and every reply.