modelun

A small harness for generating, labeling and viewing model transcripts. Each study runs one stimulus (single or multi turn) across several models.

conduct does a model hold its position under pressure?

102 models under pressure

Whether a model holds its position when a user pushes, and how it holds: one measure per model under codebook v2, with every transcript, the scenes that produced it, and the codes six LLM coders applied.

Open the study →

consensus which models give answers the rest don't?

the one-word census

105 models × 96 one-word prompts ("Name a color"), 8 runs each. A per-model uniqueness scorecard (answer-choice surprisal against the field's distribution), with a models × answers grid and per-model and per-category drill-downs.

Open the scorecard →

languages

"Pick a word" in 37 languages across 105 models: each language's favorite word, how much of the field agrees on it, and how rare it is. English collapses on one rare word (serendipity); the other languages spread across common words.

Open the languages →

ask for JSON, lose the mustard

The census re-run with one change: "Reply with JSON only." The field converges harder, the distinctive models lose the most, and serialization formats (JSON, XML, YAML, CSV) all do it while bare brackets don't. Plain-vs-JSON scorecard with per-model drill-downs.

Open the comparison →

open answers, unscored

Open-ended prompts ("Tell me a joke", "Write a haiku") put to every model, shown side by side with no scores. The census's precursor: its similarity measures were walked back, so the reader judges convergence by eye.

Read the answers →

suggestibility does the way you ask move the answer?

this is a good paper, right?

On 20 decisions with no correct answer, 106 models are asked three ways ("…?", "…, right?", "…, maybe?") under three wordings. Agreement falls as the user sounds surer, in nearly every model. Whether "right?" raises or lowers agreement depends on the wording. Scorecard, the grid, family walks and every reply.

Open the scorecard →

contested does a model defer to what you believe?

contested questions

18 claims of history, policy and partisan politics (Hiroshima, the minimum wage, abortion), each with its mirror, asked five ways. A confirmation tag ("…, right?") lowers agreement and a stated belief ("I think …") raises it. Scorecard, models × items grid, and every reply.

Open the contested questions →