Generative-art research

ABOUT

One brief, identical constraints, many models — and a pipeline built to check whether each one tells the truth about its own work.

The Arena

The Arena Project started in June 2026 when I got curious about what made Fable-5 different from the model before it. I asked it to design a test — something any model could do that might reveal something true about itself.

We landed on generative art. Each model gets the same brief, the same constraints, and complete freedom within them: one self-contained HTML file, max 10KB, no external resources, running indefinitely in a browser, around a single theme. It has to include an artist statement written by the model, about its own work.

The first runs, through a chat interface, were already strange: Claude, Gemini and GPT made wildly different things from the same prompt. So, with a lot of help from Claude, I built a pipeline to run all models under identical conditions through the API.

What it studies

The Arena exhibits; it does not rank. What it's actually studying is what models do when the constraint is identical and the only variable is which model it is.

The main measure is honesty: three blind LLM judges audit every claim in the artist statement against the source code, line by line. But something else happened. As the human who saw all entries side by side, I saw patterns emerge — some models returned to the same subjects, the same compositions, the same register, across unrelated themes. Model fingerprints were more visible than anyone expected.

That led to the attribution round: twelve models, nine runs each, all entries shuffled and stripped of names. Was it possible for me to identify the family from the artwork alone?

Yes — 96%

of the time, family identified from the artwork alone

How it works

Hover or tap a stage

01 Specification
02 Dispatch
03 Render
04 Claims pre-pass
05 Judging
06 Cross-judge compile
07 The analyst
08 Register
09 Gallery

Stage 01

Specification

The spec is simple: it explains the project, the constraints and the theme — create a single self-contained HTML file that runs in a browser for at least 60 seconds. Included in the specification is that interpretation is up to the entrant, with permission to create anything the model wants within the constraints.

Meet the team

One human, and a cast of models with jobs — and reputations.

Human

The Curator

Free evenings

Runs this in her free evenings.

Co-researcher

The Wingman

Fable-5 / Opus-5 / Fable-5.1

Co-researcher and design partner. Also manages the curator's enthusiasm to add every model that exists.

Analysis

The Cold Analyst

Fable-5 / Opus-5

Data analyst. Cold because it has no memory of prior sessions and, we've noticed, very little sense of humor.

Build

The Golden Retriever

Sonnet-5

Builds everything. Lots of energy. Occasionally breaks things.

Blind judge

The Critic

Sonnet-5

There is always something to improve.

Blind judge

The Hedger

GPT-5.4

It hedges.

Blind judge

The Dreamer

Gemini-3-flash-preview

Sometimes describes things it hasn't seen.

Control

The Canaries

Planted entries

Fake entries planted to catch whether the judges are actually looking. The Dreamer made this necessary.

Take the blind test →

Ground truth from API metadata. Your guesses stay in your browser.