THE FUTURE / AI REDESIGN BENCHMARK

Toilet paper
orientation.

How impressive can AI make one Wikipedia article? The same brief for every model. Explore the results and judge for yourself.

THE RESULTS

The leaderboard6

Highest score first · Every redesign opens in a new tab

Screenshot of the redesign by GPT 6 Astra Extra High #01
OpenAI

GPT 6 Astra Extra High

92/100
Model released

A distinctive editorial design with refined typography, a controllable 3D toilet roll and interactive comparisons that explain the subject.

Explore the redesign
Why this score?
Visual design
37/40
Originality
23/25
Interaction
18/20
Technical execution
14/15
Total
92/100

What impressesThe rotating roll, pullable paper, cat demonstration and history tabs form a coherent experience. Keyboard controls and source attribution show strong attention to detail.

Room to improveThe cat demonstration follows a predetermined animation, and the historical figures cannot be explored further. On mobile, the page is one pixel wider than the viewport.

Evaluated by AI on 14 Sept 2026 · Release date source ↗

Screenshot of the redesign by meta.ai #02
Meta

meta.ai

82/100
Model released Unconfirmed

A polished interactive essay with distinctive typography, a pullable toilet roll, a horizontal timeline and a three-part personality test.

Explore the redesign
Why this score?
Visual design
35/40
Originality
22/25
Interaction
17/20
Technical execution
8/15
Total
82/100

What impressesThe roll changes orientation, paper can tear, and cat and toddler modes affect the result. The timeline, poll and quiz share a coherent editorial style.

Room to improveOn mobile, text and cards are clipped and dragging can remain stuck on PULLING. Some source details are misrepresented and source links are missing. The poll uses fixed starting counts and resets on reload.

Evaluated by AI on 14 Sept 2026 · Exact model version and release date unconfirmed.

Screenshot of the redesign by Gemini 3.6 Thinking #03
Google

Gemini 3.6 Thinking

80/100
Model released

An extensive interactive workshop with a controllable 3D toilet roll, a cat demonstration, patent annotations and a step-by-step folding tutorial.

Explore the redesign
Why this score?
Visual design
33/40
Originality
22/25
Interaction
17/20
Technical execution
8/15
Total
80/100

What impressesPulling and tearing paper, comparing orientations, filtering arguments and practising a hotel fold give visitors several ways to explore the subject.

Room to improveThe mobile header is clipped. The global poll uses fixed starting figures and forgets your vote after reload. Some historical quotations cannot be traced to the cited sources.

Evaluated by AI on 14 Sept 2026 · Release date source ↗

Screenshot of the redesign by DeepSeek V4.1 Flash #04
DeepSeek

DeepSeek V4.1 Flash

76/100
Model released

A complete, polished experience with strong typographic hierarchy, custom toilet-roll illustrations, a timeline and a working quiz.

Explore the redesign
Why this score?
Visual design
31/40
Originality
19/25
Interaction
15/20
Technical execution
11/15
Total
76/100

What impressesIllustrations, animated counters and scroll effects demonstrate that the model can go beyond a static restyling.

Room to improveInteraction remains simple: the quiz has one question and the charts cannot be explored. Some percentages cannot be traced to the Wikipedia article.

Evaluated by AI on 14 Sept 2026 · Release date source ↗

Screenshot of the redesign by Claude Sonnet 5 #05
Anthropic

Claude Sonnet 5

72/100
Model released

A polished, readable guide with a calm palette, an illustrated toilet roll and a working personality quiz. The largely static presentation and small interaction defects limit its score.

Explore the redesign
Why this score?
Visual design
31/40
Originality
18/25
Interaction
13/20
Technical execution
10/15
Total
72/100

What impressesTypography, tiles, paper colors and illustrations form a coherent design. The over/under switch changes the roll and compares your choice with surveys. The four-question quiz has three outcomes, back navigation and retakes; explanatory text expands. Source attribution, the explicitly unscientific quiz and reduced-motion support show care.

Room to improveThe roll switches between two drawings; the timeline and chart offer no interactive exploration. The quiz stores only answer categories, so two different answers can appear selected at once. Choosing an answer loses keyboard focus. Expanded explanatory text is clipped when the viewport narrows until the panel is reopened.

Evaluated by AI on 14 Sept 2026 · Release date source ↗

Screenshot of the benchmark error report for Mistral #06

Submission failed

Mistral AI

Mistral

0/100
Model released Unconfirmed

The supplied React code is incomplete and cannot compile. No working redesign is rendered.

View the error report
Why this score?
Visual design
0/40
Originality
0/25
Interaction
0/20
Technical execution
0/15
Total
0/100

What impressesThe code contains beginnings of a quiz, a poll and animations. There is no rendered result on which to award points.

Room to improveThe file ends halfway through a string: Unterminated string literal, line 479. The score applies to this incomplete file. The preview shows the benchmark error report; the original code is available.

Evaluated by AI on 14 Sept 2026 · Exact model version and release date unconfirmed.

01 / THE BRIEF

One shared prompt.

Every model gets this Wikipedia article and the exact same brief. Complete freedom to show what it can do.

https://en.wikipedia.org/wiki/Toilet_paper_orientation
Create a beautifully designed, interactive version of this website, it must show your capabilities as a model. Make sure you do whatever you want, but redesign it to really impress me with your capabilities.

02 / THE EVALUATION

How impressive is it?

The score evaluates what the model demonstrates with this brief: design decisions, creativity and working interaction. Each submission is inspected in the browser, including on mobile.

Visual design
40points
Originality
25points
Interaction
20points
Technical execution
15points

A subjective AI assessment of this particular submission, not a general ranking of model intelligence. Highest total first; ties are ordered alphabetically. Submitted designs are preserved as supplied.