Claude Sonnet 5.5 impresses in a website test, struggles with a game

News

Robbert van Empel has tested Claude Sonnet 5.5 against its predecessor in two hands-on creative tasks. In a video published on September 29, 2026, he asks whether the new model’s practical results justify its cost. The comparison focuses on a Wikipedia page redesign and a small LEGO South Park game, rather than a broad set of coding or business benchmarks.

The first task asks both Sonnet versions to redesign a Wikipedia page. Van Empel says Sonnet 5.5 impressed him with an interactive website. The second task asks the models to build a LEGO world with South Park characters. Here, the result was less convincing: the video description says Sonnet 5.5 struggled with the game. Van Empel tries a follow-up prompt to improve that output, making the video a look at both an initial response and an iteration. These are examples from two prompts, not a controlled evaluation across many tasks.

The video also considers whether the model is cheaper in practice. Anthropic’s launch material says Sonnet 5.5 can cost up to 30 percent less per task than Sonnet 5, even though the listed token prices are unchanged. Van Empel looks at cost-per-task figures alongside Artificial Analysis intelligence scores. Those measures can help frame a comparison, but cost depends on the prompt, output length and number of retries; a small set of tasks cannot settle what a model will cost for every user. The video asks viewers to weigh those economics against the quality of the outputs.

The results point to an uneven upgrade: a strong interactive website can coexist with trouble completing a game prompt. That matters to people choosing a model for creative coding, where a polished first result, working interactions and repair work all affect value. For companies, the practical lesson is to compare complete tasks, including follow-up prompts and token use, rather than treating a provider’s average saving as a guaranteed reduction in every workflow. The test offers a concrete demonstration, not a universal ranking of Sonnet 5.5.