Claude Opus 5.5: What Anthropic’s new model can actually do
Anthropic introduced Claude Opus 5.5 on September 22, 2026, with a focus on coding, computer use and longer-running knowledge work. In this six-minute video, Robbert van Empel looks beyond the launch announcement and tests the model on creative and practical tasks. The video also reviews Anthropic’s benchmark and pricing claims, then asks how those claims compare with results that viewers can inspect.
Anthropic says Opus 5.5 improves on its previous Opus model in several evaluations and can complete more work with fewer tokens. The company lists API prices of $4 per million input tokens and $20 per million output tokens, below the previous Opus rates. It also reports stronger results on selected coding and computer-use tests. These are vendor-reported figures, so they do not establish how the model will perform in every application. The video puts the release claims alongside examples rather than treating a benchmark table as a complete answer.
The practical tests include an interactive redesign of Wikipedia and a South Park-inspired Lego world. The video also looks at animations and a game made by other users. Robbert says the model does particularly well with animations, while his own redesign and Lego experiments leave room for improvement. These are individual trials, not controlled comparisons: the prompts, time spent iterating and amount of human editing all affect what appears on screen. The examples are useful for seeing the model’s strengths and limits, but they cannot predict every user’s results.
That distinction matters for anyone deciding whether to use Opus 5.5 for creative work or software tasks. A polished demo can show what is possible, while a benchmark can compare a narrow task under defined conditions; neither alone captures the effort needed to reach a usable result. The video gives viewers concrete outputs to judge and connects them to price and performance claims from the launch.
For makers and teams, the useful next step is to try the model on a representative task and track the full cost: retries, review time, token use and corrections. The video is an independent hands-on perspective from its creator, not an Anthropic evaluation or a universal ranking. Its examples help set expectations, while decisions about adoption still depend on each workflow’s quality requirements, budget and tolerance for manual review.