AI News: GPT-6 Astra, Claude, Gemini and a very busy week

News

Robbert van Empel has published a 17-minute video about an unusually busy week in AI, bringing together new models from OpenAI, Anthropic, Google and Meta. The video opens with GPT-6 Astra, Claude Fable 5.1 and Mythos 5.1, Gemini 3.8 Flash and Flash Cyber, and Meta’s Muse Spark 1.3. The focus is not only on launch claims: the comparison also asks what the models cost, how they appear in Artificial Analysis, and what happens when they are asked to produce something people can actually inspect.

One of the most practical segments is a redesign challenge. Several free models are asked to turn the Wikipedia page about toilet-paper orientation into a more usable, interactive experience. The experiment reveals different trade-offs between visual polish, interaction and whether generated code actually builds and works. It is not a definitive benchmark, but a quick way to see how model behaviour changes when the same creative task is given to different systems.

Another demonstration turns a prompt into a small South Park-inspired world. A free European model struggles to create the world in one shot, while other attempts look more promising but still show the difference between an attractive concept and a working game. These examples keep the comparison grounded in outputs that viewers can see, test and question rather than in a single leaderboard.

The second half is a series of shorter updates. ChatGPT Images 2.5 has been released with new ways to work from sketches and edit images. Three hikers rescued from California’s Mount Shasta after using Google Gemini while planning their expedition provide a more serious reminder: a chatbot can be useful for ideas, but safety-critical decisions still need human judgement and reliable local information. The video also looks at Tesla’s Cybercab, presented as a car without a steering wheel, and at a factory robot that combines a humanoid body with an industrial robot arm.

Taken together, this is a snapshot of a fast-moving week rather than a final ranking of AI models. Its value is in showing where differences become visible: in price, design, code, reliability and the boundary between a convincing demo and a useful result. The video is a practical tour through those examples, with a simple conclusion that remains relevant beyond this week’s releases: keep thinking for yourself when an AI system sounds confident.