GPT-6 Sol and Luna cost less, but hands-on tests show mixed results
In a video published on September 24, 2026, Robbert van Empel compares OpenAI’s GPT-6 Sol and GPT-6 Luna with their earlier versions through two practical creative tasks. OpenAI says the new models’ API prices are 50 percent lower than the previous promotional prices. The video tests whether those savings produce visible practical gains.
The first challenge is a redesign of the Wikipedia page about toilet paper orientation. Both models produce a cleaner layout than the original page. Luna adds an animation that shows the roll going over or under. Sol’s version has a little more interaction and a timeline, but the video finds that neither redesign includes the richer animation or visual detail seen in some other model outputs. The results show that both can create a presentable small website, while giving a user less to explore than the brief might suggest.
The second challenge asks each model to build a South Park Lego world with working physics, stackable bricks and humor. In the video, Luna’s world is less consistently made of Lego, its characters are harder to recognize and the navigation is awkward. Sol’s version looks more like a Lego world and its characters are more recognizable. In both examples, blocks can be picked up, but neither model lets the reviewer stack them reliably. Both generate jokes in the style of the show and use recognizable characters. That observation describes the outputs; it is not a legal assessment of their use of copyrighted material.
The video also compares score and cost-per-task figures shown by Artificial Analysis. At the time of the comparison, Sol is only one point above GPT-5.6 Sol on the intelligence score cited, while Luna and its predecessor both score 37. The task-cost figures are lower for the new versions in the reviewer’s examples. These are snapshots from a third-party site and costs for specific tasks, not a complete or controlled comparison of every use case.
The practical result is mixed: the lower API prices are clear in OpenAI’s pricing table, but the two demonstrations show modest gains in interaction and visual coherence rather than a decisive leap. Teams considering Sol or Luna should test their own prompts, track retries and token use, and judge the output against their actual requirements. Two creative challenges can reveal strengths and weaknesses, but they cannot establish which model is best overall.