34 hands-on tests show GPT-6 Astra’s strengths and limitations
An early hands-on review of GPT-6 Astra documents 34 projects covering games, 3D worlds, websites, research, presentations, writing, code and browser use. Published around the September 3 launch, it combines playable builds, recordings and project notes. It is not a controlled comparison, however: the prompts, hardware and amount of human correction were not standardized across models.
The most striking results came from 3D work. Seven Little Worlds contains seven explorable miniature biomes with terrain, water, weather and animals. AFTERHOURS is a walkable city rendered as ASCII characters. The most ambitious build, Newhaven, ran for five days through Codex’s goal mode and implemented roads, utilities, traffic, population and politics. The city simulator remained unfinished, and its visual direction and controls still needed intervention.
The tests also covered work closer to everyday office tasks. Astra researched video topics, developed arguments with sources and turned them into presentations with speaker notes. The recorded browser tests show it drawing an Excalidraw workflow, comparing eBay listings without buying anything and planning a walking route through Kyoto. The review reports clear improvements over GPT-5.6 in browser control, presentations and research. It also describes the writing as more natural and responsive to direction, while noting that familiar AI writing patterns remained.
The limitations are as informative as the highlights. Early tasks often stopped after roughly 30 minutes unless the requested scope and longer-term goal were explicit. Slides repeated layouts, designs fell back on familiar colours, speaker notes required another pass and one rendering task was stopped after it slowed the test computer. Some game controls also needed fixing. Stronger output therefore did not remove the need for clear instructions, review and iteration.
OpenAI positions Astra as a model for computer use and complex professional work. These visible artifacts let users inspect some results directly, but they do not establish how Astra will perform on other hardware or workflows, nor do they show an overall failure rate. For makers and businesses, the practical lesson is to test representative assignments, define completion in advance and retain approval points around purchases, publishing and consequential changes. The model may carry projects further, while responsibility for judging the result stays with the user.