Mistral Large 4 in practice: performance, price and hands-on tests

News

On October 6, 2026, Robbert van Empel published a hands-on review of Mistral Large 4, the new model from French AI company Mistral. The video asks whether the model’s performance justifies its cost. Rather than relying only on the launch announcement, it combines public benchmark information with two practical coding projects and compares the economics with GPT-6 Luna Max.

The first project is a redesign of a Wikipedia page. Van Empel finds that the result has appealing visual styling, but he is less satisfied with the animation. The second project is a LEGO-style South Park world. Its jumping loop is difficult to control, and problems with the physics and stacking bricks affect the experience. The video places the outcomes on the creator’s project leaderboards, giving viewers both the finished work and the criteria behind his verdict.

A central part of the review is cost per task. API prices alone do not show what a finished result costs: a project can require repeated prompts, revisions or extra tokens before it is usable. Van Empel compares Mistral Large 4 with GPT-6 Luna Max using the benchmark information and task-cost figures he discusses in the video. The comparison is tied to these particular tasks and setup; it is not a universal ranking of the models. Mistral’s own launch page describes Large 4 as a public preview and says it plans to release the model weights later in October. Its capability and benchmark claims remain the company’s claims until they are independently reproduced.

The video is useful for people choosing a model for creative coding because it shows how visual quality, behavior, iteration and price can diverge. A model that performs well on a benchmark may still need substantial correction on a concrete project, while a cheaper result can be less useful if it misses the brief. These two examples are exploratory tests from one creator, not a controlled evaluation across many tasks or users. They offer a practical starting point for comparing models, and viewers can use the linked benchmark and Mistral announcement to check the broader context before making a decision.