Legora reviews 41 financial documents in minutes with GPT-6 Astra

Legora reviews 41 financial documents in minutes with GPT-6 Astra
News

OpenAI has published a customer story about Legora, a legal-technology company using GPT-6 Astra for a financial-statement tie-out. The task compares figures in draft accounts with trial balances, a consolidation schedule and the previous year’s accounts. Legora says this work can take an evening or several days when done manually. In one run, its Agent reviewed 41 documents within minutes, checked balances against supporting schedules, surfaced breaks and recorded each check for a professional to review.

The result is presented as an example of AI handling connected financial context rather than answering a single question. OpenAI says Legora measured the workflow with its Benchmark for Agentic Reasoning, or BAR, which uses end-to-end tasks based on real-world legal work. On this financial-statement workflow, Legora reports that GPT-6 Astra improved performance by nearly 40 percent against the previous model. Across the full BAR task set, the reported average improvement was about 3 percent. These figures come from Legora’s evaluation, not an independent audit.

Legora also planted four errors in the accounts for the test. OpenAI reports that Astra found all four, including a £500,000 gap hidden in a revenue note. The Agent retained the checks that the previous model had completed correctly and added around 50 more checks. That combination matters because a review process needs both exception-finding and a traceable record of what was examined. The figures do not establish a general error rate for financial work.

A key part of the workflow is the division of responsibility. The Agent performs the exhaustive comparison, while the legal or financial professional makes the judgment call on each result. OpenAI presents that human review as central to Legora’s plans to extend its platform into audit, tax, compliance and risk. The case illustrates a bounded use of an agent: accelerate a repetitive first pass, preserve evidence of the checks and leave consequential interpretation with an expert.

For businesses, the practical lesson is not simply that more documents can be processed faster. Teams also need a representative test set, clear definitions of completion, access controls for sensitive accounts and review points before a conclusion is accepted. For AI users and makers, the Legora example shows where stronger context handling may be useful: workflows that require comparisons across many files and a clear handoff from automated checking to accountable human judgment.

Source openai.com