Basis evaluated two large language models on a complex accounting task and reported notable differences in speed and behavior. In tests comparing GPT‑6 Astra and GPT‑5.6 Sol, Basis found that GPT‑6 Astra completed a 50-tab tax workbook in about half the time required by GPT‑5.6 Sol.
Faster completion and early decision-making
According to Mitch Troyanovsky, Co-founder of Basis, GPT‑6 Astra “does a better job of really understanding the intent of the user and the problem.” Basis said that improved initial decisions helped agents follow a more direct path through the workbook and reduced time spent correcting mistakes. The company also noted that this early accuracy made the model more efficient in token usage.
Adaptive reasoning and cost implications
Basis described a capability in GPT‑6 Astra to adjust the amount of reasoning applied as a task progresses, increasing computation for difficult steps and reducing it for simpler ones while keeping the model’s cache intact. Basis said this dynamic adjustment helped lower response times and cost for long-running tasks.
In internal evaluations, Basis reported roughly a 20% improvement in scores with GPT‑6 Astra. The company attributed this gain to better recognition of user intent, including clearer decisions about when to ask clarifying questions, flag assumptions, and follow instructions.
Basis evaluates agents by examining how they obey templates, consult primary sources for tax questions, and perform self-checks on their work. The firm said GPT‑6 Astra can infer these expectations from broader context with fewer explicit rules, reducing the need to write specialized rules for individual cases.
Mitch Troyanovsky emphasized that these findings reflect Basis’s internal tests and assessments of agent behavior and final outputs, including reliability on templates and source consultation.
Original source: OpenAI News