Ai2 published details of AstaBrief 8B, an 8-billion-parameter model trained to turn a research question and retrieved literature excerpts into a cited report. The team says the model is available in Asta’s Generate a report feature as Fast mode and has been open-sourced alongside training data and an example local workflow for producing reports from PDFs.
How AstaBrief was trained
The developers started from Qwen3-8B and focused on post-training data, evaluation, and the surrounding report-generation pipeline. Training emphasized supervised fine-tuning (SFT) and direct preference optimization (DPO) rather than reinforcement learning, with the authors noting RL can be unstable and costly.
Training drew on real research queries filtered from Asta logs. After removing non-scientific, bot, and privacy-sensitive traffic, the pool contained about 90K research-focused queries. The SFT dataset included roughly 47K synthetic full-report targets produced by a multi-step ScholarQA pipeline that used a mix of backing models, including Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1.
DPO pairs were constructed from a separate set of queries. One report in each pair came from the ScholarQA pipeline and the competing report from models such as o3, o4-mini, DeepSeek-V3, or DeepSeek-R1. Two judge models—GPT-4.1 and DeepSeek-R1—had to agree on the preferred report; after filtering this produced about 6K DPO examples, which the team reports had high judge alignment.
Grounding, filtering and evaluation
Ai2 reports that early SFT runs improved overall content but lagged on answer precision and citation quality, which led to focused data-quality work. Four statistical filters were tested: output-to-input token ratio, citation relevance, citation density, and citation diversity. The strongest gains came from removing synthetic reports with low citation density, the authors say.
Development evaluation used SQABench-CS2 (200 computer-science questions) as the primary test, with secondary checks on DeepScholarBench (63 queries) and pairwise comparisons against the Claude-powered pipeline and DR Tulu. The team reports AstaBrief was competitive with the Claude-backed pipeline and DR Tulu on several answer and citation measures. In a small human study, DR Tulu ranked highest on overall preference while two of three researchers preferred AstaBrief on citation accuracy.
Ai2 emphasizes that the results reflect the models and pipeline state at the time of development and frames the open weights and example workflow as tools institutions can use to run report generation locally. The post was published October 2, 2026 and notes AstaBrief is available today in Asta’s Fast mode alongside Claude-powered Thinking mode.
Original source: Hugging Face Blog