AstaBrief 8B is the scientific report generation model that Ai2 (the Allen Institute for AI) open-sourced on October 2, 2026: give it a research question and a set of retrieved literature excerpts, and it produces a report with citations. The model is already live as Fast mode in the "Generate a report" feature of Asta, Ai2's platform for scientific work, and its weights and training data are openly available; the model is hosted in the allenai organization's AstaBrief_8B repository on Hugging Face.
Why scientific reports need a model trained specifically for them
Scientific writing is demanding in a particular way: answers must stay close to the evidence, a study's conclusions must not be quietly broadened, and researchers need to check whether each citation actually supports the claim attached to it. Ai2 set out to test whether an 8B open model, built on Qwen3-8B and post-trained only for report generation, could approach the report quality of the proprietary pipelines it had been using — while cutting generation time and serving costs. By design, it writes the full report in one pass, skipping the excerpt-summarization and clustering stages of the older pipeline.
The key was not the algorithm — it was how the data was filtered
AstaBrief uses a conventional recipe of supervised fine-tuning plus direct preference optimization (SFT + DPO); the effort went into the data. The SFT examples came from real researchers' queries: roughly 90K candidate queries were filtered for quality, relevance, and privacy, then run through the ScholarQA pipeline to generate full reports as targets, leaving about 47K usable examples generated by a mix of Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. The DPO stage produced about 6K preference pairs, keeping only pairs where two judge models, GPT-4.1 and DeepSeek-R1, agreed — judges that matched human preferences 95% of the time. Of four statistical filters tested, the most effective was also the simplest: citation density, whether a synthetic report consistently attached citations to its claims; more elaborate filter combinations added nothing. In the full Asta pipeline, Fast mode averages 51.1 seconds per report against 178.5 seconds for Thinking mode — about 3.5 times faster — and on SQABench-CS2, DeepScholarBench, and a small human study, its answer and citation quality sat at the same level as the Claude-powered pipeline.
Where it can be used, and where the limits are
Open weights mean institutions can run AstaBrief on their own infrastructure — sensitive questions about unpublished work never have to leave the building. Ai2 also provides an adaptable example workflow so researchers can generate reports from their own PDFs. The limits are stated plainly: the training data and comparison baselines reflect the frontier models of 2025, and Ai2 says explicitly that it has not rerun the full evaluation against today's frontier models. And a citation can point to the right paper while the claim still overstates what that paper showed; that kind of subtle broadening of a finding's scope is still hard for current metrics to catch. For people building research tools, there is a broader lesson here: making a model professional does not necessarily mean pouring domain text into pretraining — the composition of the post-training data, and whether that data demonstrates citations that actually land, can materially change output quality.