ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
NaiveAI Open-Sources 309B Model N0.5-Flash: Zero Full-Attention Layers, Native 1M Context

NaiveAI Open-Sources 309B Model N0.5-Flash: Zero Full-Attention Layers, Native 1M Context

AI information • Admin • • 9 views

On September 27, 2026, NaiveAI released the open-weight model Naive-N0.5-Flash: a 309B-parameter mixture-of-experts (MoE) model that activates only 15.5B parameters per token, supports a native 1-million-token context window, and is published on Hugging Face under the MIT license. It is the boldest engineering bet yet on the "no full attention" path — across all 48 layers of the network, not a single one uses conventional full attention.

48 Layers, Not One of Them Full Attention

The model splits its 48-layer Transformer into eight six-layer modules, each pairing five sliding-window attention layers with one lightweight DeepSeek Sparse Attention (DSA) layer: 39 sliding-window layers plus 9 sparse layers in total. The sliding window is just 128 tokens; each sparse layer selects the top 2,048 tokens for backbone attention, combined with grouped-query attention using 4 KV groups. The official model card puts it bluntly: the whole network stays local or sparse — no full-attention layers — yet still delivers a native 1M context window.

3.25T Tokens of Training, Built on Xiaomi's Open Model

Training ran on 3.25T tokens, all at the native 1M context length: a 50B-token warmup for the sparse indexer, 3T tokens of sparse-attention training, then a 200B-token learning-rate decay. The base is Xiaomi's open-weight model MiMo-V2.5 — a continuation of the same model family as MiMo-V2.6, which Xiaomi open-sourced earlier; the acknowledgments also credit the DeepSeek sparse-attention design team and the SGLang inference framework. The HySparse2 core architecture of MiMo-V3, which Xiaomi unveiled later, follows the same sparsification trend — "dropping full attention" is becoming an explicit direction among China's open models.

"Building AI with AI": The Research Pipeline Is Part of the Pitch

More noteworthy than the parameter count is how it was built. In its accompanying technical blog, NaiveAI writes that the research and engineering pipeline was "substantially executed by AI systems": models wrote code, ran experiments, monitored results, and proposed the next round of iterations, while human researchers set objectives, constraints, and evaluation standards and made the key decisions. The technical report is titled "Naive-N0.5-Flash: Building Frontier AI with AI." NaiveAI was founded in Beijing this February by Tsinghua professor Jifeng Dai — just seven months ago — making this model something of a first report card for its "AI builds AI" methodology.

Impressive Speed Numbers — All Self-Reported

Alongside the model comes the in-house inference stack NaiveRT: mega-kernel fusion, Programmatic Dependent Launch, speculative decoding, and FP8 mixed precision. The company claims 50 tokens/s per user in standard mode and up to 2,000 tokens/s in Ultrafast mode. Independent reviewers have already cautioned that these are narrow decode-speed figures, not end-to-end measurements, and that accuracy retention over long contexts remains unverified. Hosted API pricing is $0.10 per million input tokens, $0.40 per million output tokens, and $0.01 per million cache reads, aimed squarely at coding and AI R&D workloads.

Two reasons this deserves attention. First, it turns "can sparse attention fully replace full attention?" from a paper debate into downloadable engineering fact: if a 1M context window genuinely holds up with zero full-attention layers, it rewrites the cost equation for long-context models. Second, the MIT license means no barrier to commercial use, and the 309B-total / 15.5B-active ratio puts real deployment cost near that of a mid-size dense model — a new option for teams building coding agents and long-document RAG. The sober side: every impressive number is company-reported for now, and real performance awaits community testing — which is precisely the point of open weights. Skeptics can just run it and check.

Recommended Tools

More