Summary
NCP-ArchPreview (arXiv 2609.10715, Fri 11 Sep digest) is a latent-space language model that pushes autoregressive pretraining beyond next-token prediction: alongside NTP it learns Next Concept Prediction (NCP) — predicting discrete multi-token concepts via a product-quantized concept vocabulary built from its own hidden states, with a dedicated Concept Module whose predicted concepts feed back to guide token-level generation (NTP and NCP trained jointly end-to-end). It scales to 8.9B parameters on 5.73T tokens of Dolma-3.
Why it matters
A serious small-scale probe of the beyond-next-token question: if concept-level prediction improves with scale, it points at a training objective reasoning over larger units — with obvious implications for long-context efficiency and for speculative decoding. At 8.9B it is a research preview; watch whether the concept vocabulary transfers at scale.
Technical details
| Arxiv | 2609.10715, Fri 11 Sep 2026 digest (complete 91/91 at 18:46Z retrieval; arrived early — cross-lists may accrete) |
|---|---|
| Architecture | NTP + NCP jointly; product-quantized concept vocabulary from hidden states; dedicated Concept Module |
| Scale | 8.9B params; 5.73T tokens of Dolma-3 |
Tags
architecturelatent-spacenext-concept-predictionpretraining-objectives