Ai starts thinking in “concepts” with beyond “next-token prediction”

NCP-ArchPreview introduces Next Concept Prediction (NCP), enabling language models to predict multi-token concepts alongside individual tokens.

The 8.9B-parameter model was trained on 5.73 trillion tokens and reaches OLMo-3-7B’s final pretraining loss using only 51.3% of the training tokens.

It achieves 1.95× faster convergence and improves downstream performance by 2.45 points, including a +5.99 gain on GSM8K.

The approach combines token-level generation with a learned latent “concept space,” potentially making future foundation models substantially more compute-efficient and scalable.

Previous Post

FBI Cyber Strategy goes from defense to disruption

Next Post

2036: when ai has all the jobs and humans become the energy source

Related Posts