Sebastian Raschka 7/16/2026

Inkling: A New Open-Weight 975B MoE with a Few Surprises

Read Original

This article covers the surprise release of Inkling, a 975B-parameter sparse Mixture-of-Experts (MoE) open-weight large language model from Thinking Machines Lab. It compares Inkling's performance against GLM-5.2 on benchmarks like IFBench and SimpleQA Verified, noting strengths and weaknesses in reasoning and coding tasks. The architecture is detailed: 41B active parameters, 1M token context window, regular Transformer decoder with GQA, and unique features like small convolution layers for local token mixing, an additional RMSNorm after embedding, and a learned input-dependent relative-position bias. The article also discusses implications for fine-tuning via the Tinker platform and speculates on token throughput versus competitors like Kimi K2.5 and Nemotron 3 Ultra.

Inkling: A New Open-Weight 975B MoE with a Few Surprises

Comments

No comments yet

Be the first to share your thoughts!

Browser Extension

Get instant access to AllDevBlogs from your browser

Top of the Week

No top articles yet