Inkling
Thinking Machines Lab · United States · 2026
The first model from Mira Murati's Thinking Machines Lab — a 975-billion-parameter multimodal mixture of experts released under Apache 2.0 with open weights.
Inkling, published on 15 July 2026, is the first model released by Thinking Machines Lab, the company founded by former OpenAI chief technology officer Mira Murati. It is a 66-layer decoder-only transformer with a sparse mixture-of-experts feed-forward backbone: 975 billion parameters in total, of which roughly 41 billion are active for any one request, combined with hybrid local and global attention. The context window reaches one million tokens. The model is natively multimodal on the input side. It accepts text, images between 40 and 4096 pixels and audio in 16 kHz WAV up to twenty minutes long, and it answers in text only. Thinking Machines reports frontier-class results on reasoning and coding — 97.1% on AIME 2026 and 77.6% on SWE-bench Verified — placing it near the closed leaders rather than at the top of them. A smaller companion model, Inkling-Small, carries about 12 billion active parameters for cheaper and faster inference. What makes the release unusual is the licence. The weights are open under Apache 2.0 and distributed through Hugging Face, so they can be downloaded, inspected and fine-tuned without restriction, and the company does not monetise the model through metered API access; it is also served via the in-house Tinker API and third-party inference providers. The company's stated thesis is that no single model fits every organisation and that customers should be able to adapt one to their own data. The model card is candid about limits, noting that the model can still comply with harmful requests wrapped in roleplay and advising builders to add their own content filtering rather than trusting built-in refusals.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!