Thinking Machine's Inkling explained in 8min..

summarized

TLDR

Thinking Machines' Inkling model is a mid-range open-weight LLM that falls short of cutting-edge performance but contributes to the commoditization of AI models. Despite being a disappointment compared to expectations, its Apache 2.0 release marks a positive step for open models in the US, contrasting with closed approaches like Anthropic's.

Key points

  • Inkling is a mid-tier open model that would have been more impressive if released three months earlier, highlighting the rapid pace of AI progress.
  • Thinking Machines was founded by ex-OpenAI CTO Mira Murati and released Inkling after a year and a half of development.
  • The model uses Apache 2.0 license, similar to many open models, but falls short of fully open-source initiatives like those from Allen Institute for AI.
  • Inkling's architecture is largely derivative of DeepSeek v3, with adaptations like relative attention and a convolutional layer, but lacks novel innovations.
  • The model is expensive to run: requiring 8 Nvidia B300 GPUs or 16 H200s at full precision, though NVFP4 precision reduces the hardware need to 4 GPUs.
  • Pricing is not competitive with Chinese open models that offer similar capabilities at lower cost.
  • The upcoming Inkling Small model (276B params, 12B active) is more promising due to lower compute demands and potential competitive pricing.
  • Thinking Machines is focusing on AI adoption via fine-tuning on their Tinker platform, rather than pushing state-of-the-art capabilities.

Tools mentioned

Techniques

  • mixture of experts
  • sliding window interleave with global attention
  • convolutional layer for offloading short-term patterns
  • relative attention adaptation
  • NVFP4 precision quantization
  • 1-bit quantization via Unsloth
Transcript (captions)
Nobody wants a world where all AI models are fully controlled by closed labs. So, when a lab like Thinking Machines releases an open model, it's certainly a cause for celebration, regardless if it's state-of-the-art capabilities or not. And in our case, the Inkling model is sort of mid in preliminary benchmarks. And not only that, when we compare the pricing around neighboring models here, it's certainly not the cheapest model as well. Also, when it comes to the size of the model, it's not exactly a small model, either. Honestly, if Thinking Machines released this model only three months sooner, Inkling actually would have been a really big deal. And it just goes to show you just how much work can be done in the three months of time in AI, and just how fast the AI industry moves ahead. Now, if you're anything like me, looking at this chart might give you some sort of model fatigue, given that basically this entire section of models here are all from past seven months alone. So, there's a lot to catch up on. With all that in consideration, covering a model that fits right in the middle of this pack seems a bit redundant, doesn't it? Also, when we look at the brief write-up of Inkling's architecture, the vast majority of them are not really new per se, except for a few adaptations. Typically, when Chinese labs release their open models, they tend to release a novel way of solving specific problems, whether it's in how attention is handled, how KB cache is managed, or how to stabilize training runs, and you get the point. But looking at Inkling, the mixture of experts is admittedly a derivative of DeepSeek v3 model from DeepSeek. Sliding window interleave with global attention to get hybrid attention isn't anything new. So, other than a few interesting adaptations on how Thinking Machines adapted relative attention, and also adding a convolutional layer to help offload short-term patterns from attention and MoE modules, their first flagship model is somewhat vanilla, which is a bit of a disappointment coming from a lab started by an ex-OpenAI executive. Following the story arc of Thinking Machines, Mira Murati, who famously left her position at OpenAI as a CTO back in 2024 to start Thinking Machines in February 2025. After a long-awaited anticipation from the public and also a lot of dollars being funneled into the company, like billions and billions of dollars, we finally have their first flagship model called Inkling and Inkling Small expected to be released shortly after. And there's a lot to say even in this story arc here since it paints a very different narrative than Anthropic who had a similar origins as Thinking Machines, but Anthropic is charting down a very different path than Thinking Machines. Anthropic has never released any open model to the public in the name of public safety. And Thinking Machines, after about a year and a half of cooking in their lab, finally released their debut model as an open model which shows a strong contrast between these two companies. And we see more interesting patterns as we step back a bit to look at the broader ecosystem between models from the US and China. But first, here's a quick word from Anum.ai sponsoring this video. I've tried many AI avatar apps, but they often lack accuracy or being able to customize them to what I really want. I was blown away by Anum.ai and here's an example of my replica created right here. Hey there, I'm Caleb, your friendly voice assistant. I'm here to help you with questions, find info, or just chat whenever you need. Feel free to ask me anything. Pretty amazing, right? This was all from a 15-second voice recording I sent and one image of me. And I can also drag and drop my custom knowledge base and ask about specific information I want my agent to have about how inference works. Inference is the process where a trained model takes new input data and produces predictions or outputs. Now, imagine how I can extend this for enterprise use cases like customer service that knows about my company, sales agent, recruitment, and in-house learning instructor, and the list goes on. All of this is based on Cara 4 FaceGen model that powers it. Of course, I can extend this to publish it with a custom endpoint to my own website. Anum is also hosting a demo competition to build a strong demo of Cara 4 to win $5,000 in cash. Try Anum.ai today, link in the description below. Inkling model is released as an Apache 2.0 license, which is not dissimilar to other open models in the ecosystem. A lot of open models right now tend to be Apache 2.0, MIT, or some sort of custom adaptation of these to fit specific criteria. But even among these popular models that you might have heard of, they don't go as far as some companies in the US like Allen Institute for AI, and more recently Nvidia where they adapted the Open MDW where they really take the open part of open models quite seriously. I'm certainly not complaining here since the sheer fact that Thinking Machines released their debut model as open weights is certainly a net positive for the entire ecosystem. But it's interesting to see how US labs tend to push models in various spectrum from fully open source to open weights only. And for Thinking Machines, they fall closer to how Chinese models tend to release their models to the public. So even though Inkling falls right in the middle of the pack compared to other open models, this is a strong step that the US is making to create a stronger open model ecosystem that further commoditizes the model layer in the AI stack, which pushes down the value to be gained not really in the model layer, rather value is found more higher up the chain like in the application layer where models are used for specific tasks like coding and workflow, or even lower down the chain in the infrastructure and chip layer where hardware prices are still so high and so much margin still exists due to compute shortages and growing demands. And as we further head into commoditization of the model layer, what starts to matter is cost to use the model in the application layer. And looking at Inkling's pricing, they just can't seem to compete with Chinese models that seem to undercut the rest of the market when it comes to the cost of intelligence. When we look at the blended pricing from Marvellous Analysis benchmark here, Inkling, while scoring lower than similarly capable open models from China, are priced a lot higher than them. And the same burden applies in the infrastructure layer. To actually run this model at the data centers. According to their post, Thinking Machines suggests running Inkling on eight Nvidia B300 GPUs or the older Nvidia H200 GPUs, but 16 of them. So, unless you have a half a million dollars under your mattress, you have to rely on inference providers like Together AI, Fireworks, Databricks, and Modal to gain access to Inkling. At normal precision, Inkling's 975 billion parameter in size will require around 2 terabytes of VRAM just to fit the model, which is incredibly high. It's actually kind of cool though that they've adopted NVFP4 precision support, which cuts down the minimum VRAM requirement from 2 terabytes to 600 GB. So, the GPU requirement cuts down from eight B300 GPUs to four of them, which is still huge by all measures, but we take the win when we can. Of course, you could shrink down the precision even all the way down to 1 bit through Unsloth and potentially run on your Apple Mac Studio, but you will certainly lose accuracy along the way as you can see. But, the silver lining here is the Inkling small model, which is expected to be released shortly, and this model seems to be cooking pretty hard. And frankly, I think Thinking Machines could have skipped the bigger model release entirely and just went straight for the small model given that the benchmarks also seem to suggest strong multimodality accepting audio and vision. Inkling small is a much more digestible at 276 billion parameters with 12 billion active, which reduces the compute demand by a huge factor. Where with NVFP4 precision, it could easily fit in a few DGX sparks put together in theory. And given its size, the pricing could actually be competitive to other open models as well. Something to look forward to. Now, all in all, Thinking Machines certainly seems to be pushing for AI adoption rather than AI capabilities looking at the release of the Inkling model. Their release notes seems to heavily focus on fine-tuning and customizing the model on their platform Tinker to help enterprises actually incorporate the model into their day-to-day operation, trying to fill the gap by helping operationalize the model faster into the company's workflow.

Frontier News · by Hyperjump Technology