Thinking Machines Lab Releases Inkling, a 975 Billion Parameter Open Weights AI Model Under Apache 2.0 | Free Download

Thinking Machines Lab on Wednesday released Inkling, an AI model with 975 billion parameters. The company, founded in early 2025 by former OpenAI CTO Mira Murati, is providing the model under the Apache 2.0 license, which allows developers to freely fine-tune and use it.

Inkling is currently the largest open-weight model in the United States and is positioned as an alternative to Chinese open-weight models such as DeepSeek V4, GLM 5.2, and KMK2.6. It is now available on Thinking Machines’ Tinker platform and for direct download via Hugging Face.

Model Specifications, Hardware Requirements, and Positioning Against Chinese Open Weights

Inkling combines a blend of expert architecture with the following features:

  • 975 billion total parameters
  • 256 routed experts, and two shared experts
  • Six experts are activated for each token, resulting in approximately 41 billion parameters during inference. It supports a token reference window of one million
  • The model was trained on 45 trillion tokens consisting of text, images, audio and video and was developed using the Nvidia GB300 NVL72 system.

Inspired by DeepSeek-v3, Inkling was trained from scratch by Thinking Machines, rather than being fine-tuned with existing weights.

Running Inkling at native 16-bit precision requires more than two terabytes of GPU memory. A practical hardware setup consists of approximately eight Nvidia B300 accelerators or sixteen Nvidia H200 accelerators.

For users with less GPU power, Thinking Machines offers an NVFP4 scaled version of the model that requires about half the GPU capacity. This version trades some precision for lower memory usage while maintaining most of the capabilities of the model.

Reasoning Capabilities, Apache 2.0 Licensing, and Developer Access

Like most modern Frontier models, Inkling is a reasoning model that is trained with reinforcement learning to use chain-of-thought thinking before responding.

Thinking Machines claims Inkling matches Nvidia’s Nemotron 3 Ultra, which was previously the largest US open weight model at 550 billion parameters.

It reportedly achieves comparable results on Terminal Bench 2.1 using approximately one third the number of thinking tokens. This efficiency means that Inkling can deliver the same performance at a lower cost when the price is lower on a token basis.

Users should note that THINKING tokens are billed like any other token, so longer chains of THINKING may increase the cost of each reaction. Inkling is entering a market where Chinese labs have dominated open-source AI development. Competitors include:

  • DeepSeek V4 and V4-Pro from Huawei Ascend,
  • GLM 5.2, and
  • Like K2.6.

Thinking Machines claims the Inkling is competitive with these Chinese models across a variety of workloads. However, users are advised to independently verify such claims through their own testing, as gaming industry benchmarks for AI are often unreliable.

Inkling’s benchmark charts also show it lagging behind proprietary models like Anthropic’s Cloud and OpenAI’s GPT. The main tradeoff for users is gaining access to weight and full customization versus relying on a closed model with high raw performance.

Fine-tuning, self-modification and access points for Inkling are also part of the broader discussion.

Thinking Machines describes Inkling as highly adaptable to developers building AI applications and general-purpose applications such as chatbots. The Apache 2.0 license allows commercial use, modification, and redistribution.

The company also claims that Inkling can write its own fine-tuning scripts to refine its behavior, teach itself new skills, and evaluate its abilities. The self-modification capability is intended to make optimization more accessible to developers without deep machine learning expertise.

The company’s Tkinter platform provides tools for customization and fine-tuning. For developers wishing to use Inkling:

  • Tkinter Platform: Now Available for API Access and Fine-Tuning Tools
  • Hugging Face: Model Weight Available for Direct Download
  • Third-party API providers: Thinking Machines is working with TogetherAI, Fireworks, Model, Databricks, and Besten to add models to their services.

Inkling supports a wide range of inference engines at launch, including vLLM, SGLang, Miles, TokenSpeed, and Llama.cpp.

What users should do, hints—mini preview and availability

For developers interested in Inkling for their projects: Review the model’s benchmark performance in relation to your specific use case rather than relying solely on general claims. Consider whether a 1 million token context window justifies the hardware requirements for your workload.

If full precision hardware is not possible, evaluate a quantized NVFP4 version. Test models through the Tkinter platform before committing to a self-hosted deployment. It is also useful to compare token efficiency claims with real-world workloads in your domain.

For users exploring Inkling for hobby or evaluation purposes: Download the model via Hugging Face. Install an inference engine like Llama.cpp for local testing. Even with quantization, be prepared for significant hardware requirements.

Along with Inkling, Thinking Machines is previewing Inkling-Small, a 12 billion-parameter mix of experts (MOE) model with 276 billion active parameters.

This smaller model targets users who prioritize low latency over throughput and quality. It has not been released yet, but the company plans to release the weight once testing is over. No specific timeline has been provided for release.

Inkling’s launch comes at a time when the frontier AI market is increasingly dominated by proprietary models with limited access. OpenAI recently postponed the public release of GPT-5.6 at the request of the US government.

Anthropic’s Fable 5 and Mythos 5 were temporarily suspended before being restored. Alibaba also restricted access to cloud code for employees over tracking concerns.

In this context, Thinking Machines’ decision to release Inkling under the Apache 2.0 license provides an option for developers who want access to frontier-scale capabilities without relying on a single provider. The name references the fictional supercomputer creator from Jurassic Park.

Inkling is now available on the Tkinter platform and Hugging Face. Third-party API access is expected to follow through platforms such as TogetherAI, Fireworks, Modal, Databricks, and Basetain.

The company has not yet announced pricing for the Tkinter platform. Interested users can follow the company’s official channels for updates on Inkling-Small and future releases.

Thanks for being a Ghax reader. The post Thinking Machines Lab releases Inkling, a 975 billion parameter open weight AI model, under Apache 2.0 appeared first on Ghacks.

Source:Ghacks

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top