IA · 16 July 2026 · 3 min read

Against the AI Giants: Thinking Machines Launches Inkling, a 975-Billion-Parameter Open-Source Model

In brief: Thinking Machines Lab has released Inkling, its first 975-billion-parameter open-weight Mixture-of-Experts model. Founded by former OpenAI executives Mira Murati, John Schulman, and Lilian Weng, the startup aims to disrupt the 'one-size-fits-all' approach of tech giants, offering businesses a highly customizable, decentralized multimodal alternative.

by Team Mocchi's

Against the AI Giants: Thinking Machines Launches Inkling, a 975-Billion-Parameter Open-Source Model

The Defector-Led Rebellion Against Closed AI

The generative artificial intelligence landscape is undergoing a major realignment, driven by one of the most heavily funded and talked-about startups in the sector. Thinking Machines Lab, founded in early 2025 by a prominent group of executives and researchers who defected from OpenAI, has released its first native model. Named Inkling, the model represents the first public proof of concept from a company that, despite securing a record-breaking $12 billion initial valuation, has spent the last year and a half developing its technology in strict secrecy.

The initiative is led by prominent industry figures: Mira Murati, former Chief Technology Officer of OpenAI; John Schulman, co-founder and a primary developer of ChatGPT; and Lilian Weng, former VP of safety and robotics under Sam Altman. With the launch of Inkling, the startup is posing a direct challenge not only to their former employer but also to industry giants like Google and Anthropic.

Under the Hood of Inkling: 975B Parameters and Native Multimodality

From a technical standpoint, Inkling is a behemoth. It is built on a Mixture-of-Experts (MoE) architecture boasting 975 billion total parameters. However, to maintain compute efficiency and keep operational costs manageable, the system activates only a fraction of this power — roughly 41 billion parameters — for any given task.

As reported by TechCrunch, Inkling was trained from scratch on a massive dataset of 45 trillion tokens comprising text, images, audio, and video. This native multimodal architecture allows the model to deeply comprehend disparate data streams, though its current outputs are restricted to structured text, programming code, and formatted artifacts.

An intriguing phenomenon observed during training relates to its reasoning optimization. During the later training stages, the model's internal chain of thought underwent a natural condensation process: Inkling began shedding grammatical overhead and extra words to solve problems faster, without losing any logical coherence or affecting the final output's accuracy. Furthermore, the system introduces a "thinking effort" toggle, allowing developers to scale reasoning depth up or down depending on whether they prioritize raw speed or analytical depth.

The Bet on Open-Weight and Decentralized Intelligence

The most significant strategic aspect of Inkling lies in its licensing model. Unlike the closed, proprietary APIs distributed by OpenAI or Anthropic, Thinking Machines has opted for an open-weight release. This means that researchers, startups, and large enterprises alike can download the model weights and host them on their own private infrastructure for full customization.

According to analysis by Wired, this decision reflects a deeply decentralized vision of AI, standing in stark contrast to the centralized, "one-size-fits-all" paradigm pushed by the market's current dominant players. Thinking Machines aims to prove that frontier AI should not be controlled by a narrow oligopoly but should instead be highly adaptable to the unique operational requirements of individual organizations.

In terms of performance, while the startup openly admits that Inkling is not yet the strongest overall model across every commercial benchmark, early evaluations indicate remarkable efficiency. For instance, to achieve the same accuracy in code generation, Inkling consumes only a third of the tokens compared to Nvidia's recent open-weight alternative, Nemotron 3 Ultra. Additionally, the model provides highly calibrated responses, explicitly flagging logical uncertainty rather than generating false information or hallucinating.

Mocchi's take

At Mocchi's, we believe the release of Inkling marks a major turning point for businesses seeking to adopt artificial intelligence while maintaining strict data governance. The arrival of a 975-billion-parameter open-weight model gives companies a powerful tool to deploy custom software agents and reasoning engines on proprietary servers or GDPR-compliant European cloud systems. This flexibility allows organizations to bypass the cost volatility and architectural limitations of closed APIs, enabling deeper customization of business logic. For Italian enterprises looking to build sustainable and secure AI solutions, mastering and deploying these customizable open-source giants is the clearest path to true technological autonomy.

Further reading

All articles on the Mocchi's blog