AI Models

Thinking Machines Releases 975B Parameter Inkling Model

Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, has released Inkling, an open-weights multimodal model with 975 billion parameters. The model leads U.S. open-source models on the Artificial Analysis Intelligence Index but trails top Chinese models. It shows strong agentic performance but has a 63 percent hallucination rate.

Neura News

Neura News

Neura Market Editorial

July 16, 20264 min read
Thinking Machines Releases 975B Parameter Inkling Model

Thinking Machines Lab Ships First Production Model

Thinking Machines Lab, the AI startup founded by former OpenAI CTO Mira Murati, has released its first production-ready language model. The model, named Inkling, is a Mixture-of-Experts Transformer with 975 billion total parameters. Of those, 41 billion parameters are active at any given time.

Murati previously served as OpenAI's chief technology officer and played a key role in developing ChatGPT. Her new venture has now entered the competitive open-weights model space with Inkling.

Multimodal Capabilities and Fine-Tuning Focus

Unlike many open-source AI models, Inkling natively handles text, images, and audio. It supports a context window of up to one million tokens. The model weights are freely available on Hugging Face. Thinking Machines also offers access through Tinker, its platform for adapting AI models to specific tasks.

The company is positioning Inkling as a flexible base model for customization. In its announcement, Thinking Machines stated that Inkling is not the strongest overall model available today. The company expects the combination of multimodal support, efficient processing, and fine-tuning options to set the model apart.

Training Data and Synthetic Data Use

Thinking Machines pre-trained Inkling on 45 trillion tokens of public and synthetic data. That data includes text, images, audio recordings, and videos. The training set also includes public data that may be subject to intellectual property protection.

The company used the Chinese AI model Kimi K2.5, among other methods, to generate synthetic data. Kimi K2.5 also served as the basis for Cursor's coding model. More technical details are available in the model card published by Thinking Machines.

Benchmark Performance and Rankings

According to AI benchmarking platform Artificial Analysis, Inkling debuts with a score of 41 on the Artificial Analysis Intelligence Index. That makes it the leading open-weights model from a U.S. lab. It ranks three points above the previous U.S. leader, Nemotron 3 Ultra at 38. It also sits well ahead of Gemma 4 31B at 29 and gpt-oss-120b at 24.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

On GDPval-AA v2, an agent-based benchmark that simulates knowledge-work tasks, Inkling reaches an Elo rating of 1,238. It beats Kimi K2.6 at 1,190 and DeepSeek v4 Flash max at 1,189. Inkling also scores 24 percent on the Tau-3 banking benchmark, ahead of Kimi K2.6 at 21 percent and DeepSeek v4 Flash max at 23 percent.

Factual Accuracy and Hallucination Issues

Inkling performs rather poorly on factual accuracy. Artificial Analysis gives the model a score of just +2 on its AA Omniscience benchmark. That puts it below the leading open-weights models, though still above other U.S. models such as Nemotron 3 Ultra at -1. Inkling's accuracy is 40 percent, while its hallucination rate is 63 percent. Those results are likely to limit its use in applications that need highly accurate information.

Pricing and Token Efficiency

With a 64K context window, Inkling costs $1.87 per million input tokens and $4.68 per million output tokens. That is slightly more than open-source Chinese models such as GLM-5.2 and DeepSeek v4, which offer similar or better performance on text and code tasks. For context windows up to 256,000 tokens, pricing rises to $3.74 for input, $0.748 for cached input, and $9.36 for output.

However, Inkling uses fewer output tokens than comparable open-weights models. According to Artificial Analysis, it averages 25,000 output tokens per Intelligence Index task. GLM-5.2 max uses 43,000, Kimi K2.6 uses about 38,000, and DeepSeek v4 Pro max uses about 37,000 tokens on the same tasks.

Thinking Machines says Inkling offers continuously adjustable thinking effort. Users can choose their preferred balance between cost and performance, reducing token use while maintaining the same result quality.

Inkling-Small Preview

Thinking Machines is also previewing Inkling-Small, a more compact model with 276 billion total parameters and 12 billion active parameters. The smaller model delivers similar or better results than Inkling on several benchmarks.

Inkling-Small scores 88.3 percent on GPQA Diamond, compared with 87.2 percent for Inkling. On the HLE benchmark with tools, it scores 46.6 percent, slightly ahead of Inkling at 46.0 percent. Thinking Machines credits changes to the pre-training data and training process for the results. The company plans to publish the full weights once testing is complete.

Related on Neura Market

More from Neura News

Industry

Monday.com Joins Tech Layoff Trend Citing AI as Factor

Monday.com announced it will lay off about 20% of its workforce, or over 600 employees, citing a restructuring tied to its AI-driven growth strategy. The Tel Aviv-based work management software company joins a growing list of major tech firms, including Amazon, Meta, and Microsoft, that have cited artificial intelligence as a factor in job cuts this year. A new Financial Times analysis shows U.S. tech companies have slashed nearly 140,000 jobs since January, with AI often cited as a reason.

Jul 26·12 min read
General

Open-weight AI mirrors Kubernetes ecosystem shift

Tobi Knaup, co-founder of Mesosphere, draws parallels between the rise of Kubernetes and the current trajectory of open-weight AI models. He argues that open-weight models are becoming a neutral substrate for innovation, attracting a global ecosystem of developers, startups, and enterprises. The piece warns against US restrictions on Chinese open-weight models, advocating instead for American leadership through open releases, procurement strategies, and standards.

Jul 25·7 min read
General

Open-weight AI mirrors Kubernetes rise, US warned on bans

The author, a Mesosphere co-founder, draws parallels between the rise of Kubernetes and the current open-weight AI ecosystem. He argues that open-weight models are becoming a neutral platform for innovation, and warns that US restrictions on Chinese open-weight models could isolate American developers from a global ecosystem. The piece urges the US to compete by releasing frontier models, using procurement to create demand, building the stack, and setting standards rather than imposing bans.

Jul 25·7 min read
Industry

Power line failure reveals AI data center grid risks and solutions

A fallen power line near Washington, DC caused over 3 gigawatts of data center load to vanish from the PJM grid in seconds, spiking voltage across the region. The event, which made lights flicker from Northern Virginia to Chicago, highlights a growing problem as AI data centers become larger and more concentrated. Experts warn that without better coordination or technology like ON.Energy's battery-backed uninterruptible power supply, such disruptions will become more frequent and severe.

Jul 25·5 min read