Developer

LangChain and NVIDIA Launch NemoClaw Deep Agents Blueprint

LangChain and NVIDIA have released the NemoClaw for LangChain Deep Agents blueprint, designed to help enterprises build open, governed agent systems. The blueprint combines LangChain Deep Agents Code, NVIDIA Nemotron 3 Ultra, and NVIDIA OpenShell runtime, enabling teams to tune agents for their workloads, run them securely, and optimize for quality, cost, and speed. In evaluations, Nemotron 3 Ultra with a tuned LangChain Deep Agents harness achieved an aggregate score of 0.86 at a cost of $4.48, roughly 10 times lower inference cost than the next closest performing model.

Neura News

Neura News

Neura Market Editorial

July 25, 20267 min read
LangChain and NVIDIA Launch NemoClaw Deep Agents Blueprint

LangChain and NVIDIA Launch NemoClaw Deep Agents Blueprint

For production agents, choosing a model is just one piece of the puzzle when it comes to improving agent performance. Teams also need to manage the system around the model, including the tools an agent can use, the context it sees, how it is evaluated, where it runs, and what policies apply to each action.

Today, LangChain and NVIDIA announced the NemoClaw for LangChain Deep Agents blueprint. This blueprint was developed with NVIDIA to help enterprises build open, governed agent systems. It brings together LangChain Deep Agents Code, NVIDIA Nemotron 3 Ultra, and NVIDIA OpenShell runtime. This combination allows teams to tune agents for their own workloads, run them securely, and optimize for quality, cost, and speed.

In LangChain's evaluations, Nemotron 3 Ultra with a tuned LangChain Deep Agents harness provided advanced agent performance at a much lower inference cost. The key takeaway is that agent performance improves when the model, harness, evals, and runtime are tuned together.

An Open Agent Stack for Enterprise Workloads

As enterprises move agents into production, the systems they build around the model become valuable intellectual property. Agent memory, workflows, traces, eval datasets, harness configuration, and tuning data all reflect a company's unique domain expertise. This is proprietary knowledge that can shape how the company competes. However, in closed ecosystems, teams do not fully control it. They need a way to own that work, improve it over time, and run agents with the controls their organizations require.

The NemoClaw for Deep Agents blueprint gives teams control over the full agent stack. It includes an open model layer with Nemotron 3 Ultra, which teams can run, customize, and optimize for enterprise workloads. It also features a tuned agent harness. LangChain Deep Agents Code (dcode) provides the harness layer for long-running agents, including planning, tool use, memory, and task execution. The blueprint includes a Deep Agents harness profile tuned for Nemotron 3 Ultra. Finally, it includes a governed runtime. NVIDIA OpenShell provides a secure runtime for sandboxed agent execution, with policies for how agents interact with tools, systems, and data.

Together, these layers provide a path to build and deploy agents that teams can measure, govern, and improve in production.

Reaching Benchmark-Leading Performance at 10x Lower Cost

In LangChain's agent eval suite, NVIDIA Nemotron 3 Ultra evaluated with LangChain Deep Agents achieved an aggregate score of 0.86 at a cost of $4.48. The next closest performing model cost $43.48. This makes Nemotron 3 Ultra roughly 10 times lower inference cost on this benchmark.

To reach these results, LangChain tuned how the agent uses tools, manages context, and evaluates intermediate steps. The goal was to adapt the harness around Nemotron 3 Ultra's strengths and the common patterns that show up in long-running agent tasks.

For enterprise teams, this means the entire agent system can be optimized around the requirements of their specific workloads. Teams can fine tune model weights, customize the harness, run evals, and control the runtime based on quality, cost, latency, and governance requirements. This gives companies a way to own the agent's core intelligence and decide when, how, and why the system changes over time.

Why Lower Inference Cost Leads to Better Agents

Lower inference cost is one of the biggest benefits of using open models. It reduces serving cost for end users, and it also changes how teams build and improve agents.

Evals work best when teams run them throughout the agent development lifecycle. Before deployment, teams need to test changes to prompts, harnesses, tools, models, and data. After deployment, they need to monitor real behavior, make fixes needed, and create tests so the regression does not happen again.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

When each iteration is expensive, teams run fewer evals, compare fewer variants, and avoid specialized agents because the operating cost is too high.

A more cost-efficient open stack makes it practical to run larger eval suites before deployment and in production, compare more model, harness, and tool variants, and evaluate specialized agents for specific domains.

The best system for a given workload optimizes across quality, cost, speed, and governance together. For example, the latency requirements for real-time customer support agents look very different from a coding agent running concurrent tasks in the background.

"Super agents have arrived. With an open model like NVIDIA Nemotron, a LangChain harness, the NVIDIA OpenShell runtime, and a company's own data, every enterprise can build custom agents that understand its business, use its tools, and turn knowledge into action. The future of AI won't be one-size-fits-all. Companies will use AI cloud services and build their own AI, shaped by their proprietary data, know-how, and workflows, and run it safely and securely wherever they operate.", Jensen Huang, Founder and CEO of NVIDIA

Ecosystem Support

The announcement is supported by partners across the AI infrastructure and enterprise ecosystem. EY is building an implementation practice around the software stack. Other partners include Baseten, Fireworks, Nebius, Crusoe, DeepInfra, and Together AI. These partners help enterprises serve Nemotron models in production and adapt the blueprint for business critical applications.

"EY clients in regulated industries are ready to move agentic AI out of isolated pilots and into production and are often constrained by governance, security, and the ability to prove control to a regulator or a board. Open agent architectures matter because they give enterprises transparency into how agents operate, control over where data and inference run, and the freedom to deploy on their own terms without committing to a closed stack. By delivering the NVIDIA NemoClaw blueprint, which incorporates LangChain technology, EY teams help give clients a secure, sandboxed foundation for always-on agents that can meet enterprise standards for auditability and risk from the first deployment.", Geoff Vickrey, Global Chief Commercial Officer, NVIDIA, EY

"Production agents need inference that is fast, reliable, and cost-efficient at scale. We have optimized NVIDIA Nemotron models on Baseten to deliver high throughput and low latency on NVIDIA hardware, so teams get strong price-performance without operating the infrastructure themselves. Delivering Nemotron through the NemoClaw blueprint with LangChain gives enterprises a clear path to run open agentic models in production with the performance and economics these workloads demand.", Philip Kiely, Head of Developer Relations, Baseten

"Agentic workloads make many model calls per task, so inference speed and cost directly determine whether an agent is viable in production. Fireworks serves NVIDIA Nemotron models with the throughput and price-performance that high-volume agent systems require, tuned for the tool calling and reasoning patterns these workloads depend on. Offering Nemotron through the NemoClaw blueprint with LangChain gives enterprises an efficient, open foundation they can scale with confidence.", Lin Qiao, CEO and Cofounder, Fireworks AI

"The next challenge for enterprise AI is running complex agentic workloads economically at production scale. Nebius was built for that challenge. Our AI-native cloud gives customers dedicated infrastructure optimized for high-performance inference and cost-efficient scaling. By offering NVIDIA Nemotron models through the NemoClaw blueprint with LangChain, we're making it easier for organizations to deploy and scale open agentic AI across their business.", Roman Chernin, Chief Business Officer, Nebius

Availability

The NemoClaw for LangChain Deep Agents blueprint is available today.

For deeper technical detail, read: Deep Agents Code on NemoClaw: a governed blueprint for your most sensitive code, and Tuning the harness, not the model: a Nemotron 3 Ultra playbook.

Related on Neura Market

More from Neura News

Funding

Prentis AI Lab Co-Founded by Reid Hoffman, Marc Pincus Seeks $100M

Prentis, a new AI research lab co-founded by Ritankar Das, Reid Hoffman, and Marc Pincus, is in talks to raise $100 million at a $1 billion valuation. The startup focuses on computer use models that automate office workflows. It has already signed contracts worth up to $50 million with several customers and claims its Hive-32B model outperforms rivals like OpenAI's GPT-5.4 and Anthropic's Claude Opus 4.6 on key benchmarks.

Jul 24·4 min read