According to an NVIDIA blog post, partner companies post-trained NVIDIA's open Nemotron models for specialized use: Harvey's legal AI and LangChain's Deep Agents each ran at roughly 10x lower cost than closed rivals, Arcee AI cut inference to about 90 cents per million output tokens (~20x cheaper) while ranking second on PinchBench, and H Company's computer-use agent scored over 76% on OSWorld-Verified.
Enterprise Customization Through Post-Training: Vision, Legal, and Clinical Models
According to NVIDIA's blog post, several companies took NVIDIA's open Nemotron models and post-trained them on proprietary data to build task-specific systems rather than relying on general-purpose closed models.
- H Company built Holotron 3 Nano by post-training Nemotron 3 Nano Omni on proprietary computer-use data. NVIDIA reports the resulting model achieved higher than 76% accuracy on OSWorld-Verified, a benchmark for computer-use tasks, and NVIDIA states it matched other leading frontier models "at a fraction of the cost." (E1)
- Harvey post-trained Nemotron 3 Ultra on its own legal benchmark. NVIDIA's post says Harvey reached accuracy matching leading closed models on complex legal tasks, at at least 10x lower cost per run. (E2)
- Abridge is customizing Nemotron to build what NVIDIA describes as "the first foundation model purpose-built for clinical conversations." (E7)
Across these three domains — computer use, law, and clinical dialogue — the common pattern in NVIDIA's account is the same: take an open Nemotron model and post-train it on domain-specific data rather than building or licensing a closed model from scratch.
Cost Competitiveness: Open vs. Closed Model Economics
NVIDIA's blog post frames cost reduction as a central argument for open models, citing three separate partner figures.
| Company | Base Model / Deployment | Reported Cost Difference | Benchmark Result |
|---|
| Harvey | Nemotron 3 Ultra (legal tasks) | At least 10x lower cost per run than leading closed models | Matches leading closed models on complex legal tasks |
| LangChain | Nemotron 3 Ultra (Deep Agents harness) | Approximately 10x lower cost per run than leading closed alternatives | Top agent accuracy among open models |
| Arcee AI | Nemotron on NVIDIA Blackwell | Roughly 90 cents per million output tokens — approximately 20x cheaper than comparable closed frontier models | Ranked second on PinchBench, fully open weight |
According to NVIDIA, LangChain tuned its Deep Agents framework for Nemotron 3 Ultra — adjusting prompts, tools, and middleware without retraining the model — and reported top agent accuracy among open models at roughly 10x lower cost per run than leading closed alternatives. (E3) NVIDIA also states that Arcee AI, running Nemotron on the Blackwell platform, achieved inference costs of about 90 cents per million output tokens, roughly 20x cheaper than comparable closed frontier models, while ranking second on PinchBench and keeping the model fully open weight. (E4) Taken together with Harvey's reported 10x figure, three independent partners each cite cost reductions clustered in the 10x-to-20x range against closed alternatives, according to NVIDIA's post.
New Applications in AI Agents and Enterprise Search
Beyond single-task post-training, NVIDIA's blog highlights two applications where Nemotron is paired with agentic frameworks or larger models.
- LangChain adapted its Deep Agents harness specifically for Nemotron 3 Ultra, reporting the top agent accuracy among open models at approximately 10x lower cost per run than leading closed alternatives. (E3)
- Glean built Waldo, an agentic search model that, per NVIDIA's post, "pairs Nemotron with larger closed models to deliver enterprise search at significantly lower latency and with fewer tokens." NVIDIA's blog does not disclose specific latency or token figures for this pairing. (E5)
Breakthroughs in Healthcare Specialization
NVIDIA's post cites two healthcare companies building on Nemotron for clinical use cases.
- Abridge is customizing Nemotron to build what NVIDIA calls "the first foundation model purpose-built for clinical conversations." (E7)
- Heidi Health, according to NVIDIA, "is delivering frontier-quality outcomes in clinical documentation without needing frontier-scale compute." (E8)
Both claims are NVIDIA's characterization of partner outcomes; the blog post does not provide accuracy scores, error rates, or compute-cost figures for either company's clinical deployment.
Global Expansion: Language Localization and Ecosystem Collaboration
NVIDIA's post also frames Nemotron's openness as enabling both language localization and cross-company collaboration.
- YTL AI Labs post-trained a Nemotron model for the Malay language, which NVIDIA says puts "locally customized AI in the hands of Malaysia's developer community to further its AI capabilities." (E6)
- NVIDIA describes the NVIDIA Nemotron Coalition as an effort to turn open model development into an ecosystem effort, "bringing model builders and developers together to improve Nemotron through shared data, evaluations and domain expertise." (E9)
What This Means
All of the figures and quotes above come from a single NVIDIA blog post describing its own partners' results, so they should be read as NVIDIA's account of what those companies report — not independently verified benchmarks. Within that account, a consistent economic pattern emerges: three separate companies (Harvey, LangChain, Arcee AI) each report cost reductions in the same broad 10x-to-20x range relative to closed models, while claiming accuracy at or near parity — Harvey and LangChain by matching or topping open-model leaderboards, Arcee AI by ranking second on PinchBench. Alongside this cost pattern, the customization use cases NVIDIA cites span distinct domains — computer-use agents (H Company, 76% on OSWorld-Verified), enterprise search (Glean), clinical conversation and documentation (Abridge, Heidi Health), and language localization (YTL AI Labs) — suggesting that, according to NVIDIA's telling, the post-training approach is being applied across a wider range of specialized tasks than cost figures alone capture.