fullauto.online

← Model leaderboard

Model · stat sheet

Nemo 3 Ultra

NVIDIA·frontier
Overall
86
Rank
#10 / 34
Coding89
Terminal88
Reasoning88
Tool use85
Context84
Speed72
CodingTerminalReasoningTool useContextSpeed

NVIDIA's largest Nemotron; tuned for agentic pipelines and tight tool-calling on their own stack.

What it is

Nemo 3 Ultra is NVIDIA's largest Nemotron — a frontier-tier model from the company that sells most of the hardware the rest of this board runs on. The note on this page is specific about the positioning: tuned for agentic pipelines and tight tool-calling on NVIDIA's own stack. That qualifier matters, and it is the thing to test before committing.

NVIDIA's model line ships alongside their inference tooling — NeMo, NIM, TensorRT-LLM and the hosted catalogue on build.nvidia.com — so the practical case is usually "we are already standardised there". Model cards and exact release dates are on the linked channels.

How the composite is built

Overall 86 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Capability spans 84 to 89 — a genuinely frontier-flat profile — with Speed 72 the one axis below the rest.

Coding 89
One behind Grok 5's 90 and three behind DeepSeek V4's 92. Comfortably in the top capability group here.
Terminal 88
Level with Grok 5 and behind only the top terminal tiers, from Sol's 89 to Opus 5's 95. The standout axis, and the one that matches the agentic-pipeline pitch.
Reasoning 88
Level with Sonnet 5's 88 and five behind DeepSeek R2's 93. Planning is rarely the failure mode at this level.
Tool use 85
Level with Claude Haiku 4.5's 85 and eleven behind Opus 5's 96. Strong — but the note's claim is about NVIDIA's own stack, so measure it on your harness and tool schemas, not theirs.
Context 84
Level with Qwen 4 Coder and Llama 5. The lowest of its six relative to the board; long-context work goes elsewhere.
Speed 72
The cost of the tier: behind Terra's 75 and eighteen behind the small tiers. If throughput is the constraint, this is not your model.

Where it fits

Agentic pipelines on NVIDIA infrastructure — teams already running NIM or NeMo tooling, enterprise deployments where the vendor relationship and the support path matter as much as the scores. For a hands-off coding or operations loop with a human on the escalation path, the Terminal and Tool use pair is the strongest you can buy outside the very top tier.

Limits

  • Value is rated Fair — 22.9 index points per $1.05/M blended is 21.8 per dollar. Decent, and beaten clearly by the best-cost models here.
  • Speed 72 makes it a poor fit for interactive or high-volume work.
  • Context 84 is mid-field; the model is built for operating, not for holding entire systems at once.
  • Stack-specific tuning may not transfer. The board measures behaviour on generic harnesses; NVIDIA's own pipeline is a different environment.
  • Cheaper Nemotron listings are already unscored on this board — 3.5 Lightning at $0.08/$0.20 and 3 Super at $0.08/$0.45. Re-check the catalogue before buying the top tier.

Price and access

$0.60 per million input tokens and $2.40 per million output, blending to $1.05/M. Listed on OpenRouter, and available through NVIDIA's own hosted catalogue. Last scored 15 Sep 2026.

Alternatives on this board

  • GPT-5.6 Terra — 86 at $2/$12 with Context 90. The same overall from OpenAI at over four times the blended rate.
  • Grok 5 — 87 with Terminal 88 and Reasoning 91, the nearest profile outside NVIDIA.
  • DeepSeek V4 — 86 at $0.78/$1.57 with open weights, if self-hosting matters.
  • Claude Sonnet 5 — 88 at $2/$10 with Tool use 90, the upgrade if the loop is failing on orchestration.

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.