fullauto.online

← Model leaderboard

Model · stat sheet

Amazon Nova 2 Pro

Amazon·balanced
Overall
79
Rank
#27 / 34
Coding78
Terminal75
Reasoning79
Tool use80
Context85
Speed84
CodingTerminalReasoningTool useContextSpeed

Amazon's mid tier; well integrated with Bedrock, with solid tool use and long context.

What it is

Amazon Nova 2 Pro is the mid tier of Amazon's own model family, served through Amazon Bedrock. The Nova line has always been an infrastructure play rather than a capability one: Amazon does not need the best model on any axis, it needs a competent model that sits natively inside IAM, VPC endpoints, CloudWatch, Guardrails and the rest of the AWS control plane.

Judge it on that basis. At 79 overall, 27th of 34, nobody picks Nova 2 Pro because it wins a benchmark. They pick it because the procurement, the audit trail and the network boundary are already solved.

How the composite is built

Overall 79 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. Nova 2 Pro is strongest on the two axes weighted lightest, which is why the composite reads worse than the model feels in ordinary use.

Coding 78
Below the open-weights mid-field — two behind Gemma 3 27B, fourteen behind DeepSeek V4. Usable for scaffolding and small edits; not a serious coding model.
Terminal 75
The weakest axis and twenty behind Opus 5. Nova is not built to drive an autonomous shell loop, and nothing in the score sheet suggests otherwise.
Reasoning 79
Mid-table, level with Phi-5's parameter-efficient 79 and fourteen behind Gemini 3.6 Pro. Adequate for structured business logic, thin for open-ended analysis.
Tool use 80
The strongest capability axis and the one that matters most for its actual job — Bedrock-hosted agents calling defined APIs. Level with GLM-5 and GPT-5.6 Luna.
Context 85
Genuinely good for the tier: level with DeepSeek R2 and ahead of Llama 5 (84). Long documents at an enterprise price point is a sensible thing for Amazon to optimise.
Speed 84
Fast — twenty-two points ahead of Opus 5 and level with Gemma 3 27B. Combined with the context score, that makes it a decent document-processing model.

Where it fits

Inside AWS, doing bounded work: document processing, retrieval-augmented answering over S3 content, structured extraction, internal assistants where Bedrock Guardrails and IAM-scoped tool access are the actual requirement. It is a reasonable default for a team whose compliance story is already written in AWS terms and who would rather not add a second vendor.

Limits

  • It is not a coding or agent model. Coding 78 and Terminal 75 are the two heaviest axes on this index, and they are its two weakest scores. Do not put it behind a coding harness.
  • No published price on this board, so no value rating. Bedrock pricing varies by region, throughput mode and whether you have provisioned capacity — you have to price your own workload.
  • Bedrock also sells you better models. Anthropic's tiers are available on Bedrock too, so "we need to stay in AWS" is not by itself an argument for Nova.
  • Generation details are worth checking. Amazon's Nova naming has moved quickly; confirm the exact context window, modalities and region availability in the Nova user guide rather than relying on this sheet.

Access

Amazon Bedrock, with the usual AWS surface: the Bedrock Converse API, cross-region inference profiles, provisioned throughput for predictable latency, Guardrails for content policy and Knowledge Bases for retrieval. This board records no OpenRouter or OpenCode listing, so a gateway-based harness needs an AWS integration rather than a base-URL change. Last scored 15 Sep 2026.

Alternatives on this board

  • Claude Sonnet 5 — 88 at $2/$10, and available on Bedrock. The obvious upgrade without leaving AWS.
  • Gemini 3.6 Flash — 84 at $0.75/$3.75 with Context 93 and Speed 88. The same job done better, if you can use Google.
  • GPT-5.6 Luna — 81 at $0.20/$1.20, the best value on this board, for the same class of high-volume bounded work.
  • Mistral Small 3.2 — 79 at $0.09/$0.25 with open weights, if the requirement is really "cheap and inside our boundary".

Sources

Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.