fullauto.online

Guide

AI Model Price-to-Performance Guide

A living comparison of 400+ models across five price tiers. Updated every Monday with fresh OpenRouter pricing, new releases, and benchmark data.

Published
13 Aug 2026
Reading
18 min
Class
pricing

half-life 45dfrom 13 Aug 2026

Introduction: Why Price-to-Performance Matters More Than Raw Capability

Every week a new model claims the top spot on a leaderboard. Every month prices drop 30–50%. The model that was "best" in January is overpriced by March. Chasing raw benchmarks is a losing game — the only metric that compounds is value per completed task.

This guide exists because the industry publishes capability scores and price lists separately. You get MMLU-Pro here, dollars-per-million-tokens there, and context windows somewhere else. Making a decision means stitching them together yourself, every time. We do that stitching once a week so you don't have to.

The core principle

Do not compare price per token. Compare price per passing case on your workload. A model at 1/40th the token price is not 1/40th the cost if the agent loop runs four times as long — and on agentic work, it often does.

How to Read This Guide

Price Tiers (by input cost per 1M tokens)

TierInput $/MTypical Output $/MBest For
Free$0.00$0.00Experimentation, low-stakes classification, local fallback
Budget< $1.00$0.10–$1.00High-volume extraction, routing, triage, bulk classification
Mid-Range$1.00–$5.00$2.00–$15.00Default agent loops, coding assistants, RAG, summarisation
Premium$5.00–$15.00$15.00–$75.00Complex reasoning, agentic coding, long-horizon tasks
Frontier> $15.00> $75.00Escalation tier only — measured failures of cheaper models

Key Columns Explained

  • Input $/M — Cost per 1 million input tokens (prompt + context)
  • Output $/M — Cost per 1 million output tokens (completion)
  • Cache Read $/M — Discounted rate when re-reading cached context (critical for agent loops)
  • Context — Maximum context window in tokens
  • Intelligence Index — Artificial Analysis composite (0–100, higher = more capable)
  • Coding Index — Artificial Analysis coding-specific score
  • Agentic Index — Artificial Analysis tool-use / multi-step score

Hidden Costs to Watch

  • Output token inflation — Reasoning models (o1, o3, Fable 5) generate 3–10× more output tokens for the same answer
  • Context bloat — Agent loops re-read context every turn; cache hits save 50–90% on input
  • Rate limits — Free tiers often cap at 20–50 req/min; not viable for production
  • Batch discounts — 50% off for async batch jobs (OpenAI, Anthropic, Google)
  • Tokenizer drift — Newer tokenisers (Claude 4.7+) produce ~30% more tokens for the same text

The Model Tier List

Data pulled from OpenRouter API on 13 Aug 2026. 409 models total. Router/meta-models excluded from tier counts below.

Free Tier — 18 Models

Zero cost via OpenRouter's free tier. Strict rate limits (typically 20 req/min). No SLA. Use for prototyping, evals, and non-production workloads.

ModelProviderContextBest ForNotes
Nemotron 3 UltraNVIDIA1MReasoning, codingStrongest free model; 86.8 MMLU-Pro
Nemotron 3.5 LightningNVIDIA1MFast reasoningOptimised for speed
Nemotron 3 SuperNVIDIA262KGeneral purpose120B params
Nemotron 3 Nano / Nano OmniNVIDIA256KLightweight tasks30B / multimodal variants
Gemma 4 31B / 26BGoogle262KOpen-weight localAlso available paid at $0.10–0.12/M
GPT-OSS 20B / 120BOpenAI131KOpen-weight codingAlso paid at $0.03/M
Laguna S / XS 2.1Poolside262KCode generationSpecialised coding models
North Mini CodeCohere256KCode tasksSmall specialised model
Nemotron 3.5 Content SafetyNVIDIA128KModerationSafety-specific
Nemotron Nano 12B VL / 9BNVIDIA128KVision + textMultimodal free options
Free Models RouterOpenRouter200KAuto-routingMeta-model across free tier

Budget Tier — ~150 Models (Input < $1/M)

Production-viable for high-volume, verifiable tasks. Many offer cache reads at 10–50% of base price.

ModelInput $/MOutput $/MCache $/MContextBest For
Ling 2.6 Flash (InclusionAI)$0.01$0.03$0.002262KUltra-cheap classification
Granite 4.0 H Micro (IBM)$0.017$0.112131KEnterprise lightweight
Mistral Nemo$0.019$0.03131KGeneral budget
Ling 3.0 Flash$0.021$0.063$0.0042262KCheap reasoning
Nex N2 Mini$0.025$0.10$0.0025262KBudget agent loops
GPT-5 Nano (batch)$0.025$0.20$0.0025400KBatch processing
Llama 3.2 1B$0.027$0.20160KTiny context tasks
Solar Pro4 (Upstage)$0.03$0.12$0.006524KLong context budget
Qwen 3.7 Flash$0.03$0.13$0.0061M1M context cheap
GPT-OSS 120B / 20B$0.03$0.13–0.17$0.03131KOpen-weight paid
Nova Micro (Amazon)$0.035$0.14128KAWS integration
Command R7B (Cohere)$0.0375$0.15128KRAG optimised
Qwen 3-30B-A3B$0.048$0.193262KMoE efficiency
Granite 4.1 8B$0.05$0.10$0.05131KEnterprise balanced
Nemotron 3 Nano 30B$0.05$0.20$0.025262KNVIDIA stack
GPT-5 Nano$0.05$0.40$0.005400KOpenAI cheap tier
Gemini 2.5 Flash-Lite (batch)$0.05$0.20$0.011MGoogle batch
GPT-4.1 Nano (batch)$0.05$0.20$0.01251MOpenAI batch

Mid-Range Tier — ~100 Models (Input $1–5/M)

The sweet spot for most production agent loops. Strong capability, reasonable cost, good cache pricing.

ModelInput $/MOutput $/MCache $/MContextBest For
GPT-5.6 Luna / Luna Pro$0.10$0.60$0.011.05MHigh-volume agents
GPT-5 Mini (batch)$0.125$1.00$0.0125400KBatch coding
GPT-5 Nano$0.05–0.10$0.40–0.625$0.005–0.01400KCost-optimised
Gemini 2.5 Flash-Lite$0.10$0.40$0.011MLong context cheap
Gemini 3.1 Flash-Lite$0.125–0.25$0.75–1.50$0.0125–0.0251MGoogle mid-tier
DeepSeek V4 Flash$0.14$0.28$0.0281MReasoning value
Qwen 3.6 Flash$0.1875$1.1251MQwen speed tier
Qwen 3.6 35B-A3B$0.15$1.00$0.05262KMoE balanced
Mistral Small 3.1 24B$0.351$0.555128KMistral efficient
GPT-4.1 Mini$0.40$1.60$0.101MOpenAI workhorse
Claude Haiku 4.5$1.00$5.00$0.10200KFast classification
Grok Build 0.1 (xAI)$1.00$2.00$0.20256KCheap reasoning
GPT-5.6 Terra / Terra Pro$1.00$6.00$0.101.05MOpenAI default agent
Claude Sonnet 5 (batch)$1.00$5.00$0.101MAnthropic batch
Gemini 2.5 Flash$0.30$2.50$0.031MGoogle default
GPT-4.1$2.00$8.00$0.501MOpenAI strong
Claude Sonnet 4.5 / 4.6$3.00$15.00$0.301MAnthropic strong
Gemini 2.5 Pro$1.25$10.00$0.1251MGoogle reasoning
Qwen 3.7 Plus$0.32$1.28$0.0641MQwen strong

Premium Tier — ~30 Models (Input $5–15/M)

Specialised for hard reasoning, agentic coding, and long-horizon tasks. Use as escalation, not default.

ModelInput $/MOutput $/MCache $/MContextBest For
Claude Opus 5 (batch)$2.50$12.50$0.251MBatch agentic coding
GPT-5.6 Sol / Sol Pro (batch)$2.50$15.00$0.251.05MOpenAI batch premium
Claude Opus 4.8 / 4.7 / 4.6 (batch)$2.50$12.50$0.251MAnthropic batch premium
GPT-5.5 / 5.4 (batch)$2.50–7.50$15.00–60.00$0.25–0.50400K–1.05MOpenAI previous gen
Claude Opus 5$5.00$25.00$0.501MTop-tier reasoning
GPT-5.6 Sol / Sol Pro$5.00$30.00$0.501.05MOpenAI top tier
Claude Fable 5 (batch)$5.00$25.00$0.501MAnthropic max capability
Claude Opus 4.1 / 4.5 / 4.6 / 4.7 / 4.8$5.00$25.00$0.50200K–1MAnthropic premium
GPT-5.5 / 5.4 / 5.3 / 5.2 Pro$5.00–21.00$30.00–168.00$0.50400K–1.05MOpenAI premium
Sakana Fugu Ultra$5.00$30.001MJapanese optimised

Frontier Tier — ~15 Models (Input > $15/M)

Escalation only. If you are here by default, you are overpaying. Route specific measured failures here.

ModelInput $/MOutput $/MCache $/MContextNotes
Claude Opus 4.7 Fast$30.00$150.00$3.001MFast variant
Claude Opus 4.8 Fast / Opus 5 Fast$10.00$50.00$1.001MFaster premium
GPT-5.5 Pro / 5.4 Pro / 5.2 Pro$15.00–30.00$90.00–180.00$0.50–0.75400K–1.05MOpenAI max
Claude Opus 4.1 / Opus 4$15.00$75.00$1.50200KAnthropic previous max
GPT-5 Pro$15.00$120.00400KOpenAI current max
O1 Pro$150.00$600.00$75.00200KExtreme reasoning
O3 Pro$20.00$80.00$10.00200KReasoning specialist

Best Value by Use Case

Best for Coding / Development

RecommendationModelInput $/MWhy
DefaultGPT-4.1 / GPT-5.6 Terra$1.00–2.00Strong coding, 1M context, good cache pricing
BudgetDeepSeek V4 Flash / Qwen 3.6 35B-A3B$0.14–0.15Surprisingly strong coding at budget prices
FreeNemotron 3 Ultra / GPT-OSS 20B$0.00Best free coding models available
EscalationClaude Opus 5 / GPT-5.6 Sol$5.00When default fails on complex refactoring
Local / Self-hostedQwen 2.5-Coder 32B / DeepSeek-CoderHardware onlyFull control, no API costs

Best for Reasoning & Math

RecommendationModelInput $/MWhy
DefaultClaude Sonnet 4.5 / 4.6$3.00Strong reasoning, 1M context, always-on thinking (Sonnet 5+)
BudgetDeepSeek V4 Flash / Grok Build 0.1$0.14–1.00Reasoning models at mid/budget prices
FreeNemotron 3 Ultra$0.0086.8 MMLU-Pro, 1M context
EscalationClaude Fable 5 / O3 Pro$10–20Always-on adaptive thinking, highest raw scores

Best for Long-Context Analysis (100K+ tokens)

RecommendationModelInput $/MContextWhy
DefaultGemini 2.5 Pro / Flash$0.30–1.251MBest cache pricing ($0.01–0.125/M), huge window
BudgetGemini 2.5 Flash-Lite / Qwen 3.7 Flash$0.03–0.101MCheapest 1M context
OpenAI stackGPT-4.1 / GPT-5.6 Terra$1.00–2.001MGood cache, strong retrieval
FreeNemotron 3 Ultra / 3.5 Lightning$0.001MOnly free 1M context models

Best for Chatbot / Customer Service

RecommendationModelInput $/MWhy
DefaultClaude Haiku 4.5 / GPT-4.1 Mini$0.40–1.00Fast, cheap, 200K–1M context, good instruction following
High-volumeGemini 2.5 Flash-Lite / Mistral Small 3.1$0.09–0.125Sub-$0.15/M input, decent quality
Free tierFree Models Router / Nemotron 3.5 Lightning$0.00Auto-routes to available free model

Best for Content Generation

RecommendationModelInput $/MWhy
DefaultGPT-4.1 / Claude Sonnet 4.5$2.00–3.00Natural tone, good instruction following
BudgetGemini 2.5 Flash / Qwen 3.7 Plus$0.30–0.32Strong creative writing at low cost
Long-formGemini 2.5 Pro / GPT-5.6 Terra$1.00–1.251M context for book-length generation

Best for Agentic Workflows (Tool Use, Multi-Step)

RecommendationModelInput $/MWhy
DefaultClaude Sonnet 4.5 / 5 / GPT-5.6 Terra$1.00–3.00Best tool-use training, reliable structured output
BudgetDeepSeek V4 Flash / Grok Build 0.1$0.14–1.00Strong function calling at low cost
EscalationClaude Opus 5 / Fable 5$5.00–10.00When multi-step fails on cheaper models
LocalQwen 3-Coder / Nemotron 3 Ultra (self-hosted)HardwareFull control over tool execution

Best for Multimodal (Vision + Text)

RecommendationModelInput $/MWhy
DefaultGPT-4.1 / Gemini 2.5 Flash$0.30–2.00Strong vision, 1M context includes images
BudgetQwen 3-VL 32B / Nemotron Nano 12B VL$0.00–0.104Open-weight vision models
SpecialisedGPT-5 Image / Gemini 2.5 Flash Image$0.30–2.50Image generation + understanding

The Free Models Worth Using — Detailed Breakdown

All 18 free models on OpenRouter (as of 13 Aug 2026), with specific guidance on when each is the right choice.

Tier 1: Production-Capable Free Models

ModelContextStrengthsWeaknessesUse When
Nemotron 3 Ultra1MHighest MMLU-Pro (86.8), strong reasoning, 1M contextRate limited, no SLA, slower than paidEvals, prototyping, non-prod reasoning tasks
Nemotron 3.5 Lightning1MOptimised for speed, 1M contextNewer, less benchmarkedFast free reasoning needed
GPT-OSS 20B131KOpenAI's open-weight, good codingSmaller context, 20B paramsFree coding assistant, local fallback
GPT-OSS 120B131KStronger than 20B, open-weightStill limited contextBetter free coding, can self-host

Tier 2: Specialised Free Models

ModelContextSpecialisationUse When
Nemotron 3 Super262KGeneral purpose, 120BNeed more capability than Nano
Nemotron 3 Nano / Nano Omni256KLightweight / multimodalCheap classification, vision tasks
Gemma 4 31B / 26B262KGoogle open-weight, strong baseWant Google model, can self-host
Laguna S 2.1 / XS 2.1262KCode-specialised (Poolside)Free coding-specific model
North Mini Code256KCode-focused (Cohere)Small code tasks
Nemotron Nano 12B VL / 9B128KVision + textFree multimodal needed

Tier 3: Meta-Models & Safety

ModelContextPurposeUse When
Free Models Router200KAuto-routes across free tierWant any free model, don't care which
Nemotron 3.5 Content Safety128KContent moderationNeed free safety classifier
Lyria 3 Pro / Clip Preview1MMusic generation (Google)Audio generation experiments

Critical limitation

Free tier rate limits (typically 20 req/min, 1000 req/day) make these unsuitable for production traffic. Use for development, evaluation, and internal tools only. For any user-facing production workload, budget at least $0.03/M (GPT-OSS paid tier) or $0.10/M (Gemini Flash-Lite).

Hidden Costs Deep Dive

1. Context Caching — The Agent Loop Multiplier

In a typical agent loop, the same system prompt, tool definitions, and conversation history are re-sent every turn. Without caching, a 10-turn loop with 20K context = 200K input tokens billed at full price. With caching, only the first turn pays full price; subsequent turns pay the cache-read rate (typically 5–20% of base).

ProviderCache Read DiscountExample (Sonnet 4.5)
Anthropic90% off (10% of base)$3.00 → $0.30/M
OpenAI50–95% off$2.00 → $0.10–1.00/M
Google80–90% off$0.30 → $0.03/M
OpenRouter (varies)50–95% offModel-dependent

Design implication: Put stable content (system prompt, tools, RAG context) first. Volatile content (user messages) last. This maximises cache hits.

2. Output Token Inflation — Reasoning Models

Models with always-on thinking (Claude Fable 5, Opus 5, O1, O3) generate 3–10× more output tokens for the same visible answer. A 500-token answer may cost 5,000 output tokens. At $50/M output (Fable 5), that 500-token answer costs $0.25, not $0.025.

ModelThinking ModeTypical Output MultiplierEffective Output $/M
Claude Fable 5Always on5–10×$250–500/M
Claude Opus 5Default high3–5×$75–125/M
O1 / O3 ProAlways on5–10×$300–600/M
O3 / O4 MiniDefault high3–5×$13–22/M
Sonnet 5 / Haiku 4.5Configurable1–3×$5–15/M

3. Rate Limits — Free vs Paid

Free tiers: 20–50 req/min, daily caps. Paid tiers: 1,000–10,000+ req/min. If your agent loop needs parallel execution or burst capacity, free tier will throttle you.

4. Batch Discounts — 50% Off for Async Work

OpenAI, Anthropic, and Google all offer 50% off for batch/async processing (24–48hr turnaround). Use for: evals, bulk extraction, offline report generation, nightly processing. Not for user-facing latency.

5. Tokenizer Drift — Newer Models Cost More Per Character

Claude 4.7+ tokenizer produces ~30% more tokens for the same English text vs 4.6 and earlier. GPT-4.1 tokenizer is similar. Budget 30% more input tokens when migrating to newer model generations.

How to Test Models for Your Use Case

Do not trust benchmarks. Do not trust this guide. Test on your data, your prompts, your success criteria.

Minimum Viable Evaluation (40 Cases, 1 Hour)

  1. Collect 40 real cases from your production logs. Weight toward failures — the cases that broke your current model.
  2. Define pass/fail per case. Binary is fine: "Did the agent complete the task correctly?"
  3. Run on current model. Record: pass rate, total input tokens, total output tokens, wall-clock time, total cost.
  4. Run on candidate model(s). Same 40 cases, same prompts, same tools.
  5. Compare cost per passing case. Not cost per token. Not pass rate alone. Cost per pass = total cost / passing cases.

The maths

Model A: $0.10/M, 80% pass rate, 50K tokens/case → $0.00625 per pass
Model B: $2.00/M, 95% pass rate, 20K tokens/case → $0.00421 per pass
The expensive model is cheaper per successful outcome.

What to Measure

  • Pass rate — Binary success on your task definition
  • Turns to completion — Agent loop iterations (directly drives cost)
  • Total tokens (in + out) — Actual bill driver
  • Wall-clock latency — User experience
  • Failure modes — Categorise: tool error, hallucination, reasoning error, timeout

Automation: The Weekly Eval Habit

Add this to your CI/CD. Every Monday, run your 40-case eval on the current default model + 2 candidates. If a candidate beats the default on cost-per-pass for two consecutive weeks, promote it. This guide's weekly update exists to give you the candidate shortlist.

Weekly Update Log

This section is updated every Monday with the automated pipeline. Manual annotations by DaedalusResearch.

Week of 13 Aug 2026 — Initial Publication

  • Data refresh: Pulled 409 models from OpenRouter API
  • New free models: Nemotron 3.5 Lightning, Gemma 4 31B/26B, GPT-OSS 20B/120B, Laguna S/XS 2.1, Nemotron Nano VL variants
  • Price drops: GPT-5 Nano at $0.05/M (batch $0.025), Gemini 2.5 Flash-Lite at $0.10/M, DeepSeek V4 Flash at $0.14/M
  • New premium models: GPT-5.6 Luna/Terra/Sol families, Claude Opus 5 / Fable 5, Qwen 3.8 series
  • Notable: 18 free models now available (up from ~12 in July). Budget tier deepens significantly.
  • Research note: DeepSeek V4 Flash ($0.14/M) emerges as the standout budget reasoning model — 1M context, cache at $0.028/M, strong coding.

Week of 6 Aug 2026

  • Price drops: GPT-4.1 Nano batch at $0.05/M, Mistral Nemo at $0.019/M
  • New models: Qwen 3.7 Flash/Plus, Nemotron 3 Ultra free tier
  • Research note: Mistral Nemo ($0.019/M) is the new budget floor for general-purpose — but 131K context limits long-document work.

Week of 30 Jul 2026

  • Major release: GPT-5 family (Nano/Mini/Pro) with batch pricing
  • Anthropic: Sonnet 5 / Opus 5 / Fable 5 / Haiku 4.5 launch
  • Research note: Fable 5's always-on thinking changes the cost calculus — effective output cost 5–10× listed rate. Route selectively.

Week of 23 Jul 2026

  • Google: Gemini 2.5 Pro/Flash with 1M context and aggressive cache pricing ($0.01–0.125/M cache reads)
  • Research note: Gemini 2.5 Flash-Lite at $0.10/M input + $0.01/M cache read = best value for long-context RAG.

Week of 16 Jul 2026

  • DeepSeek: V4 Pro/Flash launch — $0.14/M Flash is a budget reasoning breakthrough
  • Qwen: 3.5/3.6 series with MoE efficiency

Automation & Maintenance

This guide is regenerated weekly via an automated pipeline:

  1. Data Collection (Sunday 02:00 UTC): Python script fetches OpenRouter /api/v1/models, extracts pricing, context, cache rates. Stores JSON snapshot in versioned history.
  2. Diff Detection: Compares against previous week — flags new models, price changes >10%, removed models, context changes.
  3. Enrichment: DaedalusResearch adds Artificial Analysis benchmark scores (intelligence, coding, agentic indices) where available.
  4. Draft Generation: Script rebuilds tier tables, use-case recommendations, and update log entry.
  5. Editorial Review (Monday 09:00 BST): DaedalusContent reviews diff, adds qualitative annotations, verifies no contradictory positioning with DaedalusDesign service tiers.
  6. Publish: Updated HTML deployed to fullauto.online via build pipeline. RSS/feed updated.

Data Collection Script

/srv/samba/workspace/Projects/Daedalus Design/Development/scripts/fetch-openrouter-pricing.py

#!/usr/bin/env python3
"""
Fetch OpenRouter model pricing and metadata.
Run weekly via cron/N8N. Outputs JSON snapshot to data/pricing/YYYY-MM-DD.json
"""
import json
import requests
import pathlib
from datetime import datetime

OUT_DIR = pathlib.Path(__file__).parent.parent.parent / "data" / "pricing"
OUT_DIR.mkdir(parents=True, exist_ok=True)

def fetch():
    resp = requests.get("https://openrouter.ai/api/v1/models", timeout=30)
    resp.raise_for_status()
    data = resp.json()

    # Normalise
    models = []
    for m in data.get("data", []):
        pricing = m.get("pricing", {})
        models.append({
            "id": m["id"],
            "name": m["name"],
            "context": m.get("context_length"),
            "input_per_m": float(pricing.get("prompt", 0)) * 1_000_000,
            "output_per_m": float(pricing.get("completion", 0)) * 1_000_000,
            "cache_read_per_m": float(pricing.get("input_cache_read", 0)) * 1_000_000 if pricing.get("input_cache_read") else None,
            "modalities": m.get("architecture", {}).get("input_modalities", []),
            "reasoning": m.get("reasoning", {}),
            "created": m.get("created"),
        })

    snapshot = {
        "fetched_at": datetime.utcnow().isoformat() + "Z",
        "total_models": len(models),
        "models": models,
    }

    date_str = datetime.utcnow().strftime("%Y-%m-%d")
    (OUT_DIR / f"{date_str}.json").write_text(json.dumps(snapshot, indent=2))
    # Also write latest.json for easy access
    (OUT_DIR / "latest.json").write_text(json.dumps(snapshot, indent=2))
    print(f"Saved {len(models)} models to {OUT_DIR}/{date_str}.json")

if __name__ == "__main__":
    fetch()

N8N Workflow

Workflow ID: daedalus-price-guide-weekly on Nucleus N8N (videns.tail0d37a1.ts.net:5678).

  • Trigger: Cron (Mondays 02:00 UTC)
  • HTTP Request: GET OpenRouter models API
  • Function: Diff against previous week (stored in N8N workflow data)
  • HTTP Request: POST to DaedalusContent review endpoint (internal)
  • Wait for approval
  • Execute: Python build script to regenerate guide HTML
  • Deploy: SCP to fullauto.online webroot