Models
OpenRouter: the LLM gateway that lets you test 400+ models with one API key
OpenRouter gives you 409 models from 59 providers behind one API key. 18 free models including frontier-class options. The practical starting point for autonomous agents and AI frameworks.
- Published
- 13 Aug 2026
- Reading
- 11 min
- Class
- models
half-life 60dfrom 13 Aug 2026
Most businesses approach AI the same way they used to approach cloud: pick a vendor, sign a contract, build on their stack. Then the vendor changes pricing, deprecates a model, or you realise the model you picked is terrible at the specific thing you need it to do. Now you're locked in.
The fix is an LLM gateway — a single API endpoint that sits between your code and dozens of model providers. You write to one interface, swap models by changing a string, and never sign up for nine different API keys. OpenRouter is the most complete implementation of this idea today.
What is an LLM gateway?
An LLM gateway is exactly what it sounds like: a proxy that speaks the OpenAI API format on one side and fans out to 59+ providers on the other. You send a request to openrouter.ai/api/v1/chat/completions with model: "anthropic/claude-3.5-sonnet" or model: "google/gemini-2.5-flash" or model: "openrouter/auto" and it routes to the right provider, handles authentication, rate limiting, and fallbacks, then returns a standardised response.
The key property: your code does not change when you switch models. The prompt format, the response shape, the error handling — all identical. This is the difference between "I'll try a different model" taking five minutes versus five days of refactoring.
Why OpenRouter specifically
As of August 2026, OpenRouter aggregates 409 models from 59 providers. The catalogue includes:
- OpenAI — 94 models (GPT-4o, GPT-4o-mini, o1 series, GPT-3.5 Turbo)
- Qwen — 50 models (Qwen2.5, Qwen2.5-Coder, Qwen2.5-Math variants)
- Google — 39 models (Gemini 2.5 Pro/Flash, Gemma 2/3/4 series)
- Anthropic — 28 models (Claude 3.5 Sonnet/Haiku/Opus, Claude 3 series)
- Mistral — 18 models (Large 2, Small, Nemo, Codestral)
- DeepSeek — 13 models (V3, R1, Coder variants)
- NVIDIA — 13 models (Nemotron 3 Ultra/Super/Lightning, Nano)
- Meta/Llama — 8 models (Llama 3.1/3.2/3.3 variants)
- xAI — 6 models (Grok 2/3 series)
- Cohere — 5 models (Command R/R+, North)
- Amazon — 5 models (Nova series)
- Perplexity — 5 models (Sonar variants)
Plus models from Together AI, Fireworks, Groq, Cerebras, Novita, and dozens of smaller providers. The model list at openrouter.ai/models updates continuously — new models appear within hours of release.
The free tier — 18 models you can use today without spending a penny
This is the part most people miss. OpenRouter offers 18 free models with genuine capability. You get an API key, add a few dollars of credit (or not — some models are truly free), and start building. The free lineup as of August 2026:
| Model | Provider | Params | Context | Best for |
|---|---|---|---|---|
nvidia/nemotron-3-ultra-550b-a55b:free |
NVIDIA | 550B | 1M | Frontier reasoning, complex tasks |
nvidia/nemotron-3.5-lightning:free |
NVIDIA | — | 1M | Speed-optimised, long context |
nvidia/nemotron-3-super-120b-a12b:free |
NVIDIA | 120B MoE | 1M | Balanced capability/efficiency |
google/gemma-4-31b-it:free |
31B | 128k | Google's latest open model | |
google/gemma-4-26b-a4b-it:free |
26B MoE | 128k | Efficient MoE variant | |
cohere/north-mini-code:free |
Cohere | — | — | Agentic coding |
poolside/laguna-s-2.1:free |
Poolside | — | — | Coding agent model |
openai/gpt-oss-20b:free |
OpenAI | 20B | 128k | Open-weight GPT |
openrouter/free |
OpenRouter | Auto | Varies | Auto-routes to best free model |
liquid/lfm-40b:free |
Liquid AI | 40B | 32k | Non-transformer architecture |
nvidia/nemotron-3-nano:free |
NVIDIA | 8B | 128k | Small, fast, efficient |
The openrouter/free model is worth calling out: it automatically routes to whatever free model is currently performing best and has capacity. You don't pick — it picks for you. For prototyping, this is often the right default.
Load-bearing
The free models are not toys. Nemotron 3 Ultra is a 550B parameter frontier model with 1M context. Gemma 4 31B is Google's latest open release. You can build production agents on these without spending on inference — you only pay if you exceed the free tier limits or want premium models.
Plugging into your tools
OpenRouter speaks the OpenAI API format. Anything that works with OpenAI works with OpenRouter — you just change the base URL and API key. Three concrete examples:
Hermes Agent
In config.yaml:
providers:
- name: openrouter
type: openai_compatible
base_url: "https://openrouter.ai/api/v1"
api_key: "${OPENROUTER_API_KEY}"
models:
- "nvidia/nemotron-3-ultra-550b-a55b:free"
- "anthropic/claude-3.5-sonnet"
- "google/gemini-2.5-flash"
- "openrouter/auto"
Then reference openrouter/nvidia/nemotron-3-ultra-550b-a55b:free in your agent configs. Hermes treats it like any other provider.
Cherry Studio
Settings → Model Providers → Add Custom Provider:
- Name: OpenRouter
- Base URL:
https://openrouter.ai/api/v1 - API Key: your OpenRouter key
- Model List: fetch from
/modelsendpoint or add manually
Cherry Studio will now show all 409 models in its model picker. Switch between them mid-conversation.
OpenClaw
OpenClaw reads OPENAI_API_KEY and OPENAI_BASE_URL. Set:
export OPENAI_API_KEY="sk-or-v1-..."
export OPENAI_BASE_URL="https://openrouter.ai/api/v1"
OpenClaw now has access to every model on OpenRouter. The same pattern works for Continue.dev, Cline, LobeChat, AnythingLLM, n8n, and any OpenAI-compatible client.
Use case: testing autonomous agents across multiple models
This is where the gateway model pays off. You're building an agent loop — maybe a coding agent, a research agent, a customer-support agent. The loop makes dozens of LLM calls: planning, tool selection, execution, reflection, correction.
Different models fail in different ways. Claude 3.5 Sonnet might nail the planning but hallucinate tool arguments. GPT-4o might get the tools right but lose the thread on long contexts. Nemotron 3 Ultra might reason deeper but run slower. DeepSeek R1 might be cheaper but less reliable on structured output.
With OpenRouter, you run the same agent loop against five different models by changing one config value. You measure: success rate, latency, cost per task, token usage. You find the model that actually works for your agent, not the model that benchmarks well on MMLU.
This is how you avoid the "works on my machine with GPT-4" trap. You validate the agent architecture against model variance before you commit.
Use case: finding the best model for your specific business task
Not every task needs a frontier model. A classification task (routing support tickets, tagging content, sentiment analysis) runs fine on a 7B–20B model at a fraction of the cost. A coding agent needs the best reasoning you can afford. A creative writing task might prefer a different style entirely.
The old way: pick one model, hope it's good at everything.
The gateway way: route each task type to the model that scores best on your eval set. OpenRouter's openrouter/auto can even do this dynamically — it picks the model based on your prompt, though for production you'll want explicit routing.
Getting started — from zero to first response in five minutes
- Sign up at openrouter.ai. Email or GitHub.
- Create an API key — Settings → Keys → Create Key. Copy it; you won't see it again.
- Add credit (optional) — $5–10 gets you started on premium models. Free models work without credit.
- Pick your tool — Hermes Agent, Cherry Studio, OpenClaw, Continue.dev, or just
curl. - Configure — Set base URL to
https://openrouter.ai/api/v1, add your key, pick a model. - Send a request:
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/nemotron-3-ultra-550b-a55b:free",
"messages": [{"role": "user", "content": "Write a haiku about LLM gateways"}]
}'
That's it. You now have a working LLM gateway. Swap the model string to try something else.
The catch (there's always one)
OpenRouter adds a small per-token markup on premium models — typically 5–20% depending on the provider. For free models, there's no markup. You pay the provider's rate plus the gateway fee. For high-volume production workloads, direct provider APIs can be cheaper. But for exploration, prototyping, and multi-model workflows, the markup is the cost of not managing nine API keys, nine billing accounts, and nine rate-limit regimes.
Latency: one extra hop. Usually 50–200ms. For most agent loops, this is noise. For latency-critical paths, you'd route directly.
Data privacy: your prompts pass through OpenRouter's infrastructure. They log requests for debugging and analytics (opt-out available on paid plans). If you need zero-retention, use provider APIs directly or OpenRouter's BYOK (Bring Your Own Key) for select providers.
Conclusion
The LLM gateway pattern is the correct starting point for any team exploring autonomous agents, AI frameworks, or multi-model workflows. It decouples your architecture from any single provider's roadmap, pricing, or outages.
OpenRouter is the most complete gateway available today: 409 models, 59 providers, 18 free options including frontier-class models, OpenAI-compatible API, and integrations with every major AI tool. You can be productive in five minutes.
Don't commit to a model. Commit to a gateway. Try OpenRouter, find what works for your use case, and scale from there.