Model · stat sheet
Falcon 3
TII's open family; a permissive, widely deployable base with efficient inference.
What it is
Falcon 3 is the open-weights family from TII, the Technology Innovation Institute in Abu Dhabi — the lab that put the original Falcon 40B and 180B into the open in 2023, at a point when very little else of that scale was downloadable. Falcon 3 is a different proposition: a set of smaller models, released in late 2024, aimed at running efficiently on modest hardware rather than at competing with frontier releases.
It sits 33rd of 34 on this board. That is an honest placement against 2026 models and not a useful summary of the family, which was never trying to win a general capability index.
How the composite is built
Overall 73 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. The shape is flat and low: nothing above 83, nothing catastrophic, and one axis — Speed — well ahead of the rest.
- Coding 73
- Lowest on the board. Usable for short, conventional snippets and little beyond that. Nineteen points behind DeepSeek V4 and seven behind Gemma 3 27B.
- Terminal 69
- The lowest terminal score anywhere on this board. A model at this size does not maintain state across a long command sequence; treat shell autonomy as out of scope.
- Reasoning 74
- Lowest on the board, a point under OLMo 3's 75. Direct questions get reasonable answers; multi-step problems do not.
- Tool use 71
- Weak. Structured output will need constrained decoding or a validation-and-retry wrapper rather than trust in the schema.
- Context 76
- The best of its capability axes, which tells you where the family's engineering effort went. Modest by 2026 standards — check TII's model cards for the exact window on the size you intend to run.
- Speed 83
- The point of the model. A point ahead of OLMo 3 and one behind Gemma 3 27B's 84, from far fewer parameters. On constrained hardware the practical advantage is larger than this number implies.
Where it fits
Embedded and edge deployment, low-resource environments, and research where a permissively licensed base with a published training story matters more than raw score. It is also a reasonable fine-tuning base for a single narrow task: a 73 general score says little about what the family does after task-specific training, which is the honest way to evaluate a model this size.
Limits
- Do not give it autonomy. Terminal 69 and Tool use 71 are the two lowest relevant scores here. It belongs behind a deterministic wrapper, not in an agent loop.
- No published price on this board, so no value rating and no gateway listing recorded. In practice you host it yourself, and your cost is hardware.
- The licence is permissive but not plain Apache. TII's Falcon licence carries an acceptable-use policy. Read it before redistributing.
- Check what TII has shipped since. Falcon 3 dates from late 2024 and the lab has continued releasing; this sheet is a read on this family, not on TII's current best.
Access
Download the weights from TII's Hugging Face organisation and run them locally — llama.cpp, Ollama and vLLM all support the family, and quantised variants are published alongside the full-precision checkpoints. There is no first-party hosted API on this board and no OpenRouter listing recorded. Last scored 15 Sep 2026.
Alternatives on this board
- Mistral Small 3.2 — 79 at $0.09/$0.25 with Speed 88 and an Apache licence. Six points better and easier to deploy in every respect.
- Phi-5 — 78 with Reasoning 84 and Speed 90. The stronger choice if you want reasoning out of a small model.
- Gemma 3 27B — 79 at $0.08/$0.45, with a much larger ecosystem, if you have a GPU to spare.
- OLMo 3 — 74 from Allen AI, the other model here chosen for openness of process rather than score.
Sources
- TII's Falcon LLM site — family overview, technical reports and licence.
- The TII organisation on Hugging Face — model cards with the authoritative sizes, context windows and quantised checkpoints.
- On this site: local LLM inference, September 2026, and the full model board.
Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.