π«π· Version franΓ§aise
LLM Provider Comparison β Orkeon
Status as of 2026-09-19, derived from the source code (
src/core/Orkeon.Infrastructure/LLMs/) and from each provider's declaredLlmProviderCapabilities. Legend: β supported Β· β absent Β· β partial/generic Β· β not campaigned (declared from the vendor's documentation, pending the first real-execution campaign).
| Provider | Base class | SSE streaming | Native tool calling | Multi-turn chat (tool roles) | System message | top_p / stop | GBNF grammar | response_format | thinking | Vision | reasoning_content round-trip | Prompt cache | Timing metrics | Polly resilience | API key sanitization |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenAI | OpenAI-compat | β | β | β | β | β | β | β schema | β effort | β | β | β auto | β | β | β |
| Azure OpenAI | OpenAI-compat | β | β | β | β | β | β | β schema | β effort | β | β | β auto | β | β | β |
| Anthropic | HttpLlmProviderBase | β native | β | β | β (native, separate) | β | β | β schema | β toggle | β | β | β explicit | β | β | β |
| DeepSeek | OpenAI-compat | β | β | β | β | β | β | β object | β toggle | β | β | β metrics | β | β | β |
| Z.AI (GLM) | OpenAI-compat | β | β | β | β | β | β | β object | β toggle | β | β | β metrics | β | β | β |
| Together AI | OpenAI-compat | β | β | β | β | β | β | β schema | β | β | β | β auto | β | β | β |
| Mistral AI | OpenAI-compat | β | β | β | β | β | β | β schema | β effort | β | β | β auto | β | β | β |
| Qwen | OpenAI-compat | β | β | β | β | β | β | β object | β budget | β | β | β auto | β | β | β |
| Kimi / Moonshot | OpenAI-compat | β | β | β | β | β | β | β object | β toggle | β | β | β auto | β | β | β |
| Google Gemini | OpenAI-compat | β | β | β | β | β | β | β schema | β effort | β | β | β auto | β | β | β |
| Grok (x.AI) | OpenAI-compat | β | β | β | β | β | β | β schema | β effort | β | β | β auto | β | β | β |
| MiniMax | OpenAI-compat | β | β | β | β | β | β | β (accepted but non-binding β measured) | β (always-on inline, split out) | β | β | β auto | β | β | β |
| HuggingFace | OpenAI-compat | β | β | β | β | β | β | β object | β | β | β | β auto | β | β | β |
| OpenRouter β | OpenAI-compat | β | β | β | β | β | β | β schema (per endpoint) | β budget (reasoning object) |
β (per model) | β | β auto (+ cache_write_tokens) |
β (usage.cost exposed) |
β | β |
| Mammouth AI β | OpenAI-compat | β | β | β | β | β | β | β (undocumented) | β (undocumented) | β (per model) | β | β auto | β | β | β |
| Ollama | HttpLlmProviderBase | β | β (/api/chat) |
β (/api/chat) |
β (prepend) | β | β | β schema | β toggle | β (images) |
β | β | β | β | β |
How to read the capability columns
response_format, thinking and Vision are not hand-maintained here: each provider declares
an LlmProviderCapabilities value object, and OpenAICompatibleProviderBase translates it into
the OpenAI dialect once. Anthropic, Ollama and Qwen override the hook because their APIs speak
their own dialect.
response_formatβobjectmeans the API guarantees well-formed JSON;schemameans it validates against a JSON Schema server-side. A schema sent to anobject-only provider is downgraded with a warning, never in silence. Anthropic is schema-only: it has no equivalent ofjson_object, so a schema-less JSON request there is reported rather than sent.thinkingβeffortaccepts a level hint only;togglecan also switch reasoning on and off;budgetadditionally accepts an explicit token budget (Qwen only β Anthropic rejectsbudget_tokenswith a 400 on the current generation). On the wire, Anthropic's toggle is written asthinking: {type: adaptive|disabled}β a payload detail, not a capability level.- Anything a provider does not support is reported. An option declared in YAML on a provider that cannot honour it produces an actionable warning naming the option, the provider and the remedy. This was the actual defect the 2026-07-27 audit found: not the missing wiring, but its invisibility.
Notes
- β prompt cache (auto): the vendor caches prompt prefixes implicitly and reports the hit;
OpenAICompatibleProviderBase.ParseSuccessResponsereadsprompt_cache_hit/miss_tokensand the OpenAI-standardprompt_tokens_details.cached_tokensgenerically. β explicit (Anthropic) means the cache does nothing until acache_controlbreakpoint is placed β opt in withLlmCacheConfig/ the YAMLcache:block. - Anthropic SSE:
ChatStreamingAsyncparses the Messages API event stream natively since LLM-05; it used to fall back to a buffered emulation, so no token arrived early. - Ollama tool calling: goes through
/api/chatwhen the conversation declares tools, replays tool calls, or carries an image; everything else keeps/api/generate(NDJSON streaming, GBNF). The text-fallback protocol remains in charge for models without tool support. - Azure OpenAI: two API shapes β the dated deployment URL (default) and the v1 GA surface
(
api_version: v1), which is the only path to the Responses API and to the non-OpenAI models Azure resells. - HuggingFace: model identifiers accept a routing suffix (
:fastest/:cheapest/:preferred/:<partner>) β the only cost and latency lever on Inference Providers. - top_p / stop: Ollama only exposes
temperature+num_predict(= max_tokens). - OpenAI
max_completion_tokens: OpenAI retiredmax_tokenson its current models (the 2026-08-30 campaign failed ten modes on that one field), so the OpenAI dialect writesmax_completion_tokensβ accepted by the older generations too (verified ongpt-4o-minithe same day). The compatible vendors keepmax_tokens: the retirement is OpenAI's alone. - Gemini
response_format: undocumented on the compat surface when first audited (2026-08-18) and undeclared then; measured live on 2026-08-30, the surface acceptsjson_objectandjson_schemaand enforces the schema server-side, so the provider now declaresJsonSchema. - DeepSeek vision: arrived with
deepseek-v4-flash-vision-exp(measured 2026-08-30) and is native on the Flash tier since V4.1 Flash (2026-09-10): the defaultdeepseek-flashsees, the experimental companion is retired. Per provider vs per model as everywhere (D-03):deepseek-v4-prostays text-only and answers an image with the vendor's own error.deepseek-v4-flashis a retired model's name the API "temporarily" routes to V4.1 Flash β the default moved to the vendor's own id on 2026-09-19. - MiniMax: campaign-backed since 2026-08-30 (7/2/3, same day it was integrated). The
reasoning arrives INLINE β every reply opens with a
<think>block insidecontent, no separate field β and the dialect splits it out toreasoning_content, re-inlining it on replay as the vendor documents.response_formatis accepted but NON-BINDING (a schema is ignored,json_objectarrives fenced in markdown): the None declaration is a measurement (raw calls, 2026-08-30, recorded in the campaign catalog's MiniMax note β the archived M8 shows β because the None declaration keeps the option from ever being sent). Vision is per model (D-03):MiniMax-M2answers "I'm unable to view the image", and the VL family does not appear on the platform's/modelsβ no companion declarable yet. - Grok (x.AI): every declared capability is a live measurement β a full 12-mode campaign
passed against
api.x.aithrough the generic OpenAI dialect before the provider class existed (2026-08-30, archived underllmproviders-test/custom-endpoints/). Keys carry thexai-prefix, which the factory infers. - OpenRouter β : the model marketplace (445 models from 60 vendors on 2026-09-18) behind
one key, integrated documentation-first (LLM-09) β no campaign archived yet. Identifiers
are
vendor/model(mandatory prefix), with the:free/:nitro/:floorsuffixes and theopenrouter/autorouter slug (usable, refused as a default: the served model drifts β theserved_modelmetadata says who answered). The reasoning trace comes back asreasoning, neverreasoning_content(aReasoningFieldNamehook on the base); thinking travels as thereasoningrequest object (enabled/effort/max_tokens, so the Budget declaration is one of the transport β OpenRouter converts effort and budget into each other per model); the real charge arrives inusage.cost(thecostmetadata) with itsupstream_inference_cost/is_byok/cache_write_tokens/reasoning_tokensbreakdown; two constant attribution headers (HTTP-Referer,X-OpenRouter-Title) name Orkeon.json_schemais honoured per endpoint and the provider does not sendprovider.require_parameters: whether a schema can be silently ignored elsewhere is the first campaign's question. The advanced routing body (provider {β¦},models[],plugins[]) is not exposed. Thesk-or-v1-key prefix is documented by secondary sources only, so the factory does not infer from it yet. - Mammouth AI β : the French multi-model subscription whose included API credits drive
Orkeon, integrated documentation-first (LLM-09) β no campaign archived yet. On three
concordant clues (2026-09-18) the API is a LiteLLM proxy; nothing in the provider depends on
it. Identifiers are the vendors' own bare strings (
gpt-5.6-sol,claude-sonnet-5,gemini-3.7-flash), so the provider is reached by host (api.mammouth.ai) or by"Provider": "mammouth"and never inferred from a model name β the same string without a base URL keeps going to the vendor. Onlymessages,model,temperature,max_tokens,top_pandstreamare documented:response_formatand thinking stay undeclared (structured warning, never a silent drop) until the first campaign measures them β the MiniMax rule; vision is declared from the vendor's owntext, imagemodel list. Prices are the vendor's upper bounds (gemini-3.7-flashat 1.5 / 7.5 $/M, twice the direct price). - Anthropic identity-linked keys: refuse every request without an
anthropic-workspace-idheader (2026-08-30). SetLlmConfig.WorkspaceId(CLI:--workspace-id); classic keys need nothing. - Polly and API key sanitization: provided by
HttpLlmProviderBaseβ active everywhere.
Defaults and newer models β catalogue review of 2026-09-19
Every vendor's own model and pricing pages were read on 2026-09-19, alongside the public
catalogues of the two aggregators (OpenRouter, 447 models; Mammouth, 100). The rule is the
one LlmProviderDefaultModels states: a default changes only when the vendor retires the
name; otherwise a newer model is a candidate until a campaign has archived an M1 on it.
| Provider | Default (code) | Served on 2026-09-19 | Newer on the vendor's API | Verdict |
|---|---|---|---|---|
| OpenAI | gpt-5.6-sol |
yes β 4 / 20 $/M through 2026-11-21, the replacement target of both 2026 deprecation waves | gpt-6-astra (10 / 50, effort lowβ¦max with no none, 2Γ billing above 272K input tokens) |
stays β campaigned 2026-09-21: gpt-6-astra 10/2 with temperature: 1 pinned; it has no reasoning_effort: none, so function tools on /v1/chat/completions are refused (use /v1/responses) β not a default until the dialect speaks Responses; gpt-6-astra-pro is an aggregator label, not an OpenAI id |
| Anthropic | claude-sonnet-5 |
yes β 2 / 10 $/M made permanent, active until at least 2027-06-30 | claude-fable-5-1 (2026-09-01, 10 / 50, thinking always on, forced tool_choice refused) |
stays β campaigned 2026-09-21: claude-fable-5-1 11/1 twice (the red is M6, the text-fallback protocol; native tools, schema, vision and cache all green) β viable, at five times the price; claude-opus-5-fast is a speed flag, not an id; temperature / top_p / top_k return 400 on every model from Opus 4.7 on |
| Azure OpenAI | deployment | β | β | nothing to default to |
| Ollama | llama3.2 |
yes β 3B, text only, no thinking, a year old | qwen3.5:4b, gemma4:e4b, granite4.2:3b (tools + thinking, same size class) |
stays: a new default means a pull on every machine |
| Together AI | meta-llama/Llama-3.3-70B-Instruct-Turbo |
yes β 1.04 / 1.04 $/M, not scheduled | zai-org/GLM-5.3-Flash (1M, tools + JSON, 0.15 / 0.50), Qwen/Qwen3.5-9B, deepseek-ai/DeepSeek-V4.1-Flash; Llama 4 left serverless |
candidate GLM-5.3-Flash, a seventh of the price β campaigned 2026-09-21: 11/0/1, it reads the image (the first serverless vision model on Together, now the kit's vision companion) and its cache is read (8000 tokens, 0.99); Qwen/Qwen3.5-9B 10/0/2. Both ready; the bump is the owner's call |
| DeepSeek | deepseek-flash (was deepseek-v4-flash) |
yes β V4.1 Flash, 0.30 / 1.20 $/M peak, native vision | it is the newest | changed 2026-09-19: the old name is a retired model "temporarily" routed here; replayed 2026-09-21: 12/12 twice β M2 and M9, red for five campaigns on deepseek-v4-flash, are green on deepseek-flash |
| Kimi | kimi-k2.6 |
yes β 0.95 / 4.00 $/M, no retirement date | kimi-k3 (July 2026, 1M, 3 / 15, fixed sampling, no thinking field, always reasons) |
stays β campaigned 2026-09-21: kimi-k3 12/12 then 11/1 (M2 on the flattened shape), reasoning trace, cache 0.97, bare JSON where k2.6 now fences its json_object β ready, at three times the price; kimi-k2.5 and moonshot-v1-* retired 2026-08-31; docs now on platform.kimi.ai |
| Qwen | qwen3.7-plus |
yes β still the Plus tier, one of the three recommended models | qwen3.8-max (= qwen3.8-max-0902), qwen3.8-flash; no qwen3.8-plus |
stays β campaigned 2026-09-21: qwen3.8-max 12/12, qwen3.8-flash 11/1 (M2), both ready, and qwen3.7-plus is still 12/12; the 3.8 generation adds preserve_thinking |
| Mistral AI | mistral-medium-2604 |
yes β 1.50 / 7.50 $/M, not deprecated | nothing for chat since April (OCR 4.1 only) | stays; devstral-* / magistral ids sit in the deprecated table |
| HuggingFace | meta-llama/Llama-3.1-8B-Instruct |
yes β but tools on one of its four routed providers, and :fastest may pick another |
Qwen/Qwen3.5-9B (tools on three providers, from 0.10 / 0.15), zai-org/GLM-5.3-Flash, deepseek-ai/DeepSeek-V4.1-Flash |
candidate Qwen3.5-9B β campaigned 2026-09-21: 10/0/2, it sees and its tools hold; its first pass returned an empty M1 after 4113 tokens of thinking on the router's 4096 fallback β pin Llm:MaxTokens (or turn thinking off) before making it a default. The routing still explains the default's moving reds (M2 and M5 both red on the rc.4 pass) |
| Z.AI | glm-5.2 |
yes β still under "Latest Models", 1.40 / 4.40 $/M | glm-5.3 (2026-08-18), glm-5.3-flash (0.15 / 0.50, image + video), glm-5.3-flashx |
stays β campaigned 2026-09-21: glm-5.3-flash 12/12 twice (it sees, where 5.2 and 5.3 answer 1210; cache read) β the cleanest candidate in the fleet, at a ninth of the price; glm-5.3 10/2 (M2, M9) brings nothing over 5.2. The Toggle declaration still needs a per-model guard before a bump |
| Google Gemini | gemini-3.7-flash |
yes β "previous generation", no shutdown date, 0.75 / 3.75 $/M until 2026-12-31 then 1.50 / 7.50 | gemini-3.8-flash (GA 2026-09-02, same price, minimal thinking rejected) |
first candidate for a bump β campaigned 2026-09-21: gemini-3.8-flash 11/0/1 twice, identical to 3.7 mode for mode on the direct transport; OpenRouter and Mammouth not played (no key) |
| Grok (x.AI) | grok-4.6 |
yes β recommended, 2 / 6 $/M below 200K prompt tokens, 4 / 12 above | none | stays |
| MiniMax | MiniMax-M2 |
yes β listed as legacy, no retirement date | MiniMax-M3 (2026-06-01, 1M, image + video input, same 0.30 / 1.20) |
stays β campaigned 2026-09-21: MiniMax-M3 9/1/2 then 8/1/3 β it reads the image (now the kit's vision companion), M2 red like M2, the inline <think> block split out as on M2, 128 cached tokens once; the reasoning-format question is answered |
| OpenRouter | google/gemini-3.7-flash |
yes | google/gemini-3.8-flash (2026-09-02, same price) |
follows the direct default β not campaigned on 2026-09-21 (no key) |
| Mammouth AI | gemini-3.7-flash |
yes | gemini-3.8-flash |
follows the direct default β not campaigned on 2026-09-21 (no key) |
ModelPricingRegistry follows the same pages: the GPT-5.6 family and gpt-6-astra, Sonnet 5
at 2 / 10, Fable 5.1 / Fable 5 and Haiku 4.5.
Output caps β the documented maximum per model (LLM-10)
A request carries an output cap (max_tokens, max_completion_tokens at OpenAI, num_predict
at Ollama). The engine used to send 4096 for every model; a reasoning model spends that
budget thinking and answers empty, the tool-free retry that follows narrates the deliverable
instead of writing it, and the run goes green with no file (owner recette 2026-09-19,
kimi-k3). Since LLM-10 the cap left unpinned is the model's documented maximum, from the
LlmModelOutputLimits catalogue in Orkeon.Constants.Llm; a pinned value (Llm:MaxTokens, a
Studio profile, a crew's max_tokens) always wins; a model the catalogue does not know keeps
the 4096 fallback. Only documented figures go in, each read on 2026-09-19:
| Provider | Model | Cap sent | Source | Note |
|---|---|---|---|---|
| OpenAI (and Azure deployments of the same ids) | gpt-5.6-sol, gpt-6-astra |
128 000 | developers.openai.com/api/docs/models | reasoning counts inside the cap |
| Anthropic | claude-sonnet-5, claude-opus-5, claude-fable-5-1 (dated variants follow the family) |
128 000 | platform.claude.com/docs/en/about-claude/models/overview | thinking counts inside max_tokens; the field is required, an unknown Claude gets 4096 |
| Gemini | gemini-3.7-flash, gemini-3.8-flash (also under google/β¦ on OpenRouter and bare on Mammouth) |
65 536 | ai.google.dev/gemini-api/docs/models | includes thought tokens; 65 537 is a 400 |
| DeepSeek | deepseek-flash (and the retired deepseek-v4-flash* names it routes) |
393 216 | api-docs.deepseek.com/api/create-chat-completion | "1 to 384K"; 384 000 on Mammouth |
| Kimi | kimi-k3 |
131 072 | platform.kimi.ai/docs/guide/kimi-k3-quickstart | the vendor's own max_completion_tokens default; real bound 1M β prompt |
| Kimi | kimi-k2.6 |
131 072 | platform.kimi.ai/docs/guide/troubleshooting | a pin, not a documented cap: the bound is 256K β prompt; a rejection drops the field on the retry |
| Qwen | qwen3.7-plus, qwen3.8-flash, qwen3.8-max |
131 072 | alibabacloud.com/help/en/model-studio | the chain of thought has its own thinking_budget; 65 500 on Mammouth |
| Z.AI | glm-5.2, glm-5.3, glm-5.3-flash |
131 072 | docs.z.ai/guides/overview/concept-param | schema maximum; default 65 536 |
| Z.AI | glm-4.6v-flash |
32 768 | docs.z.ai/guides/vlm/glm-4.6v | β |
| MiniMax | MiniMax-M2 |
131 072 | platform.minimax.io/docs/guides/models-intro | "128k (including CoT)"; M3 is absent β the vendor publishes no output figure (512k appears only as a benchmark setting; 512 000 on Mammouth) |
| xAI | grok-4.6 |
128 000 | docs.x.ai/developers/rest-api-reference | no per-model cap; the vendor's own default when unset, reasoning excluded |
| Mistral | mistral-medium-2604 (mistral-medium-3-5) |
none | docs.mistral.ai/api/endpoint/chat | no output cap, only "prompt + max_tokens β€ context": the field is left out and the model writes to its window |
| Together | any | 4096 (fallback) | docs.together.ai/docs/serverless-models | no per-model output cap, only prompt + max_tokens β€ window. The window itself was the cap from 2026-09-19 to 2026-09-21, on the assumption that context_length_exceeded_behavior: truncate clamps it to window β prompt; the campaign of 2026-09-21 measured that two of three serverless engines refuse regardless (max_new_tokens, context_length_exceeded) and the third clamps on the buffered path only. The fallback holds, the flag is still sent, and the retry net drops the cap on those wordings; omitted, the field means 2048 (finish_reason: length) |
| Ollama | any | none (num_predict left out) |
docs.ollama.com/modelfile | -1, infinite generation is the runtime's default: a local model writes to its context |
| HuggingFace, Docker Model Runner, anything else | β | 4096 (fallback) | β | the router's bound is the routed provider's context, which differs per route; pin Llm:MaxTokens |
Two things the catalogue is honest about. Where the vendor bounds the cap by window β prompt
(Kimi, Together, Mistral), a fixed value near the window fails on any real prompt β so those
rows are either a pin, the fallback, or nothing. And a cap the endpoint refuses (Qwen "Range of
max_tokens", Kimi "prompt tokens + max_tokens exceeds", Gemini maxOutputTokens, DeepSeek's
422, Together's engines with max_new_tokens / context_length_exceeded) is retried once
without the field when it came from the
catalogue, with a warning naming the model β a pinned value is the user's, and its rejection
surfaces unchanged. The Studio profile editor reads the same catalogue: the hint under the
Β« Maximum response Β» field says what an empty field means for the chosen model, and invites a
pin when the model is unknown.
Per-model mandatory parameter values
Some models refuse a request unless a parameter carries one specific value. This is a different animal from the per-model capability gaps above (D-03 β a model lacking thinking or vision): here the capability exists, but the model dictates the value, and the vendor answers anything else with a 400. The constraint is per model, never per provider β the same vendor ships models with opposite demands, so a provider-level pin is one campaign away from breaking the sibling model. Every row below was measured live; nothing is inferred.
| Provider | Model | Parameter | Mandatory value | Vendor's own words | Measured |
|---|---|---|---|---|---|
| Kimi | kimi-k2.6 |
temperature |
1 |
invalid temperature: only 1 is allowed for this model |
2026-08-03 |
| OpenAI | gpt-5.6-sol |
temperature |
1 |
'temperature' does not support 0 with this model. Only the default (1) value is supported. |
2026-08-30 |
| OpenAI | gpt-5.6-sol |
reasoning_effort |
"none" when the request carries function tools on /v1/chat/completions |
Function tools with reasoning_effort are not supported for gpt-5.6-sol in /v1/chat/completions. To use function tools, use /v1/responses or set reasoning_effort to 'none'. |
2026-08-30 |
| OpenAI | gpt-6-astra |
temperature |
1 |
'temperature' does not support 0 with this model. Only the default (1) value is supported. |
2026-09-21 |
| OpenAI | gpt-6-astra |
reasoning_effort |
no "none" exists (low, medium, high, xhigh) β the Sol workaround is impossible, so function tools stay refused on /v1/chat/completions until the dialect speaks /v1/responses |
'reasoning_effort' does not support 'none' with this model. Supported values are: 'low', 'medium', 'high', and 'xhigh'. β and without it, Function tools with reasoning_effort are not supported for gpt-6-astra in /v1/chat/completions. To use function tools, use /v1/responses |
2026-09-21 |
| Anthropic | claude-sonnet-5 |
temperature |
1, or omit the field |
`temperature` is deprecated for this model. (1 and omission pass; 0 and 0.7 do not) |
2026-08-30 |
| Mistral | mistral-medium-2604 |
reasoning_effort |
high or none only |
reasoning_effort low is not supported for this model, supported values: [<ReasoningEffort.high: 'high'>, <ReasoningEffort.none: 'none'>] |
2026-08-30 |
| Mistral | mistral-medium-2604 |
top_p |
explicit 1 when temperature is 0 and reasoning is on (omission is NOT 1 there) |
top_p must be 1 when using greedy sampling. |
2026-08-30 |
The counter-example that makes the registry per-model: gpt-4o-mini β same provider as
gpt-5.6-sol β rejects reasoning_effort outright (Unrecognized request argument supplied: reasoning_effort, measured the same day). A provider-wide "always send none" would break it.
Machine-readable twin: llmproviders-test/lib/catalog.json, key requiredParams (per
provider, keyed by model id) β the campaign scripts resolve it automatically and the report
header prints the values actually used. In application code, express the same values through
LlmConfig (Temperature, Thinking = { Effort = "none" }); on the wrong value the vendor's
error comes back attributed (CapabilityMismatchHint names the thinking knob on the
tools-with-reasoning refusal).
What this table does not prove
Every row is backed by unit tests asserting the emitted payload β against a mocked HTTP
handler. A mock proves Orkeon sends what we believe it sends; it does not prove the vendor
accepts it. Real-execution evidence is tracked separately in the test matrix journal
; run a campaign with
orkeon llm probe --provider <name> --archive <dir>.