AI models evolve rapidly. New models are usually available in Gumloop within a day of their public release, so you have the latest options even if this page has not caught up yet.
Choosing a model
Open the model dropdown in Agent Preferences. The fastest way to choose is one of the options at the top:- Recommended: the best balance of speed, quality, and cost. This is the default for new agents, and currently points to Grok 4.6.
- Smartest: maximum intelligence for complex reasoning and agentic work.
- Auto: Gumloop picks a model per message instead of pinning one.
On Enterprise plans, admins can restrict which models a Custom Role may use. Models your role does not allow are hidden from this picker (and from workflow AI nodes), so you only see the models you are permitted to select.
Auto: let Gumloop choose the model
Auto is Gumloop’s model router, not a model from a provider. Select it once, and Gumloop picks which model runs each message.Auto sits at the top of the model dropdown in Agent Preferences.
How Auto picks a model
1
It reads your message first
Before the agent starts working, Gumloop classifies the message — how hard is it, how much reasoning does it need, are there attachments — and picks the cheapest tier of models that should still nail the outcome. While that happens the chat shows Choosing the best model.
2
It runs on a real model, and tells you which
The tier is a short, ordered list of models; Gumloop runs the best one you’re allowed to use. Once the run starts, the chat names it (Using Claude 5 Sonnet), and that same model is what shows up in the chat’s models list and its credit breakdown.
3
It can move up mid-run
If a run turns out harder than the message looked, Auto can escalate to a stronger model part-way through — the chat shows Switching to …. Escalation sticks for the rest of that run, so a run never climbs and then quietly drops back down. Auto never de-escalates a run it already promoted.
Where Auto can be used
Auto and restricted models
On Enterprise plans, Auto stays inside your organization’s model rules — it is not a way around them.- Both layers of restriction apply. Organization-wide model restrictions and a Custom Role’s model allow-list are both enforced on the routing decision, on the model Auto runs, on any mid-run escalation, and on fallbacks.
- A blocked model is never used. Auto skips it and takes the next permitted model in that tier.
- A fully blocked tier steps aside. If nothing in the chosen tier is permitted, Auto moves to the closest tier that does have a permitted model.
- If nothing is permitted, the run fails. Auto stops with a model-restricted error rather than running on a model your organization disallowed.
- Admins can restrict Auto itself. It appears in the restriction and Custom Role model lists like any other picker entry, so an allow-list that leaves it out hides it. Agents already saved on Auto then simply run on your organization’s normal model choice, with no routing.
- Models with 30-day data retention are never routed to. Claude Fable models have to be chosen deliberately, so Auto never selects one.
Auto FAQ
Is Auto a model?
Is Auto a model?
No. It’s a router. Selecting Auto means “decide for me”: Gumloop classifies each message and runs a real provider model — Claude, GPT, Gemini, Grok, and others — behind it.
Does Auto always use the same model?
Does Auto always use the same model?
No. The decision is made per message, so the same agent can answer a quick question on a small fast model and a deep research prompt on a frontier model.
Which model did it actually use?
Which model did it actually use?
The chat tells you. It shows Choosing the best model while deciding, then names the model it runs (Using …) and any mid-run switch (Switching to …). The chat’s models list and credit details record the real models, never the routing guess.
Can it change models in the middle of a run?
Can it change models in the middle of a run?
Yes, upward. If the work turns out harder than expected, Auto escalates to a stronger model and stays there for the rest of that run.
What happens when my admin restricts a model?
What happens when my admin restricts a model?
Auto won’t use it. Organization restrictions and Custom Role allow-lists apply to every model Auto could pick, including escalations and fallbacks, so a restricted model is skipped in favor of a permitted one. If no permitted model is left anywhere, the run fails with a model-restricted error instead of bypassing the rule.
Can an admin turn Auto off?
Can an admin turn Auto off?
Yes. Auto is listed in organization restrictions and Custom Role model lists, so it can be blocked or left out of an allow-list. Agents saved on Auto fall back to running on a normal model, without routing.
Can I use Auto in a workflow AI node?
Can I use Auto in a workflow AI node?
No. Auto is for agents (including Gumball). Workflow AI nodes, the three model presets, and organization fallback models always name a specific model.
How is Auto billed?
How is Auto billed?
Exactly like the model that ran: token-based Chat & Reasoning, plus Compute, Tool Calls, and the orchestration fee. The routing decision itself appears as a small separate Model Routing charge in the chat’s credit breakdown.
What if routing fails?
What if routing fails?
The run continues on Gumloop’s baseline model for Auto instead of failing. Routing errors, timeouts, and content Auto can’t inspect all fall back to that baseline rather than escalating on a guess.
Should I use Auto or a preset?
Should I use Auto or a preset?
Use Auto for agents whose messages vary a lot in difficulty — it saves money on the easy ones without capping the hard ones. Pin a specific model or preset when you need every run to be predictable and identical, such as a compliance-sensitive agent or an eval baseline.
Reading the model card
Selecting or hovering a model opens a detail card so you can compare options at a glance:- Description: what the model is best at.
- Speed and Intelligence: relative ratings across the catalog.
- Provider and Context: who makes the model and how much it can read at once (in tokens and approximate words).
- Tool Calling and Vision: capability checks.
- Badges: extra labels that call out how a model is served. Models marked US-provider hosted are served by US providers under Zero Data Retention (ZDR) policies, and models marked 30-day data retention are not covered by ZDR.
For agents that read images or screenshots, pick one with Vision.
Available models
These are the models you can choose for an agent, grouped by provider. Every one supports tool calling, so it can use your apps and workflows. Models marked Vision can also read images and screenshots. New models are added continually, so the in-product picker is always the most up-to-date list.GPT-OSS 120B is exclusive to agents and is not available in workflow AI nodes. The reverse also happens: a few models don’t support tool calling, so they only appear in workflow AI nodes and never in the agent picker.
All open-source models (such as GPT-OSS 120B, LLaMA, DeepSeek, Qwen, Kimi, MiniMax, and GLM) are accessed through US-based providers under Zero Data Retention (ZDR) policies, and carry a US-provider hosted badge in the model picker. Your data is never used for model training and is not stored after inference.
Models with 30-day data retention
Gumloop serves every model under Zero Data Retention with one documented exception: the Claude Fable family. Anthropic keeps prompts and outputs sent to Claude Fable models for 30 days to check for misuse, then deletes them. Anthropic does not train on that data, and every other model — including every other Claude model — is unaffected.
These models carry a 30-day data retention badge in the picker, and:
- They are off by default. An organization admin has to explicitly allow non-zero data retention models before anyone can select one. See AI Model Governance.
- They are never chosen for you. They are excluded from the Recommended / Smartest presets, from Auto, and from fallback models — the only way one runs is if you pick it.
- If the acknowledgement is later revoked, they become unavailable immediately and agents set to them fail until you switch models.
How agents are charged
In agents, model cost is token-based and variable. You are charged for the tokens each message uses, which depends on the model, the length of the conversation, and the tools available. There are no fixed per-message tiers. Model calls bill at cost, converted to credits at $0.005 per credit, and each run also carries Compute and an Orchestration Fee. Select the credit count in any chat’s header to see its exact usage, and see Credits for the full breakdown.Bring your own key (BYOK)
Provide your own provider API key to cut model costs. BYOK makes AI model calls cost 0 credits in both agents and workflow AI nodes, and it also applies to image generation and voice transcription. Since the model spend leaves your credit balance entirely, agent chats on BYOK carry a 16% orchestration fee instead of 8%, calculated on what the run would have been worth — still a large net saving. See Credits.Each key only waives the models it can actually serve. A Fireworks AI key covers the open models Gumloop serves through Fireworks — DeepSeek V4 Flash, DeepSeek V4 Pro, Kimi K3, Kimi K2.7 Code, GLM-5.2, MiniMax M3, and Qwen3.8 Max. The other open models in the picker (GPT-OSS 120B, Gemma 4 26B, GLM-5.3-Flash) run elsewhere, so a Fireworks key does not affect them.
Setting up a key
Setting up a key
Requires a Pro plan or higher and your own OpenAI, Anthropic, Google AI, Perplexity, SpaceXAI, or Fireworks AI account.
- Pro
- Enterprise
Add a key under your personal credentials at Connectors so your own calls route through it, or add a shared team key so the whole team can use it without managing individual keys.Pro users cannot set keys at the organization level.
Using these models in workflows
The same models power workflow AI nodes, and they are billed the same way as agents: by token usage. Nodes that don’t need tool calling also offer a few models the agent picker leaves out.How workflow AI nodes are charged
How workflow AI nodes are charged
Workflow AI nodes (such as Ask AI, Analyze Image, and Generate Report) are billed by token usage, based on the model you pick and how many input and output tokens each call uses. There are no fixed per-call tiers, so a short prompt costs far less than a long-context one.
- Smaller, faster models cost less per token than frontier models.
- With BYOK, workflow AI node calls cost 0 credits.
- Image generation is billed at a flat 30 credits per image (free with BYOK), regardless of size, quality, or model. Audio transcription is billed by audio length at a small per-minute rate that depends on the model (roughly 1 to 2 credits per minute), and is also free with BYOK.
