AI models evolve rapidly. New models are usually available in Gumloop within a day of their public release, so you have the latest options even if this page has not caught up yet.
Choosing a model
Open the model dropdown in Agent Preferences. The fastest way to choose is one of the three presets at the top:
- Recommended: the best balance of speed, quality, and cost. This is the default for new agents.
- Smartest: maximum intelligence for complex reasoning and agentic work.
- Fastest: optimized for speed and low latency on simple, high-volume tasks.
On Enterprise plans, admins can restrict which models a Custom Role may use. Models your role does not allow are hidden from this picker (and from workflow AI nodes), so you only see the models you are permitted to select.
Reading the model card
Selecting or hovering a model opens a detail card so you can compare options at a glance:- Description: what the model is best at.
- Speed and Intelligence: relative ratings across the catalog.
- Provider and Context: who makes the model and how much it can read at once (in tokens and approximate words).
- Tool Calling and Vision: capability checks.
- Badges: extra labels that call out how a model is served. Models marked US-provider hosted are served by US providers under Zero Data Retention (ZDR) policies.
For agents that read images or screenshots, pick one with Vision.
Available models
These are the models you can choose for an agent, grouped by provider. Every one supports tool calling, so it can use your apps and workflows. Models marked Vision can also read images and screenshots. New models are added continually, so the in-product picker is always the most up-to-date list.GPT-OSS 120B, Qwen3.5 397B, and Kimi K2.6 are exclusive to agents and are not available in workflow AI nodes.
All open-source models (such as GPT-OSS 120B, LLaMA, DeepSeek, Qwen, Kimi, MiniMax, and GLM) are accessed through US-based providers under Zero Data Retention (ZDR) policies, and carry a US-provider hosted badge in the model picker. Your data is never used for model training and is not stored after inference.
How agents are charged
In agents, model cost is token-based and variable. You are charged for the tokens each message uses, which depends on the model, the length of the conversation, and the tools available. There are no fixed per-message tiers. Model calls bill at cost, converted to credits at $0.005 per credit, and each run also carries Compute and an Orchestration Fee. Open Chat Details on any conversation to see its exact usage, and see Credits for the full breakdown.Bring your own key (BYOK)
Provide your own provider API key to cut model costs. BYOK makes AI model calls cost 0 credits in both agents and workflow AI nodes, and it also applies to image generation and voice transcription. Since the model spend leaves your credit balance entirely, agent chats on BYOK carry a 16% orchestration fee instead of 8%, calculated on what the run would have been worth β still a large net saving. See Credits.Each key only waives the models it can actually serve. A Fireworks AI key covers the open models Gumloop serves through Fireworks β DeepSeek V4 Flash, DeepSeek V4 Pro, Kimi K3, Kimi K2.7 Code, GLM-5.2, MiniMax M3, and Qwen3.7 Plus. The other open models in the picker (GPT-OSS 120B, Qwen3.5 397B, Kimi K2.6, Gemma 4 26B) run elsewhere, so a Fireworks key does not affect them.
Setting up a key
Setting up a key
Requires a Pro plan or higher and your own OpenAI, Anthropic, Google AI, Perplexity, SpaceXAI, or Fireworks AI account.
- Pro
- Enterprise
Add a key under your personal credentials at Connectors so your own calls route through it, or add a shared team key so the whole team can use it without managing individual keys.Pro users cannot set keys at the organization level.
Using these models in workflows
The same models power workflow AI nodes, and they are billed the same way as agents: by token usage.How workflow AI nodes are charged
How workflow AI nodes are charged
Workflow AI nodes (such as Ask AI, Analyze Image, and Generate Report) are billed by token usage, based on the model you pick and how many input and output tokens each call uses. There are no fixed per-call tiers, so a short prompt costs far less than a long-context one.
- Smaller, faster models cost less per token than frontier models.
- With BYOK, workflow AI node calls cost 0 credits.
- Image generation is billed at a flat 30 credits per image (free with BYOK), regardless of size, quality, or model. Audio transcription is billed by audio length at a small per-minute rate that depends on the model (roughly 1 to 2 credits per minute), and is also free with BYOK.
