Skip to main content
Performance in the agent sidebar answers three questions about one agent: what is it costing, is it doing good work, and what has it actually been doing. Those are the three tabs. This page is the map of that screen. Two of the things you switch on from it — Evaluations and Reflections — have their own guides, linked below.
Performance page with the Insights tab showing Credits Over Time, Credits by, Usage by, overview metrics, impact report, Evaluations, and Reflections
Performance covers the tasks you are allowed to see. Owners see everything on the agent; a User’s numbers follow the agent’s Task Visibility. For organization-wide reporting, see Insights.

Insights

Pick a date range, and optionally Compare to previous period to see the change. The chart’s interval is chosen for you to suit the range, and you can switch to any other interval that still makes a readable chart:
Intervals that would produce a single bar or thousands of them are not offered, so a short range cannot be charted by month and a year cannot be charted by the hour.
Credits Over Time charts three series together — credits, tasks, and cost per task — so you can tell a spike in spend apart from a spike in usage. The overview on the right gives you the headline numbers for the range: Under the chart, two breakdown tables:
  • Credits by — Models, Sources (where tasks came from, such as a trigger, Slack, or the app), Members, and Connectors.
  • Usage by — tool calls and reads per Connector, Skill, Artifact, Subagent, and Knowledge Source.
Reading it. High credits with few tasks usually means long tasks on an expensive model: check Credits by Models and consider a cheaper model or a lower summarization trigger. Lots of tasks with almost no connector usage often means people are asking the agent things it has no tools for.

Task outcomes

Under the breakdowns, Task outcomes charts how tasks ended over time, on the same range and interval as the rest of the tab:
Read it next to Credits Over Time. Cost falling while Didn’t get done rises is a cheaper model doing worse work, not a saving. A steady band of Needed a human usually points at a missing tool or an instruction gap for one specific kind of request.

What changed, and what it did to cost

Markers along the bottom of Credits Over Time show when the agent was changed, so you can line a change up against the spend and volume around it. Each marker covers one interval of the chart; select one to see the changes inside it, newest first, with who made each one:
  • Agent created
  • Model changed — the model it moved to, and the one it came from
  • Instructions edited
  • Connectors, skills, and knowledge sources added or removed
  • Reflections the agent applied — select one to open the suggestion
When enough data exists on both sides, the card also shows how cost per task moved after the change, so a model swap’s effect is visible without doing the arithmetic. It appears only once both the before and after windows have completed, and it stays quiet when the movement is within the agent’s normal noise.
A change made before the agent ran any tasks is marked No tasks yet, since there is nothing yet to compare against.

Impact report

Have Gumloop send this summary out on a schedule, by Email or to a Channel, to the people you add under To. Useful for keeping a team lead in the loop without giving them the builder.

Evaluations and Reflections

Both features live in the summary panel on the right of this page, and both are switched on from here: Everything else about them — criteria, tags, data points, schedules, reports, and the API — is in those two guides.

Evaluations tab

Once evaluations are on, this tab shows the grades: how completed tasks scored against your criteria, which ones failed, and the trend over time. That is where you catch a regression after changing the instructions or the model.

Tasks tab

The full task history for the agent, filterable, with who started each one, where it came from, and what it cost. Open a task to read the whole thing. This is the fastest way to answer “what did it actually do yesterday?” and to spot the pattern behind a bad evaluation score.

Evaluations

Define criteria and grade the agent’s answers.

Reflections

Let the agent propose its own improvements.

Credits

What the agent is charged for.

Organization Insights

The same picture across every agent and team.