×

Manu Sharma•October 8, 2026

Recursion Managed Agents: AI agents that get better every run

Recursion Managed Agents· Sessions
Recursion Managed Agents· Sessionsincident INC-4821
INC-4821 · SEV-2 · checkout-api · 5xx at 6× baseline since 14:07
Incident: checkout is failing for some customers after the 14:05 deploy
Find the root cause, stop the errors, ship a tested fix, count the customer impact and keep the status page current.
running
8 agents · 4 models
outcome satisfied 5/5 · mitigated in 26 min · fix in review
  • PagerDuty0
  • Datadog0
  • GitHub0
  • LaunchDarkly0
  • Stripe0
  • Zendesk0
  • Statuspage0
  • Slack0
Incident coordinatorClaudecoordinatorrunningLogs and tracesGeminispecialistwaitingDeploy diffopen weightsspecialistwaitingFeature flagsGPTspecialistwaitingReproduceClaudespecialistwaitingCustomer impactGPTspecialistwaitingFix authorClaudespecialistwaitingOutcome graderClaudegraderwaiting
coordinator · PagerDuty INC-4821 acknowledged · checkout-api 5xx at 6× baseline since 14:07 · plan: 5 checks
Checkout errors spike after a deploy.
PagerDuty pages the on-call: checkout is failing at six times the normal rate since 14:07, two minutes after the 14:05 deploy. The coordinator reads the alert, the error graphs and the deploy, then plans five checks.

8 agents · 4 models. Every graded run makes the next one better.

Every company has work that keeps coming back. The eval suite that reruns on every new checkpoint. The flaky test nobody owns. The forecast due Monday morning. The brief a rep needs an hour before the call. None of it is hard to describe. All of it is hard to keep up with.

For the last two years, the answer has been AI assistants. Your people use them all day, for chat and for code, and they get more done for it. But every session still starts with someone at the keyboard, and the result still needs a person to check it and carry it where it belongs. When people log off, the work waits.

So we built something different, for ourselves first. At Labelbox, agents now start on their own, work inside the tools we already use, hand back finished work, and get better every run. In just a few months, every function in the company, from engineering and research to sales, finance, and operations, has put them to work.

Today, we’re opening that system to every team. It’s called Recursion Managed Agents.

Chat and coding assistantsworking hours
A person starts and steers every session
Recursion Managed Agentsall 168 hours
SchedulesEventsAPI calls
One week of AI work. Chat and coding assistants are busy through every working day, but a person starts and steers every session, so they go quiet at night and on the weekend. With Recursion Managed Agents, agents start on schedules, events and API calls and work all 168 hours of the week, many sessions at once.

A representative week. Assistants are busy whenever your people are. Managed agents start on schedules, events, and API calls, so the work keeps moving nights and weekends too.

Why learning agents are a business imperative

We started Labelbox in 2018 around one observation: an AI system is only as good as the signal it learns from. It’s why 90%+ of leading US AI labs work with Labelbox today, and why we built Horizon, our RL environments for post-training frontier models.

Here’s what that work has taught us. Frontier models get better and cheaper every quarter, and every company gets the same ones. If everyone rents the same intelligence, the intelligence isn’t the advantage. What you do with it is: the expertise, workflows, and judgment only your company has, and a system that turns them into agents that keep getting better at your work.

That system is the learning loop, and we believe it’s the new IP of the firm. We launched Recursion in June so enterprises could build one. Managed Agents is how the loop starts: put agents to work, grade every run, and let the work itself make the next run better.

One goal in. A fleet of specialists out.

You don’t write a script for a Managed Agent. You give it a goal, what starts it, and what good looks like. When the work is done, you get the deliverable, the evidence behind it, and a grade against the bar you set.

What good looks like is a rubric. It’s a short checklist your team writes in plain language, the same one you’d hand a new hire. A grader agent scores every run against it, line by line, and points to the evidence for each call. When the work misses, the agent goes back and revises it. No rubric yet? Start with a plain definition of done and tighten it as you review the first runs.

Multi-agent by design

A coordinator plans the work and brings in specialists as it needs them, so a bigger problem gets more agents, not a longer prompt.

Built to run for days or weeks

Sessions pause on dependencies, recover from failures, and pick up exactly where they left off, for as long as the work takes.

Any model, per agent

Claude, GPT, Gemini, and open-weight models run side by side, each specialist on the model it does best.

Starts on its own

A schedule, an alert, a Slack message, a webhook, or an API call starts the work. Nobody has to press go.

Recursion Managed Agents· Sessions
Recursion Managed Agents· Sessionscoordinator · 6 specialists · 1 grader
Inference coordinatorClaude
JiraENG-2291
Cut the ranking model’s p95 latency by 30% without losing accuracy.
planreading ENG-2291 and the code
Profiler
Quantization
Batching
Distillation
Load test
Accuracy guard
recommendation
Ship int8 with a 4 ms batch window: p95 −38%, accuracy −0.1.
ProfilerGPTDatadog
not started
QuantizationClaudeGitHubSandbox
not started
BatchingOpen weightsGitHubSandbox
not started
DistillationGeminiGoogle Cloud
not started
Load testGPTSandboxGrafana
not started
Accuracy guardClaudeGoogle BigQuerySandbox
not started
Outcome graderClaude
your rubricwaiting for a recommendation
Bottleneck found in production
evidence: Datadog traces, 10k requests
Change is tested and reviewed
evidence: PR #5120, CI green
No new hardware
evidence: same GPU pool, 2.3× throughput
p95 down at least 30%
evidence: load test at 2× peak traffic
Accuracy within 0.2 points
evidence: 50k held-out queries, 3 seeds
satisfied 5/5ship to a 5% canary

Three ways agents work together

Not every job needs more than one agent, and the ones that do don’t all need the same shape. Every agent can work in three modes.

  • Multi-agent. A coordinator chains specialists. Each one is its own agent, with its own model, instructions, tools, and permissions, so the agent reading your billing data can stay read only while another one issues the refund.

  • Team. A swarm on one board. The lead agent splits the work, copies of it take the pieces in parallel and check each other’s results, and the lead makes one call. It fits work with independent pieces: three approaches to benchmark, five checks on a vendor, a dozen accounts to research.

  • Auto. The default. The agent decides for itself when a job needs specialists or a team.

In every mode, the agents share one sandbox and one record, a failure in one doesn’t stop the others, and the cost rolls up to a single number.

A coordinator chains specialists, each with its own model, tools, and permissions.

Three ways agents work together. Multi-agent: a Support Escalation coordinator chains three specialists, an Order Investigator with read-only access to Stripe, a Refund Issuer that can change data in Stripe after approval, and a Customer Reply agent in Zendesk, each on its own model. Team: copies of a Catalog API agent claim tasks on a shared board to prototype three caching approaches, one teammate checks another's numbers, and the leader records the decision. Auto: a Vendor Risk Review agent answers a quick question alone, then forms a team of five for a full vendor review.

Works in the apps your business already runs on

Agents work across 100+ business apps and any MCP server, limited to exactly the tools you allow. Every tool is scoped as read only, changes data, or destructive, and credentials stay in a vault, so the model never sees them. Slack, GitHub, and Jira are built in, so you can start an agent from a Slack message and get the result back where your team already works.

Recursion Managed Agents
GitHub
GitHub
Native integration · vaulted credential
Tools this agent can call4 of 5 on
  • get_pull_request
    read only
  • list_check_runs
    read only
  • create_branch
    changes data
  • create_pull_request
    changes data
  • merge_pull_request
    destructive

Credentials stay in the vault. The agent sees only the tools you turn on, and every call lands in the session transcript.

Connect once. Turn tools on per agent.

Where there’s no API, the agent gets its own browser. Every Monday at 07:00, an RFP Sweep agent signs in to six government bid portals, downloads the new bid documents, and adds each match to Salesforce with a go/no-go summary. Every click is saved as a screenshot you can replay.

Recursion Managed Agents· Sessions
Recursion Managed Agents· SessionsRFP Sweep · scheduled run · computer use
New tab
New tabSearch or type a URLAgent driving
Agent
Computer · every action saved with a screenshot0 frames
Live
RFP Sweepsession 7318running
every Monday 07:00computer useSalesforceSalesforce
Agent narration
Monday 07:00 run. Checking 6 bid portals for new RFPs that fit our fleet software. None has an API, so I’ll sign in to each one.
Runsevery Monday 07:00
Sep 21
0 new
Sep 28
1 new
now · Oct 5
running
next · Oct 12
07:00
Every Monday at 07:00, with no one watching, the RFP Sweep agent signs in to six government bid portals with the team's vendor login, searches for fleet telematics RFPs, downloads the bid documents and adds each match to Salesforce with a go/no-go summary. The run finds three new RFPs, skips one already tracked, checks all six portals and schedules the next run for the following Monday.

Every run, on the record

Every session is a durable record: the agent version it ran, the environment it ran in, every model turn and tool call, the files it produced, what it cost, and how it was graded. You can follow it live, step in mid-run, or open it a month later for an audit.

  1. Start

    Agent, environment and task

  2. Provision sandbox

    Files, skills, tools and credentials

  3. Agent works

    Model turns and tool calls

  4. Grade

    Scored against your rubric

  5. Revise if failing

    Back to work on what missed

  6. Deliver artifacts

    Tied back to the session

Session events0 events
    A session starts with an agent, an environment and a task. Recursion Managed Agents provisions the sandbox, the agent works through model turns and tool calls, and a grader scores the result against the rubric. The first grade misses a criterion, so the agent revises and the second grade passes. The artifact is delivered and every step is saved as an event in the session.

    When an agent is stuck, it asks. If it keeps failing the same way, the session stops for review instead of burning more hours, with the transcript and the grader’s notes attached. Someone fixes what blocked it and replies, and the run picks up where it left off. Work that still misses the bar after its allowed revisions comes back marked that way, never passed off as done.

    Across the fleet, the week’s sessions sit on one board: what shipped, what’s still running, what’s blocked, and how each one scored.

    Recursion Managed Agents· Sessions
    Recursion Managed Agents· Sessions144 sessions this week · 127 completed · 10 running
    • Incident Investigation
    • Flaky Test Fixer
    • Eval Regression Watch
    • Experiment Sweep
    • brighter = better graded

    Brighter rows mean better-graded work.

    How we use it at Labelbox

    We hold our own agents to the bar we’d set for any customer. Here’s how we use them at Labelbox: 51 agents across 11 teams, from software engineering, reliability, and security to ML research, evals, product, finance, sales, strategy, and procurement. Each one starts on its own, works in the tools its team already uses, and hands back a finished result.

    • CI failure on mainGitHubJenkinsSlack
      Delivers: A pull request with the fix and 50 green runs in a row
      Software engineering

      Flaky Test Fixer

      Reproduces each flaky test, finds the race and fixes it.

    • PagerDuty alertPagerDutyDatadogSentryGitHub
      Delivers: A root-cause draft and the suspect commit before on-call is up
      Reliability & infra

      Incident Investigation

      Correlates alerts, logs and recent deploys the moment an alert fires.

    • New pull requestGitHubJira
      Delivers: Real vulnerabilities only, each with proof and the fix
      Security engineering

      Code Security Review

      Reviews every pull request and reproduces what it finds.

    • Issue labeled sweepLinearGitHubDatabricks
      Delivers: A results table with the winning config, rerun to confirm
      ML research

      Experiment Sweep

      Runs dozens of experiment configs at once and compares every run with the baseline.

    • Pull request touching promptsGitHubGoogle BigQuerySlack
      Delivers: A per-task score diff, and a blocked merge if anything regresses
      Evals

      Eval Regression Watch

      Reruns your eval suite on every model or prompt change.

    • Request tagged for the roadmapLinearGongNotion
      Delivers: A draft spec with the problem, the customer evidence and open questions
      Product

      Spec Writer

      Writes the first spec for each request that makes the roadmap, from the calls and tickets behind it.

    • Mondays 06:00SalesforceSnowflakeGoogle Drive
      Delivers: An updated forecast with what moved since last week, and why
      Finance

      Forecast Refresh

      Rebuilds the revenue forecast from pipeline, usage and bookings.

    • Meeting on the calendarSalesforceGongGoogle Calendar
      Delivers: A one-page brief in the rep’s inbox an hour before every call
      Sales

      Account Research

      Builds a brief from the CRM, past calls, filings and news.

    • Mondays 08:00WebGongNotionSlack
      Delivers: A weekly brief on competitor moves and what they mean for open deals
      Strategy

      Competitive Intelligence

      Monitors competitor sites, pricing pages, job posts and mentions in your sales calls.

    • API call from the procurement portalSalesforceDocuSignGoogle Drive
      Delivers: Approve, approve with conditions or reject, with evidence for every check
      Procurement

      Vendor Risk Review

      Runs sanctions, legal, security, financial and reference checks in parallel.

    • A scheduleAn alertAn eventAn API call
      And many more

      Any job that keeps coming back

      These are a few of the jobs teams run. Describe yours, what starts it and what good looks like, and an agent takes it from there.

    The surprise wasn’t that agents could do the work. It was how quickly they got better at it.

    Same model, a third cheaper. Take the agent that reviews pull requests on the repositories behind our RL environments. In under two days, it ran 39 graded reviews on one agent version and one model, so its instructions never changed. What changed was its memory, updated 15 times along the way. Counting the 33 reviews that passed the grader, the median cost fell from $13.15 for the first 11 to $8.90 for the last 11, 32% less per review. Every pull request is different, so this isn’t a lab test, but the cost fell in each stretch.

    It learned the job, not a new model. It learned the codebase: how the shared task framework sets its defaults and what its CI enforces. It learned its tools: which ways to authenticate git work in its sandbox, and that wide searches and whole-file dumps can burn a whole session without producing a review. And it learned the bar: after the grader rejected the same finding four times, it stopped flagging a default the framework sets on purpose. Time per review held at about 27 minutes. The savings came from model spend: it reached the same bar with less wasted work.

    PR Review· one agent version · one model
    Median cost per review that passed the grader, in run order
    median cost per review
    −32%
    median cost per review
    memory updates along the way
    15
    memory updates along the way
    per review, about the same throughout
    27 min
    per review, about the same throughout
    From its learned memory
    • /workflows/pr-review/repo-harness-facts.mdhow the shared task framework sets its defaults
    • /environment/sandbox-tooling.mdwhich ways to authenticate git work in its sandbox
    • /workflows/pr-review/session-budget-failures.mdwide searches that burned whole sessions
    • /workflows/pr-review/validator-pitfalls.mdwhat the result checker rejects, and why
    Our pull-request review agent ran 39 graded reviews on one agent version and one model. Median cost per review that passed the grader fell from $13.15 over the first 11 reviews to $11.34 over the next 11 and $8.90 over the last 11, 32% less, while time per review stayed near 27 minutes. Its memory was updated 15 times along the way, with files on the shared task framework, sandbox tooling, wasted searches and result-checker pitfalls.

    Recursion session records for our PR review agent, Sep 19–21, 2026. Medians over the 33 reviews that passed the grader, on one agent version and one model.

    Gets better every run

    This is what makes Managed Agents more than orchestration. A chat agent starts every task from zero. A Managed Agent starts from everything it has already learned doing your work.

    Every run teaches the agent. Each session is graded against your rubric and leaves a short report on what worked, what failed, and what the grader caught. The lessons go into a memory of the agent’s own that it brings to every run after. When it brings in specialists or a team for a job, the learning still belongs to the agent that owns the job.

    Better shows up on the bill. An agent that remembers where the answer usually is wastes fewer steps finding it. Fewer wasted steps mean fewer tool calls and fewer tokens, and since one bill covers compute and model spend, the same job costs less every week it runs. No model is retrained along the way.

    Faster only counts if it’s still right. Every run is graded against the same rubric, so cost and quality sit side by side in the record. A shortcut that costs quality shows up in the next grade.

    Memory you can inspect. Lessons don’t pile up unchecked. Once five new session reports are in, and at most once an hour, a curator agent reads them and updates the agent’s memory, following standing instructions your team can write, like what to keep and what to ignore. You can browse everything an agent has learned, see the last 30 days of changes, and delete any lesson that no longer holds. Reference material you want every run to use goes in a separate store your team edits directly, and memory can be turned off for any agent.

    Incident Response · learning from its own sessions
    Incident Response learning from its own sessions. For each incident it hands pieces of the work to three subagents, a Log Analyst, a Deploy Checker and a Status Writer, and the graded session goes to reflection. Reflection draws a lesson and saves it to Incident Response's own memory, which no other agent reads: check flag changes from the last hour first; count failed checkouts in Stripe, not error logs; test tax changes with orders to Hong Kong and the UAE. Each new session starts from that memory.

    Subagents do pieces of the job. The agent that owns it learns from every graded session and starts the next one from its own memory.

    Every run captures a little more of how your company does the work. That knowledge compounds, and it stays in your workspace.

    Autonomy your security team can sign off on

    None of this matters if your security team can’t approve it. Every session runs in an isolated sandbox with only the tools and credentials you grant. Credentials stay in a vault and are attached to approved calls as short-lived tokens, so the model never sees them, and a call to a host off the allowlist is blocked, not attempted. Sensitive actions wait for a person to approve, and spend, token, and time limits pause a run for review.

    An inference-optimization agent runs in an isolated sandbox with no credentials inside. Each tool call goes through Recursion Managed Agents, which attaches a short-lived token from the vault to approved calls to Datadog (read only) and GitHub (changes data), blocks a request to a host that is not on the allowlist, and records every call in the session transcript.
    Inference Optimizer· session controls
    Approval gatewaiting on a person
    github.merge_pull_requestchanges data

    Merge PR #5120 and serve the int8 model to a 5% canary

    approved by the on-call

    Limitspauses for review at any limit
    Spend
    $0.00 of $40.00
    Time
    0h 00m of 4h 00m
    Tokens
    0.0M of 10.0M
    An inference-optimization run asks to merge a GitHub pull request, a write action, and waits at an approval gate until the on-call engineer approves it. Spend, time and token use climb toward their limits; the run pauses for review if it reaches any of them.

    Recursion runs with zero data retention at every LLM provider, so model providers never keep your prompts or outputs. Every model turn and tool call lands in an audit-ready transcript in your workspace, backed by Labelbox’s security and compliance program: SOC 2 Type II, GDPR, CCPA, NIST 800-171, and PCI DSS.

    Security and governance

    • Zero data retention with every LLM provider
    • Credentials vaulted, never reachable from agent code
    • Every session in an isolated sandbox
    • Immutable agent versions
    • Calls limited to allowlisted hosts
    • Audit-ready session logs
    • SOC 2 Type II
    • GDPR
    • CCPA
    • NIST 800-171
    • PCI DSS
    Inference Optimizer · session 2291
    An inference-optimization session sends each model turn to an LLM provider with zero data retention, so nothing is kept after the response. The agent runs in an isolated sandbox with no credentials inside; approved calls to Datadog and GitHub get a short-lived token from the vault, a request to a host off the allowlist is blocked, and every call is written to the audit log in your Recursion workspace.

    Pricing

    Pay by the agent session hour, with compute and model spend in one estimate you see before you launch. No seats, plans, or subscriptions.

    Existing Labelbox enterprise customers get a one-time $100 credit and can apply additional spend to their committed Labelbox spend, or pay by card instead if they prefer. See pricing.

    Where this goes

    Within a few years, we think every company will run a fleet of agents. That part feels inevitable. What isn’t inevitable is who gets better. Companies that rent intelligence will get the same results as everyone else. Companies that own their learning loop will compound: every run sharpens the next, every agent gets better at its job, and the gap widens every quarter.

    That’s what we’re building Labelbox to do: turn the work a company already does into intelligence it owns. Horizon does it for the labs that build frontier models. Recursion does it for the companies that put them to work.

    Bring us the work

    Most enterprises already use AI to answer questions, and chat has its place. But running our own agents, we’ve found that most of the opportunity is in the work behind those questions: the escalations, vendor reviews, incident write-ups, and account research that come back every week, however many people you hire.

    That work needs agents that don’t wait to be asked. They start on their own, run for days without anyone watching, and fan out to thousands of sessions at once when the job is big, like every contract up for renewal or every flaky test in the suite. They’re graded against the bar your best people hold, and they come back with finished work, not a draft for someone to redo.

    Pick one of those jobs. We’ll help you set your first agents running on it inside your apps, under controls your security team signs off on, and everything they learn stays in your workspace. An agent that starts this quarter will know things next quarter that a new one would have to learn from scratch.

    For a limited time, new teams get $20 in free credits to get started.