Flaky Test Fixer
Reproduces each flaky test, finds the race and fixes it.
Manu Sharma•October 8, 2026
8 agents · 4 models. Every graded run makes the next one better.
Every company has work that keeps coming back. The eval suite that reruns on every new checkpoint. The flaky test nobody owns. The forecast due Monday morning. The brief a rep needs an hour before the call. None of it is hard to describe. All of it is hard to keep up with.
For the last two years, the answer has been AI assistants. Your people use them all day, for chat and for code, and they get more done for it. But every session still starts with someone at the keyboard, and the result still needs a person to check it and carry it where it belongs. When people log off, the work waits.
So we built something different, for ourselves first. At Labelbox, agents now start on their own, work inside the tools we already use, hand back finished work, and get better every run. In just a few months, every function in the company, from engineering and research to sales, finance, and operations, has put them to work.
Today, we’re opening that system to every team. It’s called Recursion Managed Agents.
A representative week. Assistants are busy whenever your people are. Managed agents start on schedules, events, and API calls, so the work keeps moving nights and weekends too.
We started Labelbox in 2018 around one observation: an AI system is only as good as the signal it learns from. It’s why 90%+ of leading US AI labs work with Labelbox today, and why we built Horizon, our RL environments for post-training frontier models.
Here’s what that work has taught us. Frontier models get better and cheaper every quarter, and every company gets the same ones. If everyone rents the same intelligence, the intelligence isn’t the advantage. What you do with it is: the expertise, workflows, and judgment only your company has, and a system that turns them into agents that keep getting better at your work.
That system is the learning loop, and we believe it’s the new IP of the firm. We launched Recursion in June so enterprises could build one. Managed Agents is how the loop starts: put agents to work, grade every run, and let the work itself make the next run better.
You don’t write a script for a Managed Agent. You give it a goal, what starts it, and what good looks like. When the work is done, you get the deliverable, the evidence behind it, and a grade against the bar you set.
What good looks like is a rubric. It’s a short checklist your team writes in plain language, the same one you’d hand a new hire. A grader agent scores every run against it, line by line, and points to the evidence for each call. When the work misses, the agent goes back and revises it. No rubric yet? Start with a plain definition of done and tighten it as you review the first runs.
A coordinator plans the work and brings in specialists as it needs them, so a bigger problem gets more agents, not a longer prompt.
Sessions pause on dependencies, recover from failures, and pick up exactly where they left off, for as long as the work takes.
Claude, GPT, Gemini, and open-weight models run side by side, each specialist on the model it does best.
A schedule, an alert, a Slack message, a webhook, or an API call starts the work. Nobody has to press go.
Not every job needs more than one agent, and the ones that do don’t all need the same shape. Every agent can work in three modes.
Multi-agent. A coordinator chains specialists. Each one is its own agent, with its own model, instructions, tools, and permissions, so the agent reading your billing data can stay read only while another one issues the refund.
Team. A swarm on one board. The lead agent splits the work, copies of it take the pieces in parallel and check each other’s results, and the lead makes one call. It fits work with independent pieces: three approaches to benchmark, five checks on a vendor, a dozen accounts to research.
Auto. The default. The agent decides for itself when a job needs specialists or a team.
In every mode, the agents share one sandbox and one record, a failure in one doesn’t stop the others, and the cost rolls up to a single number.
A coordinator chains specialists, each with its own model, tools, and permissions.
Agents work across 100+ business apps and any MCP server, limited to exactly the tools you allow. Every tool is scoped as read only, changes data, or destructive, and credentials stay in a vault, so the model never sees them. Slack, GitHub, and Jira are built in, so you can start an agent from a Slack message and get the result back where your team already works.
Credentials stay in the vault. The agent sees only the tools you turn on, and every call lands in the session transcript.
Connect once. Turn tools on per agent.
Where there’s no API, the agent gets its own browser. Every Monday at 07:00, an RFP Sweep agent signs in to six government bid portals, downloads the new bid documents, and adds each match to Salesforce with a go/no-go summary. Every click is saved as a screenshot you can replay.
Every session is a durable record: the agent version it ran, the environment it ran in, every model turn and tool call, the files it produced, what it cost, and how it was graded. You can follow it live, step in mid-run, or open it a month later for an audit.
Agent, environment and task
Files, skills, tools and credentials
Model turns and tool calls
Scored against your rubric
Back to work on what missed
Tied back to the session
When an agent is stuck, it asks. If it keeps failing the same way, the session stops for review instead of burning more hours, with the transcript and the grader’s notes attached. Someone fixes what blocked it and replies, and the run picks up where it left off. Work that still misses the bar after its allowed revisions comes back marked that way, never passed off as done.
Across the fleet, the week’s sessions sit on one board: what shipped, what’s still running, what’s blocked, and how each one scored.
Brighter rows mean better-graded work.
We hold our own agents to the bar we’d set for any customer. Here’s how we use them at Labelbox: 51 agents across 11 teams, from software engineering, reliability, and security to ML research, evals, product, finance, sales, strategy, and procurement. Each one starts on its own, works in the tools its team already uses, and hands back a finished result.
The surprise wasn’t that agents could do the work. It was how quickly they got better at it.
Same model, a third cheaper. Take the agent that reviews pull requests on the repositories behind our RL environments. In under two days, it ran 39 graded reviews on one agent version and one model, so its instructions never changed. What changed was its memory, updated 15 times along the way. Counting the 33 reviews that passed the grader, the median cost fell from $13.15 for the first 11 to $8.90 for the last 11, 32% less per review. Every pull request is different, so this isn’t a lab test, but the cost fell in each stretch.
It learned the job, not a new model. It learned the codebase: how the shared task framework sets its defaults and what its CI enforces. It learned its tools: which ways to authenticate git work in its sandbox, and that wide searches and whole-file dumps can burn a whole session without producing a review. And it learned the bar: after the grader rejected the same finding four times, it stopped flagging a default the framework sets on purpose. Time per review held at about 27 minutes. The savings came from model spend: it reached the same bar with less wasted work.
Recursion session records for our PR review agent, Sep 19–21, 2026. Medians over the 33 reviews that passed the grader, on one agent version and one model.
This is what makes Managed Agents more than orchestration. A chat agent starts every task from zero. A Managed Agent starts from everything it has already learned doing your work.
Every run teaches the agent. Each session is graded against your rubric and leaves a short report on what worked, what failed, and what the grader caught. The lessons go into a memory of the agent’s own that it brings to every run after. When it brings in specialists or a team for a job, the learning still belongs to the agent that owns the job.
Better shows up on the bill. An agent that remembers where the answer usually is wastes fewer steps finding it. Fewer wasted steps mean fewer tool calls and fewer tokens, and since one bill covers compute and model spend, the same job costs less every week it runs. No model is retrained along the way.
Faster only counts if it’s still right. Every run is graded against the same rubric, so cost and quality sit side by side in the record. A shortcut that costs quality shows up in the next grade.
Memory you can inspect. Lessons don’t pile up unchecked. Once five new session reports are in, and at most once an hour, a curator agent reads them and updates the agent’s memory, following standing instructions your team can write, like what to keep and what to ignore. You can browse everything an agent has learned, see the last 30 days of changes, and delete any lesson that no longer holds. Reference material you want every run to use goes in a separate store your team edits directly, and memory can be turned off for any agent.
Subagents do pieces of the job. The agent that owns it learns from every graded session and starts the next one from its own memory.
Every run captures a little more of how your company does the work. That knowledge compounds, and it stays in your workspace.
None of this matters if your security team can’t approve it. Every session runs in an isolated sandbox with only the tools and credentials you grant. Credentials stay in a vault and are attached to approved calls as short-lived tokens, so the model never sees them, and a call to a host off the allowlist is blocked, not attempted. Sensitive actions wait for a person to approve, and spend, token, and time limits pause a run for review.
Merge PR #5120 and serve the int8 model to a 5% canary
approved by the on-call
Recursion runs with zero data retention at every LLM provider, so model providers never keep your prompts or outputs. Every model turn and tool call lands in an audit-ready transcript in your workspace, backed by Labelbox’s security and compliance program: SOC 2 Type II, GDPR, CCPA, NIST 800-171, and PCI DSS.
Pay by the agent session hour, with compute and model spend in one estimate you see before you launch. No seats, plans, or subscriptions.
Existing Labelbox enterprise customers get a one-time $100 credit and can apply additional spend to their committed Labelbox spend, or pay by card instead if they prefer. See pricing.
Within a few years, we think every company will run a fleet of agents. That part feels inevitable. What isn’t inevitable is who gets better. Companies that rent intelligence will get the same results as everyone else. Companies that own their learning loop will compound: every run sharpens the next, every agent gets better at its job, and the gap widens every quarter.
That’s what we’re building Labelbox to do: turn the work a company already does into intelligence it owns. Horizon does it for the labs that build frontier models. Recursion does it for the companies that put them to work.
Most enterprises already use AI to answer questions, and chat has its place. But running our own agents, we’ve found that most of the opportunity is in the work behind those questions: the escalations, vendor reviews, incident write-ups, and account research that come back every week, however many people you hire.
That work needs agents that don’t wait to be asked. They start on their own, run for days without anyone watching, and fan out to thousands of sessions at once when the job is big, like every contract up for renewal or every flaky test in the suite. They’re graded against the bar your best people hold, and they come back with finished work, not a draft for someone to redo.
Pick one of those jobs. We’ll help you set your first agents running on it inside your apps, under controls your security team signs off on, and everything they learn stays in your workspace. An agent that starts this quarter will know things next quarter that a new one would have to learn from scratch.
For a limited time, new teams get $20 in free credits to get started.