Feature Builder
Builds each ready ticket into a working feature, with tests, in its own sandbox.
Labelbox builds the environments frontier labs use to train AI and the platform enterprises use to put agents to work.
Agents for the work that keeps coming back
Recursion isn’t a chatbot. Its agents start on their own, on a schedule, an alert or an event, and work in the background until the job is done: every ready ticket built, every new bug fixed, every dependency kept current, dozens of experiments run at once. Here are a few of the jobs teams hand it.
Describe the job, what starts it and what a good result looks like. From then on Recursion runs it in the background: it plans the work, brings in as many specialist agents as it needs, runs them in parallel and delivers the result into your apps, with the evidence attached. Models, sandboxes, credentials, memory and grading are handled for you, so most teams use it out of the box.
48 configs ran in parallel overnight. Config 31 beats the baseline by 2.1 points, confirmed on a rerun.
Every run lands on one fleet board.
Gets better, and cheaper, the more it works.
Every graded run leaves notes and skills the next run starts from. When a job runs often enough, Recursion turns its graded runs into training environments, the kind Labelbox builds for frontier labs, trains a specialist model on them and switches over once it beats the model you use today.
5–10×lower cost per task once a specialist takes over
The data and environments frontier models learn from
Over 90% of leading US AI labs post-train and evaluate on Labelbox.

Problem
Meta needed a benchmark that remained discriminative as existing LLM evaluations saturated. The team wanted tasks grounded in practical reasoning rather than obscure knowledge or synthetic puzzles, with enough rubric detail to capture partial credit and enough quality control to support a public-private contamination diagnostic.
Solution
Labelbox produced the data foundation for GIM: 820 expert-authored problems across seven cognitive categories, including 229 multimodal items and 528 rubric-graded prompts. The work included original prompt creation, structured scoring criteria, review, expert feedback, and quality assurance, enabling Meta to calibrate a 2PL IRT model over more than 200,000 prompt-response pairs.
Result
Meta released GIM-615, calibrated item parameters, and an evaluation framework that benchmarked 22 models across 47 reporting configurations. The paper found GIM remains far from saturated, with roughly 20% of items above frontier ability, giving researchers a durable way to compare model capability, thinking budgets, and future systems.




Our applied research team publishes the benchmarks and methods we use to train and evaluate frontier models.
Each agent gets only the access its job needs, in a workspace of its own. The model never sees your passwords or keys, and every step is recorded.
Your teams already ask AI for answers. Recursion takes on whole jobs: it starts on its own, works across your apps and hands back finished, graded work. Bring the job that eats your experts’ week, and we’ll set it running.
Frontier AI labs already train on Labelbox environments and expert data. Tell us where your models fall short, and we’ll scope the environments, data and evals.