Where AI learns to do real work
Labelbox builds the environments frontier labs train in and the platform enterprises run their agents on.
- Tracker0
- Cluster0
- Evals0
- GitHub0
- Notion0
- Slack0
- Docs0
For enterprises
Agents that do the work, and get better at it.
Recursion Managed Agents are a new way to use AI in the business. Set a goal; Recursion plans it, spawns as many agents as the problem needs, and works it in parallel across your systems, autonomously. Every run makes the next one better.
For frontier AI
The data, environments, and evaluation infrastructure the world's frontier AI labs build on.
Hundreds of AI teams build with Labelbox

How Meta built GIM with Labelbox data to evaluate frontier AI reasoning
Problem
Meta needed a benchmark that remained discriminative as existing LLM evaluations saturated. The team wanted tasks grounded in practical reasoning rather than obscure knowledge or synthetic puzzles, with enough rubric detail to capture partial credit and enough quality control to support a public-private contamination diagnostic.
Solution
Labelbox produced the data foundation for GIM: 820 expert-authored problems across seven cognitive categories, including 229 multimodal items and 528 rubric-graded prompts. The work included original prompt creation, structured scoring criteria, review, annotation, and quality assurance, enabling Meta to calibrate a 2PL IRT model over more than 200,000 prompt-response pairs.
Result
Meta released GIM-615, calibrated item parameters, and an evaluation framework that benchmarked 22 models across 47 reporting configurations. The paper found GIM remains far from saturated, with roughly 20% of items above frontier ability, giving researchers a durable way to compare model capability, thinking budgets, and future systems.

Human preference signal for evaluating LLMs inside Vertex AI

Higher-quality training signal for personalized shopping AI


Tracking surgical instruments in video to advance robotic surgery
Latest work from Labelbox Research
Labelbox's world-class applied research team pioneers frontier AI data generation and evaluation methods. Through scientific precision and co-innovation, we help customers achieve real-time AGI breakthroughs.