Robots learn from what people do. We turn it into training data at scale.

Terra is Labelbox’s data engine for robotics: egocentric and teleoperated demonstrations captured on purpose-built rigs, annotated and validated by an AI diversity engine, and delivered from pre-training through evals.

  • 1PB+
    produced as of January 2026
  • 7
    capture rig configurations
  • Up to 4
    synced camera streams per episode
  • Pre-training → evals
    one pipeline, one episode format
Why this is hard

Physical AI is blocked by data, not models

Language models had the internet. Robots have nothing of the kind: every demonstration has to be recorded on purpose, in the real world, by a person. Whoever solves that at scale, with the variety the world actually has, sets the pace for the field.

  1. 01

    Scale

    One task needs thousands of demonstrations to generalize. No lab can record that in-house.

  2. 02

    Diversity

    Ten thousand episodes in one kitchen is one episode. Coverage across environments, objects, operators and lighting is the whole point.

  3. 03

    Ground truth

    Raw video is not training data. Actions, poses, contacts and success labels have to be right, or the model learns the noise.

What ships

One pipeline, pre-training to evals

Every rig writes the same episode format, so the data that pre-trains a model is the same shape as the demonstrations that post-train it and the episodes that evaluate it.

01

Pre-training

Diverse egocentric and third-person video across environments, objects and tasks, with auto-labels for scene, object and action segments.

ego monoego stereoego trio
02

Post-training

Expert teleoperation and instrumented-gripper demonstrations with synchronized action streams and success labels, for imitation learning and RL fine-tuning.

teleoptrio + gripperego + wrist
03

Evals

Standardized manipulation and navigation tasks with scoring rubrics, run against your policy and reported per environment.

overheadprotocolper-environment report
Capture rigs

Seven ways to see a task

Every rig is hardware-synchronized, ships to the operator pre-configured, and writes to the same episode format. Pick the viewpoints your policy needs and mix them in one program.

RigCamerasModalityBest for
Ego mono1 head-mountedRGBScene understanding, pre-training
Ego stereo2 head-mounted, syncedRGB + depth3D scene, navigation
Ego trio3 head + peripheralRGB wide FOVFull-context manipulation
Ego + wristHead + 1–2 wristRGB close-upGrasping, fine manipulation
Trio + gripper3 head + instrumented gripperRGB + gripper stateAction-labelled demonstrations
Overhead1–2 fixed above the workspaceRGBBird’s-eye view, bimanual tracking
TeleoperationRobot-mounted + operatorJoint states, actions, videoPost-training, imitation learning
How it works

An engine that knows what it has not seen yet

Every episode is ingested, auto-categorized by environment, object, task and operator, validated, and counted against a target distribution. Where coverage is thin, collection is steered there next.

  1. 1Ingest · Raw streams land from operators worldwide
  2. 2Categorize · Environment, objects, task and operator style auto-tagged
  3. 3Validate · Automated checks, human review on edge cases
  4. 4Steer · Gaps found, collection redirected to fill them
Sample episodes

Five episodes from the corpus

Real footage from Terra rigs across tasks, sensors and environments. Bimanual and gripper-instrumented episodes first; egocentric captures follow.

Bimanual stationary

Stacking plates

Human egocentric: Trio + Gripper

Packing garments

Human egocentric: Mono

Washing dishes

Human egocentric: Stereo

Groceries

Human egocentric: Trio

Folding clothes

Inside the lab

Dwarkesh visits the San Francisco robotics lab

Where rigs are built, operators are trained and episodes are validated before they ship. A first-hand look at how Terra data gets made.

FAQ

Questions robotics teams ask us

Does Labelbox build robots?

No. Terra produces the data robots learn from and the evals that measure them. We are embodiment-agnostic and work with your hardware or ours.

What do you deliver?

Synchronized multi-camera video, action and state streams, annotations and success labels, in the formats your training stack expects.

Who records the data?

Collection runs through the Alignerr network and specialist third-party partners, on rigs we supply, in homes, workplaces and our San Francisco robotics lab. Labelbox specializes in what happens next: turning that raw footage into validated, training-ready datasets with our technology.

Can we commission a task?

Yes. Most programs start with a task list and a target distribution; the diversity engine steers collection until coverage is met.

How is quality checked?

Every episode is auto-categorized and validated, with human review on edge cases. Episodes that fail are re-recorded, not shipped.

Tell us what your robots need to learn

Share the tasks and environments you need covered. We will scope the program and show you the first episodes.