See what great looks like with Labelbox
Explore success stories from a range of verticals and real-life use cases.

Human preference signal for evaluating LLMs inside Vertex AI
Problem
As LLMs grow more sophisticated, accurately evaluating their performance becomes critical. Automated metrics give insight, but human judgment remains the standard for nuances like relevance, bias, and overall quality. Producing large-scale, high-quality human preference signal is the hard part for most enterprises — it takes time, resources, and domain expertise.
Solution
After seeing the impact of using Labelbox internally, Google Cloud selected it to deliver LLM evaluation as a managed solution inside Vertex AI. Customers launch a human evaluation job and set criteria (question-answer, multi-turn chat, summarization), and Labelbox's platform produces expert-graded preference signal across customizable dimensions like instruction following, verbosity, and relevance.
Result
Customers develop and ship LLM applications with confidence. They receive quality-reviewed evaluation results within days, and launch evaluation jobs in minutes.

How Meta built GIM with Labelbox data to evaluate frontier AI reasoning

Benchmarking agentic models on 1,000+ real-world tool-use tasks

Expert-graded audio signal for emotion and speech-style models

Hardening an LLM's STEM reasoning with expert multimodal signal

Encoding legal judgment into a specialist AI agent

Expert preference signal for a financial-reasoning frontier model


Daily evaluation signal that retrains Speak's speech models

Expert speech signal for voice-preserving translation models

RLHF preference signal for a text-to-image reward model


Curating training signal from 1B+ images for farm robotics

Higher-quality training signal for personalized shopping AI


Higher-quality NLP signal that cut Dialpad's cost per datapoint

Computer vision signal that optimizes truck loading and unloading

Expert-graded signal for NASA JPL's Martian frost model


Contextual signal that powers Criteo's brand-safe ad targeting


P&G makes owned, trusted AI data an enterprise standard


Extracting clinical signal from millions of unstructured medical records


Expert-verified signal behind Nayya's personalized benefits recommendations


Predicting brand trust by encoding expert judgment into models


Tracking surgical instruments in video to advance robotic surgery

How Ancestry trains models to read historical records faster


Training Walmart's conversational AI on higher-quality language signal


Burberry predicts campaign engagement from its own marketing imagery


Detecting utility defects from drone imagery with CV signal

Encoding stylist judgment into AI-generated fashion ads

Unified training signal that shipped generative AI 5x faster

Clinical-grade ground truth for diagnostic medical imaging models


Active learning that targets a geospatial model's blind spots


Targeting model weaknesses to improve Deque's accessibility AI

Scaling text signal for an edtech question-answering model


A self-retraining data engine for fall-prevention AI

Teaching a spacecraft model to spot life-like motion


Segmentation signal that teaches farm robots weed from crop

Enhancing millions of rental listings with visual intelligence to improve discovery and quality

Enabling a safe launch for a frontier text-to-image model with scaled content moderation
