Robot teams can see when a robot drops a cup, misses a handoff, or fails to recover. What they cannot quickly determine is whether the failure came from missing data, bad annotations, an eval gap, a policy change, or sim to real.
The evidence is scattered across the robot learning stack, so every failure becomes a manual investigation. Teams spend weeks guessing what to collect, relabel, retrain, or redesign.
What we are building
We are building the debugging and evaluation layer for robot learning.
Our platform ingests real world rollouts alongside the data, annotation, eval, and policy versions that produced them. It groups recurring failures, turns them into repeatable evals in simulation, and identifies the pipeline changes most likely to have caused an improvement or regression.
Researchers can see what changed, test why performance moved, and decide what to fix next.
Initial wedge
We are starting with robotics data quality.
Data suppliers need to demonstrate that their datasets improve policy performance. Robotics labs need to reject low value data before spending time and compute training on it.
We connect datasets to measurable eval outcomes, allowing suppliers to validate valuable data and buyers to identify missing, redundant, or low signal data.
Why now
Robot policies are improving, real world deployments are growing, and simulation and world models are becoming useful enough to test real failures at scale.
The bottleneck is shifting from simply collecting more data to learning from failures faster.
Traction
We have:
- $100K in contracted ARR from Vision Lab
- $10K in paid pilot revenue from Contracted AI
- Paid pilots beginning with DoorDash, Instawork, and Humaid
- A broader pipeline across robotics companies, labs, and data providers
Team and vision
The founding team combines robot policy and data quality research at Berkeley BAIR with experience building large scale robotics data, annotation, and quality systems at Amazon FAR AGI lab.
We intend to become the system of record for robot learning: the place teams decide what data to collect, what models to train, and what policies to ship.
We are raising $1.5 million, and half of the round has already been requested, to convert current pilots into recurring contracts and prove that the platform improves data value and policy performance across multiple customers.



