XDOF Nears $1.2B Series B Just Three Months After Coming Out of Stealth
XDOF, a robotics training-data startup founded in 2024, is in late-stage talks for a Series B at a $1.2B valuation led by 8VC, with annualized revenue near $50M.

XDOF, a startup that builds training-data pipelines for physical robots, is in late-stage talks to raise a Series B at a valuation of about $1.2 billion led by 8VC, according to TechCrunch sources. The company only left stealth a few months ago, closed a $70 million Series A in June 2026, and was not planning another raise so soon. Investors came to XDOF after its annualized revenue approached $50 million, a pace that persuaded backers not to wait for a standard fundraising cycle.
What happened
| Detail | Fact |
|---|---|
| Round | Series B (late-stage talks) |
| Valuation | ~$1.2 billion |
| Lead investor | 8VC |
| Series A size | $70 million (closed June 2026) |
| Series A investors | Thrive Capital, Andreessen Horowitz, Lux, Spark Capital |
| Annualized revenue | Approaching $50 million |
| Current customers | 20, including several frontier AI labs |
| Founded | 2024, by Philipp Wu (CEO) and Fred Shentu (CTO) |
XDOF and 8VC did not respond to requests for comment. The deal terms are not final and could still change, and TechCrunch was unable to confirm the total amount being raised or whether the $1.2 billion figure includes new capital.
What does XDOF actually do?
The company builds what investors are calling the data-supply chain for physical robotics. That means collection tools, annotation systems, and data pipelines that AI labs and robotics companies need to train general-purpose robots but cannot easily build in-house.
The analogy being floated in venture circles is Scale AI or Mercor for robots. Scale AI became a multi-billion-dollar business by labeling text and image data for large language models. XDOF is betting the same bottleneck exists for physical AI: unlike LLMs, which trained on text scraped from the web, robots have no equivalent real-world dataset to draw from.
To fill that gap, XDOF uses two collection methods. First, remote teleoperation: human operators steer robotic arms from a distance to generate movement data. Second, egocentric capture: workers wear body sensors to record everyday tasks like folding laundry or flattening cardboard boxes. The startup plans to hire and train data-collector teams around the world at scale.
XDOF is also partnering with UC Berkeley’s AI Research lab to release what it says is the largest collection of high-quality robot training data ever assembled, a dataset internally called ABC.
Where the founders came from
Co-founders Wu and Shentu are both UC Berkeley researchers. Wu was studying how robots learn from large datasets when he ran into a familiar wall: not enough data to work with. He and Shentu built GELLO, a low-cost teleoperation system that lets a human operator control a robotic arm remotely to produce training data. Their work produced an influential robotics paper and became the technical foundation for XDOF.
Why it matters
Physical AI is attracting enormous capital precisely because the data problem is hard and structural. Software AI had the internet. Robots need humans in the loop, collecting motion data in the real world, and that is slow and expensive work. Companies that solve it at scale hold a durable position in the supply chain, not just a product advantage.
The speed of this fundraise is also telling. XDOF did not go out looking for a Series B. VCs came to them, reportedly driven by the revenue trajectory. Reaching $50 million in annualized revenue within months of leaving stealth puts XDOF in the same conversation as the fastest-growing infrastructure startups in recent memory. For context, check our ongoing coverage of AI funding rounds to see how this compares to peers.
Competitors in this space include Mecka AI and broader data platforms like Scale AI and Micro1, which are expanding from LLM data into robotics. XDOF’s head start with frontier AI lab customers and its Berkeley research roots give it credibility, but the field will get crowded fast as the prize becomes clearer.
Our take
The “Scale AI for robots” framing is accurate enough to be useful. Real-world motion data has no Wikipedia equivalent. Someone has to collect it, clean it, and deliver it at production quality. That is an operations problem as much as a technology problem, and operations businesses tend to be stickier than model businesses once they have enterprise contracts.
What we would watch: how XDOF maintains data quality as it scales collector teams globally, and whether the ABC dataset with Berkeley becomes a genuine moat or a one-time PR moment. If you are building AI-powered products and wondering where your training data comes from, the robotics sector is showing us what that supply chain looks like when it is built deliberately. Businesses integrating AI into physical workflows should pay attention. Our AI integration work increasingly runs into exactly this data-availability problem.
Frequently asked questions
What is XDOF and what does it do?
XDOF is a startup founded in 2024 by UC Berkeley researchers Philipp Wu and Fred Shentu. It collects real-world teleoperation and sensor data to train general-purpose robots, building the data pipelines and annotation systems that AI labs and robotics companies need but cannot easily build themselves.
How much is XDOF raising in its Series B?
XDOF is in late-stage talks for a Series B at a valuation of about $1.2 billion, reportedly led by 8VC. The total amount being raised has not been disclosed and the deal terms are not yet final.
How fast is XDOF growing?
XDOF's annualized revenue is approaching $50 million. The company only emerged from stealth a few months ago and raised a $70 million Series A in June 2026, making this growth trajectory unusually rapid.
Who are XDOF's competitors in the robot training data space?
Competitors include Mecka AI and larger data platforms such as Scale AI and Micro1, which are expanding their services from LLM data labeling into robotics training data.


