YouMind
Iniciar sessão

Robotics Data is a Broken Business

164K
1.2K
104
78
1.3K

TL;DR

Sam Padilla details the shutdown of Eidon AI, explaining that while embodied AI faces a data bottleneck, the business model fails due to complex logistics, a tiny buyer network, and unsustainable unit economics.

A technical bottleneck does not necessarily imply a good business.

We recently shut down @eidon_ai after 2+ years in the data space.

We started in early 2024 with a gamified multimodal data collection app for video and images. In early 2025, we pivoted into robotics data.

The thesis from the beginning was simple:

Embodied AI suffers from a critical data bottleneck.

So we built the Eidon Tracker, an IMU-based kit for collecting high-precision 7-DOF arm tracking data alongside first-person video.

Sam Padilla - inline image

We deployed the tracker to 50+ collectors across several countries and collected 1,000+ hours of household task data, now all open source:

https://x.com/LeRobotHF/status/2102011968026554652

Two years in, the premise still holds:

Embodied AI suffers from a critical data bottleneck.

But it may be that the value of the data doesn't justify building a business around it.

The Logistics Nightmare

Useful data needs standards. Visible hands, consistent lighting, clean POV framing, task compliance, and enough duration to make the recording usable.

This implies QA on every recording, on top of recruitment, communication, compensation, and moderation of your collectors.

The industry is now also so saturated that, unless you have massive scale, video-only data is a commodity. The more valuable data comes from custom collection hardware.

And when hardware gets involved, you graduate from just managing collectors to shipping hardware, writing instructions, debugging setup issues, checking calibration, rejecting bad recordings, answering support messages, coordinating with customs, and trying to keep quality consistent across different people, homes, countries, and time zones.

Sam Padilla - inline image

Preparing our first shipment

We worked with third-party collector logistics companies like BeeWork. We shipped kits to them, and they coordinated end-collector logistics while enforcing our standards.

This helped, but it did not make the logistics disappear. It increased the per-hour cost of collection, and it still left us with the weird edge cases that only show up once hardware is in the field.

Deployed hardware very much follows the Anna Karenina principle: All working devices work alike; each broken device breaks in its own way.

The Buyer Network

There are only a handful of companies in the world for whom this data is valuable. And inside those companies, only a few people can decide to spend hundreds of thousands of dollars on a dataset.

In practice, the buyer network for robotics data is tiny. Maybe a couple hundred people globally. At this level, it is not even a marketing problem. You can market all you want; at the end of the day, you either have the relationship, or you don't.

And it is not enough to just know people at the company or even know robotics engineers at the company. You need to have a direct line to the teams making the final call.

Sam Padilla - inline image

This access is so hard to get that it becomes a product of its own. Data brokers and aggregators fill that gap. We worked with companies like Defined AI, who already held previous commercial relationships with big labs and basically brokered deals between the end buyers and us.

While helpful, it also introduced the normal costs of selling through an intermediary: slower communication, filtered requirements, diluted margins, misaligned incentives, and less direct feedback from the people who actually needed the data.

Getting in front of end-buyers is possible. We did it. But those conversations open a different set of problems: scale.

The Chicken and Egg of Scale

We reached a total of 3,000 hours of data. Half of it was POV video only (OS here: https://huggingface.co/buckets/eidon-ai/egocentric-pov). The other half was the aforementioned POV video + tracker datasets.

That was a hard-earned figure. It took us about 5 months from the day we shipped our first batch of trackers. But when we sat down to converse with end-consumer labs, the request was closer to 3,000 hours monthly, preferably weekly.

Sam Padilla - inline image

Admin dashboard showing almost 2k hours collected

Data has a short useful lifespan. Once a customer trains on a batch, that batch becomes borderline useless. Labs don't need a one-time dataset, they want a reliable, recurring data pipeline.

The problem is that many labs will not treat you as a serious supplier until you can prove you already have that pipeline. So you are forced to make an early bet: what kind of data will you collect?

But that bet is almost impossible to make cleanly, because no two customers want the same dataset.

We bet on IMU-based joint-angle derivation of human arm joints. That immediately cut out labs that only cared about hand dexterity, believed IMUs were the wrong approach, or wanted gripper teleoperation instead.

Even among labs open to our data, requirements were not standard. Some wanted final joint angles from our kinematics chain. Others wanted raw IMU measurements. And if you collect for one format without preserving the other, you can't just go back later and recover the missing data.

Sam Padilla - inline image

Yours truly

The R&D, manufacturing, logistics, and QC costs of bootstrapping for a given data shape are so great for a startup that making the wrong bet early is often an early death sentence.

And a very easy way to be almost guaranteed to make the wrong bet is to choose the easy option. We did that.

Hour Economics

We sold 1k-hour checkpoints of video+tracker data for ~$60/h. Depending on the quality of your data (i.e. finger dexterity), you could command up to $100/h.

This works on paper. But from that sticker price, you need to subtract:

  • Collector pay (roughly 10% of the final hour price)
  • Hardware BOM
  • Loss/breakage/replacement
  • Customs
  • Assembly time
  • QA / rejected recordings

We figured out how to build low-cost hardware and deploy it at scale. So we ate every cost to build and get the hardware into our collectors' hands.

All of this was financed by us, up front. You may spend time and money collecting 1k hours of data for a given buyer, to realize that the lab next door wants different sensors, tasks, and signals.

We had samples that we socialized early and ensured the data was useful, but there was only so much we could adjust in a timely fashion once the hardware was in the field.

Getting contracts up front is also borderline unfeasible. The collector industry is so saturated that the labs and aggregators hold all the leverage. If you don't meet their need for scale/specs, they can just go to the other dozen data suppliers lining up to do business with them.

Solve Hard Problems

There is a consensus in tech that robotics is the next big thing. And as it often happens with "the next thing," a lot of people are flocking into the field. People with software backgrounds are making that transition and bringing over software startup heuristics that can kill you in hardware.

We developed the Eidon Glove around the same time as we developed the trackers. The glove was a very low-cost 19-DOF finger and hand tracker based on Project Homunculus by @LeRobotHF and @nepyope.

Sam Padilla - inline image

Eidon Gloves

It used a Hall-effect sensor to capture high-resolution finger joint angles and an IMU to measure wrist motion.

The data it could collect was obviously more valuable than arm tracking. Dexterity is where a lot of robotics actually breaks. But the glove was also much harder to manufacture, calibrate, support, and deploy.

Coming from software myself, I made the software-founder choice: pick the simpler product, ship faster, iterate later.

So we chose the trackers, without knowing that it would cost thousands of dollars and months to deploy that at scale, and basically put us on a path of no return.

We made the same, but bigger, mistake when we first started Eidon.

From the beginning, we knew the end goal was to break into robotics. But in early 2024, this was not an obvious bet. So again, we chose the easy way out: a gamified data collection app with quests and crypto-synergy focused on videos and images.

Sam Padilla - inline image

Our first quest-based app

Had we chosen the hard thing from the start, or had we at least chosen to double down on the gloves instead of the trackers, Eidon may still be around today.

The Shrinking Opportunity Window

Since @clementdelangue and team made the announcement of our final dataset earlier this week, I have received messages from at least 5 different early founders who are exploring the robotics data space. Accelerators like YC are also cranking these out, backing anything robotics-adjacent.

https://x.com/ClementDelangue/status/2102046770947613026

But I think a lot of people don't realize that data collection companies have a limited window of opportunity.

The ultimate goal of any robotics lab is to build a data flywheel. Deployed instances collect their own data and, in turn, contribute to the iteration of their world models or VLAs.

Robotics data providers are simply the bootstrapping force of that data flywheel. And once these labs can fulfill their own data needs - be it via an internal data collection team or data feeds from their fleet - the need for third-party data collectors will evaporate.

Think of robotics data companies as the drivers on Waymos before they are allowed to operate (and collect their own data) autonomously. They play an important, but very time-bound role.

We had this clear since we started. And the goal was always to move past data collection into hardware, simulation environments, and model building. We called this "graduating".

Sam Padilla - inline image

How long the opportunity window will last is anyone's guess. I put it at 1-3 years. Regardless, I think the runway for a new startup in data to graduate is getting really short.

Parting Thoughts

This work matters. Data and robotics at large matter.

So in an attempt to contribute to this community, we have open-sourced pretty much anything of value we built in our two years:

If you build something with any of these, please let me know. I would love to see this being used.

And if you are thinking about starting something in the robotics data space, I hope you will consider these lessons:

  1. Minimize your logistics overhead as much as possible.
  2. Have clear incentives for data collectors.
  3. Choose hard things. Difficulty == moat.
  4. Have a clear path to graduate out of data.
  5. Spend time and money on networking early.
  6. Get samples and socialize them fast.
  7. Never sign data exclusivity deals.
  8. Give back to the community.

There is more to this. Economics, building custom hardware, BOMs, the synergies with open source, and much more. But I will expand on those points in other essays if there is interest.

Huge shoutout to the team; it was a good run:

@RobertDaleSmith now at neuralink and @peter_van_toth continuing some of this work.

Sam Padilla - inline image
Guardar com um clique

Faça leitura aprofundada de artigos virais com IA no YouMind

Guarde a fonte, faça perguntas específicas, resuma o argumento e transforme um artigo viral em notas reutilizáveis num único espaço de trabalho com IA.

Explorar o YouMind
Para criadores

Transforme o seu Markdown num artigo 𝕏 impecável

Quando publica os seus próprios textos longos, formatar imagens, tabelas e blocos de código para o 𝕏 é uma dor de cabeça. O YouMind transforma um rascunho completo em Markdown num artigo 𝕏 impecável e pronto a publicar.

Experimente Markdown para 𝕏

Mais padrões para decifrar

Artigos virais recentes

Explorar mais artigos virais