Plutux

Sub-system

Teleoperation and data collection: where the training examples come from

Every learned manipulation policy rests on demonstrations, and demonstrations are produced by people operating robots. That makes data collection an industrial process with a headcount, a cost per hour and a quality problem — the same shape as the annotation industry behind language models, with hardware attached.

In one sentence

Teleoperation is the remote or direct human operation of a robot to perform a task, recorded as paired observation and action data for training control policies.

The interfaces vary in fidelity and cost. Virtual reality controllers map hand motion to the robot's end effector — cheap and intuitive, with no sense of what the robot is touching. Exoskeletons and leader-follower rigs, where the operator moves a replica of the robot's arm, give better correspondence and force feedback and cost far more. The choice sets both the data's quality and its price per hour.

Everything the annotation industry learned applies here. Guidelines have to be precise or different operators solve the task differently and the policy learns the average of inconsistent behaviour. Sessions must be reviewed. Failed episodes are valuable but only if labelled as failures. It is a managed operation, not a recording session.

How it works

Force feedback changes what can be taught

Contact tasks depend on how hard to push and when to stop, and an operator who cannot feel the contact cannot demonstrate that judgement. Interfaces that render force back to the operator produce demonstrations containing the information a policy needs for insertion and assembly — and they are substantially more expensive.

Deployment as data collection

A robot working with a human supervisor produces data continuously, including the interventions when it fails — which are exactly the corrections a policy needs. Several deployment models are explicitly structured this way: the customer gets work done, the operator gets training data, and the machine improves on the task it is actually being asked to do.

Cross-embodiment data

Data collected on one robot does not directly apply to another with different kinematics. Efforts to pool datasets across robot types, and to train policies that condition on the body they are controlling, are attempts to escape that — because the alternative is every developer collecting their own data from scratch.

What this depends on

1 of these is marked as a chokepoint: a handful of qualified suppliers, a multi-year lead time, or a single geography.

  • ResourceChokepoint

    Trained teleoperators

    A labour operation with recruitment, training and quality management, priced per hour of demonstration.

  • Resource

    Low-latency links for remote operation

    Operating at a distance needs latency low enough to close the human control loop, which limits where it can be done.

  • Technology

    Data pipelines and quality control

    Synchronised multi-sensor recording, review, labelling and storage — the same machinery as any large training set.

    Data pipelines
  • Supply chain

    Robot hardware to demonstrate on

    Every episode occupies a physical robot for as long as the task takes, so the size of the fleet rather than the willingness to record sets how much data can be produced.

    Humanoid hardware
  • Technology

    Force sensing at the robot

    An interface can only render contact back to the operator if the machine measures it; without that the demonstrations omit exactly the judgement contact tasks depend on.

    Force and tactile sensing

What depends on this

Other pages in this map that name Teleoperation and data as something they cannot do without.

Who supplies this

What each company supplies at this step, and — where a public figure exists — its share of this specific market — with what that share measures, the period it covers and who published it. Some rows also show the company’s own reported revenue for the segment covering this step, which is a different thing: it says how much this business matters to that company, not how much of the market it holds. Not a ranking and not a recommendation.

  • TeslaTSLA

    Operates in-house data collection for its humanoid manipulation programme.

  • TeradyneTER

    Supplies collaborative robot platforms widely used as research and data-collection hardware.

  • NVIDIANVDA

    Supplies simulation and data generation tooling used alongside physical collection.

  • Figure AIPrivate

    Collects demonstration data from deployed humanoids under supervision.

  • Scale AIPrivate

    Sells the operator time and annotation behind demonstration datasets, the same business it built for language.

  • 1X TechnologiesPrivate

    Runs teleoperated humanoids in real homes and uses the sessions as the training set.

  • Hyundai Motor005380.KS· Korea

    Owns the robotics group whose field deployments generate the longest-running real-world manipulation logs.

  • Physical IntelligencePrivate

    Collects cross-embodiment demonstration data as the input to models meant to run on hardware it does not build.

What would change the picture

  • Whether pooled cross-embodiment datasets reduce the per-developer data burden.

  • Whether video-derived action data substitutes for teleoperation at acceptable quality.

  • Whether supervised deployment becomes the standard data-collection model.

Questions people ask about this

Why is data collection so expensive?
Because every example takes real time on real hardware with a trained person operating it. There is no equivalent of scraping. A dataset large enough to train a general manipulation policy therefore represents a substantial number of paid operator-hours, plus the rigs and the review process around them.
Can data from one robot train another?
Only partly. Different kinematics mean the same task requires different joint motions, so raw action data does not transfer. Policies that condition on the body they control, and datasets pooled across robot types, are the current attempts to make it transfer — with meaningful but incomplete results.

How these pages are written

Each page explains one technology in plain language, states what it depends on, and names companies by what they supply at that step. Company roles are described qualitatively and deliberately carry no market shares, revenue figures or rankings — those change faster than an explainer can, and a stale number is worse than none. Ticker links point at company pages on this site and are provided for reference only.

Nothing here is investment advice, a recommendation, or a forecast. A company named on a page about a technology is not thereby a good investment, and the chokepoints described are structural facts about supply chains rather than predictions about prices. Technology moves; where a page describes something as unresolved or in development, that was true when it was written.

Plutux is not an investment adviser. Market data and AI-generated analysis are for information and education only, not investment advice. Disclaimer

© Plutux Technology Limited 2026
Teleoperation and data — Humanoids and embodied AI: How It Works and What It Depends On | Plutux