Plutux

Segment

The data layer: what models are actually made of

Compute gets the attention because it is expensive and visible. Data decides what a model knows. This branch covers where training data comes from, what has to be done to it before it is usable, and the infrastructure that lets a deployed model reach information it was never trained on.

In one sentence

The AI data layer covers the corpora models are trained on, the pipelines that clean and prepare them, the human operations that produce demonstrations and judgements, and the retrieval systems that supply information at answer time.

Almost none of the effort here is glamorous, and almost all of it matters. A frontier corpus is assembled from crawled text, licensed archives, code, and increasingly model-generated material, then filtered, deduplicated and quality-scored — with the filtering doing more for final quality than most architectural choices.

Two shifts have changed the picture. Legal and commercial constraints have pushed corpora from freely scraped toward explicitly licensed, which makes data a contracted cost and an advantage for organisations holding proprietary archives. And retrieval has partly decoupled what a model knows from what it was trained on, moving some of the burden from training to serving.

How this breaks down

Split by where the data comes from and when it is used — training time versus answer time.

What this depends on

1 of these is marked as a chokepoint: a handful of qualified suppliers, a multi-year lead time, or a single geography.

  • StandardChokepoint

    Licensing and crawl permissions

    What may legally be trained on is increasingly a matter of contracts and site policies rather than technical availability.

  • Supply chain

    General-purpose compute and object storage

    Corpus work runs on ordinary cloud machines against petabytes of object storage, not on accelerators. It is a large bill that never appears in accelerator-hour accounting.

    Data centres

Companies across Data layer

Every company named on a step below this page, ordered by how many of those steps it appears at. Compiled from the pages themselves rather than written separately, so the two cannot disagree. Not a ranking and not a recommendation.

32 more companies appear at a single step each; they are named on the pages for those steps.

How these pages are written

Each page explains one technology in plain language, states what it depends on, and names companies by what they supply at that step. Company roles are described qualitatively and deliberately carry no market shares, revenue figures or rankings — those change faster than an explainer can, and a stale number is worse than none. Ticker links point at company pages on this site and are provided for reference only.

Nothing here is investment advice, a recommendation, or a forecast. A company named on a page about a technology is not thereby a good investment, and the chokepoints described are structural facts about supply chains rather than predictions about prices. Technology moves; where a page describes something as unresolved or in development, that was true when it was written.

Plutux is not an investment adviser. Market data and AI-generated analysis are for information and education only, not investment advice. Disclaimer

© Plutux Technology Limited 2026
Data layer — Artificial intelligence: How It Works and What It Depends On | Plutux