Segment
The data layer: what models are actually made of
Compute gets the attention because it is expensive and visible. Data decides what a model knows. This branch covers where training data comes from, what has to be done to it before it is usable, and the infrastructure that lets a deployed model reach information it was never trained on.
In one sentence
The AI data layer covers the corpora models are trained on, the pipelines that clean and prepare them, the human operations that produce demonstrations and judgements, and the retrieval systems that supply information at answer time.
Almost none of the effort here is glamorous, and almost all of it matters. A frontier corpus is assembled from crawled text, licensed archives, code, and increasingly model-generated material, then filtered, deduplicated and quality-scored — with the filtering doing more for final quality than most architectural choices.
Two shifts have changed the picture. Legal and commercial constraints have pushed corpora from freely scraped toward explicitly licensed, which makes data a contracted cost and an advantage for organisations holding proprietary archives. And retrieval has partly decoupled what a model knows from what it was trained on, moving some of the burden from training to serving.
How this breaks down
Split by where the data comes from and when it is used — training time versus answer time.
- Data processing frameworksThe batch and streaming engines that clean, deduplicate and tokenise a corpus.Definition page
- Training corporaWhere the tokens come from, and why filtering matters more than volume.Definition pageChokepoint
- Data pipelinesThe batch processing and streaming layer between raw data and a training step.Definition page
- Human data operationsDemonstrations, preference comparisons and expert evaluation — a real industry, not a footnote.Definition pageChokepoint
- Embeddings and vector searchTurning text into vectors so a system can find what is relevant rather than what matches keywords.Definition page
What this depends on
1 of these is marked as a chokepoint: a handful of qualified suppliers, a multi-year lead time, or a single geography.
- StandardChokepoint
Licensing and crawl permissions
What may legally be trained on is increasingly a matter of contracts and site policies rather than technical availability.
- Supply chain
General-purpose compute and object storage
Corpus work runs on ordinary cloud machines against petabytes of object storage, not on accelerators. It is a large bill that never appears in accelerator-hour accounting.
Data centres
Companies across Data layer
Every company named on a step below this page, ordered by how many of those steps it appears at. Compiled from the pages themselves rather than written separately, so the two cannot disagree. Not a ranking and not a recommendation.
- DatabricksPrivate3 steps
Data processing frameworks · Data pipelines · Embeddings and vector search
- Hugging FacePrivate3 steps
Data processing frameworks · Training corpora · Data pipelines
- AnyscalePrivate2 steps
- Scale AIPrivate2 steps
- AppenAustralia1 step
32 more companies appear at a single step each; they are named on the pages for those steps.
How these pages are written
Each page explains one technology in plain language, states what it depends on, and names companies by what they supply at that step. Company roles are described qualitatively and deliberately carry no market shares, revenue figures or rankings — those change faster than an explainer can, and a stale number is worse than none. Ticker links point at company pages on this site and are provided for reference only.
Nothing here is investment advice, a recommendation, or a forecast. A company named on a page about a technology is not thereby a good investment, and the chokepoints described are structural facts about supply chains rather than predictions about prices. Technology moves; where a page describes something as unresolved or in development, that was true when it was written.
Plutux is not an investment adviser. Market data and AI-generated analysis are for information and education only, not investment advice. Disclaimer