DNA sequencing: reading a genome, and what it costs
Sequencing determines the order of bases in a DNA molecule. The cost of doing it for a human genome has fallen by orders of magnitude in two decades, and each fall opened a category of biology — and of clinical practice — that had previously been unaffordable.
In one sentence
DNA sequencing determines the sequence of nucleotide bases in a DNA sample, typically by reading many short fragments in massive parallel and reassembling them computationally, or by reading long single molecules directly.
The dominant approach fragments the DNA, binds the fragments to a surface, amplifies each into a cluster, and then reads all clusters simultaneously as fluorescently labelled bases are added one cycle at a time. Parallelism is the whole trick: hundreds of millions of fragments are read at once, and reads are reassembled against a reference.
Long-read methods read a single molecule continuously for tens of thousands of bases, which resolves repetitive regions and structural variation that short reads cannot. Historically less accurate per base and more expensive, they have improved enough to take a meaningful share of work where structure matters.
How it works
The flow cell is a semiconductor product
The surface holding the fragments is patterned with ordered nanowells using photolithography borrowed from chip manufacturing, which packs far more reads into the same area than random binding allowed. It is a direct case of semiconductor process technology setting the cost curve of a biological measurement.
Depth versus breadth
How many times each position is read determines confidence. A rare variant in a tumour sample needs very deep coverage; a straightforward germline genome needs far less. Cost per genome therefore depends on the question, and headline prices usually quote the cheaper case.
Analysis is a real share of the cost
An instrument produces raw signal; turning that into variants requires alignment, calling and interpretation — substantial computing plus curated databases. For clinical use the interpretation and reporting is frequently the larger cost and the harder regulatory problem.
What this depends on
1 of these is marked as a chokepoint: a handful of qualified suppliers, a multi-year lead time, or a single geography.
Supply chain
Flow cells and patterned surfaces
Consumables manufactured with semiconductor patterning; the cost curve depends on that process.
A read is only interpretable against an agreed reference assembly and shared rules for calling a variant pathogenic; without them the measurement says nothing.
What depends on this
Other pages in this map that name DNA sequencing as something they cannot do without.
What each company supplies at this step, and — where a public figure exists — its share of this specific market — with what that share measures, the period it covers and who published it. Some rows also show the company’s own reported revenue for the segment covering this step, which is a different thing: it says how much this business matters to that company, not how much of the market it holds. Not a ranking and not a recommendation.
Supplies long-read sequencing, the accuracy-first alternative to short reads.
Oxford Nanopore TechnologiesLondon
Supplies nanopore sequencing, the only platform that reads a strand directly.
MGI TechShanghai
Supplies the sequencers used across China, and the main challenge to Illumina on price.
What would change the picture
Whether competitive short-read platforms change pricing at the high-throughput end.
Whether long reads take share as accuracy and cost improve.
Whether clinical sequencing volume grows enough to shift the market's centre from research.
Questions people ask about this
Why is a genome quoted at very different prices?
Because the price depends on how deeply it is read and on whether analysis is included. A shallow germline genome on a high-throughput instrument is far cheaper than deep tumour sequencing with clinical interpretation, and headline figures usually quote reagent cost at maximum throughput.
Why are long reads useful if they are less accurate?
Because some questions cannot be answered by short reads at any depth. Repetitive regions and large structural rearrangements are ambiguous when reassembling short fragments, and a single long read spanning the region resolves it directly. Accuracy per base matters less than the span.
Each page explains one technology in plain language, states what it depends on, and names companies by what they supply at that step. Company roles are described qualitatively and deliberately carry no market shares, revenue figures or rankings — those change faster than an explainer can, and a stale number is worse than none. Ticker links point at company pages on this site and are provided for reference only.
Nothing here is investment advice, a recommendation, or a forecast. A company named on a page about a technology is not thereby a good investment, and the chokepoints described are structural facts about supply chains rather than predictions about prices. Technology moves; where a page describes something as unresolved or in development, that was true when it was written.
Plutux no es un asesor de inversiones. Los datos de mercado y el análisis generado por IA son solo informativos y educativos, no asesoramiento de inversión. Aviso legal