Custom AI ASICs: chips a buyer designs for its own workload
The largest buyers of AI compute are also, increasingly, chip designers. A custom accelerator gives up flexibility and takes on design risk in exchange for a lower cost per unit of work on a workload the buyer already knows in detail.
In one sentence
A custom AI ASIC is an accelerator designed by or for a single operator, targeting the specific models and serving patterns that operator runs, rather than sold as a general-purpose part.
The case for building one is arithmetic. If a company will spend billions of dollars a year running one narrow class of workload, a design that removes everything that workload does not use — and adds exactly the memory and interconnect it does — can beat a merchant part on cost per token by enough to pay for a chip team many times over.
The case against is that the chip takes two to three years from specification to deployment, and the workload it was specified for may not be the workload that matters when it lands. This is why custom parts have historically targeted inference and stable, high-volume internal services before they targeted frontier training.
How it works
Who actually builds it
Rarely the operator alone. The architecture and the parts that encode the workload are in-house; the physical implementation, the high-speed interfaces, the packaging and the relationship with the foundry usually come from a design-services partner. That is a real business line for the large merchant silicon vendors, and it means an in-house chip still routes through the same foundry and packaging queues.
What gets removed, and what gets added
Out go graphics blocks, most of the general-purpose programmability, and support for numeric formats the operator does not use. In go arithmetic units matched to the operator's model shapes, on-chip memory sized to its layers, and interconnect matched to how its clusters are wired. The result is efficient on that workload and awkward on anything else.
The strategic reason, separate from cost
A second source changes a negotiation. Even a custom part that never displaces the majority of a fleet changes what the merchant vendor can charge for the rest of it, which is why the programme is often worth running before the chip is competitive on its own terms.
What this depends on
3 of these are marked as a chokepoint: a handful of qualified suppliers, a multi-year lead time, or a single geography.
Supply chain
Design services and IP
SerDes, memory controllers, packaging expertise and physical implementation are bought in. Very few organisations carry all of it internally.
A custom part with stacked memory needs the same silicon interposer and the same assembly line as a merchant accelerator, and it joins the same queue for them.
Operator parts are usually assembled from several dies so that a compute die can be re-spun without redoing the interfaces. Without a working die-to-die link the design collapses back to one reticle-limited chip.
Timing, power and physical verification are done in the same licensed tool flow the merchant vendors use. A tape-out that fails sign-off costs a mask set and a quarter.
What each company supplies at this step, and — where a public figure exists — its share of this specific market — with what that share measures, the period it covers and who published it. Some rows also show the company’s own reported revenue for the segment covering this step, which is a different thing: it says how much this business matters to that company, not how much of the market it holds. Not a ranking and not a recommendation.
Takes an operator's architecture to a manufacturable part, the Japanese counterpart to the Taiwanese design-service houses.
ByteDancePrivate
Co-designs its own inference silicon with merchant partners to cut what it buys for recommendation and model serving.
What would change the picture
Whether custom parts move from inference into frontier training at scale, or stay on stable high-volume workloads.
Whether design-services capacity — not foundry capacity — becomes the limit on how many custom programmes can run at once.
How much of a fleet has to be custom before the pricing effect on merchant parts is the main return on the programme.
Questions people ask about this
Does a custom chip mean the operator stops buying merchant GPUs?
In practice, no. Custom parts have generally been additive: they take the workloads whose shape is known and stable, while merchant GPUs keep the research, the newest architectures, and anything that needs to run tomorrow rather than in the next chip generation.
Why does this still help the merchant silicon industry?
Because the custom part is not built in a vacuum. It uses bought-in interconnect IP, a design-services partner, the same foundry, the same packaging and the same memory. The spend moves along the chain rather than leaving it.
Each page explains one technology in plain language, states what it depends on, and names companies by what they supply at that step. Company roles are described qualitatively and deliberately carry no market shares, revenue figures or rankings — those change faster than an explainer can, and a stale number is worse than none. Ticker links point at company pages on this site and are provided for reference only.
Nothing here is investment advice, a recommendation, or a forecast. A company named on a page about a technology is not thereby a good investment, and the chokepoints described are structural facts about supply chains rather than predictions about prices. Technology moves; where a page describes something as unresolved or in development, that was true when it was written.