SILICON · EPISODE 9 OF 10 · 12 min

CPU vs ASIC vs FPGA vs GPU

Last time, we put logic and memory together. Now, what do we build out of them? Apple's 2024 iPhone 16 Pro answers in an odd way. Its main chip holds a six-core processor, a six-core graphics unit, and a sixteen-core "neural engine". Why three kinds of brain on one chip, instead of one good one?

▶ Watch on YouTubePlay the whole Silicon series ▶

The full story

This is the episode's narration, word for word. Headings jump to that point in the video.

The question 0:00

Last time, we put logic and memory together. Now, what do we build out of them? Apple's 2024 iPhone 16 Pro answers in an odd way. Its main chip holds a six-core processor, a six-core graphics unit, and a sixteen-core "neural engine". Why three kinds of brain on one chip, instead of one good one? And why did a chip designed for video games end up running artificial intelligence?

A machine that follows a list 0:31

Start with the chip that can do anything. A processor runs a list of instructions, stored in memory as numbers. It repeats one loop, over and over. Fetch the next instruction. Decode it, which means working out which operation it asks for and which data it means. Execute it, in an arithmetic unit built from adders like the one from episode seven. Then step to the next instruction. The clock from episode seven sets the pace, and a handful of tiny, fast memories, called registers, hold the numbers being worked on.

Eight ticks 1:11

In 1971, Intel squeezed that whole loop onto one chip with two thousand three hundred transistors: the 4004. Its data sheet spells the loop out. Three clock ticks to send the address, two to bring the instruction back, three to carry it out. Eight ticks, about eleven millionths of a second, for a basic instruction. Other groups had built processor chips around then, for a fighter jet and for a computer maker. The 4004 was the one put on open sale, to anyone who wanted it.

Same chip, any job 1:48

It was designed for a Busicom calculator. But Intel's data sheet pitched it for billing machines, terminals and test equipment. Change the list of instructions, and the same chip becomes a different product. A 2023 desktop chip runs the same loop on twenty-four cores, ticking up to six billion times a second.

The cost of reading the recipe 2:12

But flexibility has a cost, and you pay it on every single step. Think of a cook who walks over and re-reads the recipe card before every stir. The stir is cheap. The reading isn't. The Stanford engineer Mark Horowitz put rough numbers on it. Adding two numbers takes about a tenth of a picojoule. But on a simple processor, fetching, decoding and keeping track of that one instruction costs around seventy picojoules. For a simple sum, the bookkeeping costs hundreds of times more than the sum.

Bake it into the wiring 2:49

So what if you skip the recipe? Wire the transistors to do one job and nothing else, so the data flows straight through, with no instructions to fetch at all. That's an ASIC, an application-specific integrated circuit. In one study, a general-purpose chip used five hundred times more energy than an ASIC to encode the same high-definition video. Even doing over ten operations per instruction, it was still fifty times worse, because nine-tenths of its energy was overhead.

Then why not build everything that way? 3:25

Because an ASIC is expensive to start and impossible to change. Every design needs its own set of masks for the printing from episode six, plus a year or two of design and testing. For the most advanced chips, estimates run to hundreds of millions of dollars. And once it's printed, the circuit is fixed in the metal. A bug, or a new standard, means new masks. So an ASIC only pays off for a huge number of chips, doing a job that won't change.

A race you can watch 3:59

There's one race where you can watch this play out. Bitcoin mining is one fixed calculation, a scrambling calculation called a hash, repeated endlessly for money. Miners started on ordinary processors in 2009, moved to graphics chips in 2010, then to reconfigurable chips in 2011. In January 2013, the first mining ASICs arrived. One, made on a process years out of date, was forty times more energy-efficient than a then-current graphics chip. And it could do nothing else.

A chip you can rewire 4:38

Those 2011 chips were FPGAs: field-programmable gate arrays. Picture a sea of small logic blocks, with switchable wiring between them. Each block doesn't contain a fixed gate. It contains a tiny memory holding a truth table, so it can act as any gate you like. The wiring switches are transistors, each set on or off by another memory bit. Load a different set of bits, and the same chip becomes a different circuit.

Freeman's bet 5:12

In 1985, Xilinx, a company co-founded by Ross Freeman, sold the first one. It had sixty-four logic blocks, enough for a few thousand gates. Freeman's bet was that transistors would get so cheap that flexibility would matter more than perfect efficiency. Its configuration memory is last episode's static RAM: two inverters holding each other in place.

The price of the switchboard 5:40

But all that switchable wiring has a price. One careful study compared FPGAs with ASICs made on the same process. The FPGA needed about thirty-five times the silicon area, ran three to four times slower, and used about fourteen times the power when switching. Most of the chip is switches and memory, not the logic you actually wanted.

What they're for 6:08

So why use one? Because you don't need masks. Engineers test a chip design on FPGAs before paying for an ASIC. They go in products made in small numbers, and in gear like phone base stations and network routers, where the standards keep changing. Microsoft says it put one into nearly every new server in its data centres, starting with its search engine. And those Bitcoin miners built their ASICs from designs very like their FPGA ones.

The triangle 6:42

That's the trade-off. Flexible, efficient, or cheap and quick to put to a new job: pick where you sit, because you can't have all three corners. A processor is flexible, but wasteful. An ASIC can be a hundred to a thousand times more efficient, but costly and fixed. An FPGA sits in between. And then there's a fourth answer, which started with pixels.

Built for pixels 7:10

A full-HD screen has about two million pixels, and many redraw them sixty times a second. That's over a hundred million colours to work out every second, and each one needs much the same maths. So graphics chips took a different approach. Instead of a few clever cores, build thousands of simple arithmetic units, and have one instruction steer a whole group at once, each on its own pixel. Nvidia marketed its 1999 chip as the first "GPU", a graphics processing unit.

Sharing the recipe card 7:47

Remember the cook re-reading the recipe? A GPU reads the card once, and a whole row of cooks stir together. The overhead is shared, so far more of the energy goes on the arithmetic itself, especially for the heavier maths graphics needs. Nvidia's H100 has nearly seventeen thousand of these simple units, and eighty billion transistors. For AI it adds 528 Tensor Cores, specialised just for the matrix maths. The catch: each group of units must do the same thing at the same time. Give units in a group different jobs, and they take turns.

Opening it up 8:28

In November 2006, Nvidia announced CUDA, a way to program its graphics chips in C, a language programmers already knew. The toolkit was released in June 2007. By 2009, researchers were training neural networks on GPUs up to fifty-five times faster than on a processor, cutting weeks down to about a day.

What a neural network actually does 8:57

Why do neural networks fit GPUs so well? Because a neural network is mostly one operation. Each layer multiplies its inputs by a big grid of learned weights, and adds up the products. That's a matrix multiplication. Multiply two grids a thousand numbers wide, and that's a billion multiply-and-adds. And each of the million answers can be worked out without waiting for any other. For an image network, about ninety-five per cent of GPU time went on exactly that.

AlexNet 9:32

In 2012, a network called AlexNet, from the University of Toronto, won a major image-recognition contest. Each network took five to six days to train, on two off-the-shelf graphics cards. Its winning entry, an average of several such networks, got the top-five error down to fifteen per cent, against twenty-six per cent for the next best. Its authors wrote that it would improve simply by waiting for faster GPUs. Nvidia now calls it the start of the modern AI era.

Then the ASIC comes back 10:09

But once a job is huge and settled, the ASIC argument returns. In 2013, Google worked out that if people used voice search for three minutes a day, it would need to double its data centres. So it built an AI ASIC, the TPU, in fifteen months. At its heart is a grid of sixty-five thousand multipliers, all able to work at once, every tick. Running trained networks, it was fifteen to thirty times faster than the chips of its day, and thirty to eighty times better per watt. It was announced in 2016. Meanwhile, Nvidia's H100 barely does graphics at all.

All of the above 10:52

So back to your phone. It has a processor for whatever an app decides to do, a GPU for graphics, and a neural engine for AI. Apple's 2017 chip already had a neural engine doing up to six hundred billion operations a second. Other blocks handle jobs like video. Each job goes, as far as possible, to the silicon that can do it with the least energy. So which is best? All of them, each in its place.

What's stopping it? 11:25

And when one chip isn't enough, you fill a building with them. But Horowitz's conclusion was blunt: what limits computers now isn't how many transistors we can make, it's power. And every one of these chips turns its electricity into heat. So where does all this go next, and what's stopping it? That's next time.

Sources

Every factual claim in the episode is tied to one of these. Spotted an error? Tell us.

  1. Apple Newsroom: "Apple debuts iPhone 16 Pro and iPhone 16 Pro Max", Sept 2024 (A18 Pro) — apple.com
  2. Intel, MCS-4 Micro Computer Set data sheet, November 1971 (scan via Bitsavers; OCR'd) — bitsavers.org
  3. Computer History Museum, Silicon Engine: "1971: Microprocessor Integrates CPU Function… — computerhistory.org
  4. IEEE Spectrum: "Chip Hall of Fame: Intel 4004 Microprocessor" (Busicom 141-PF;… — spectrum.ieee.org
  5. Intel, "Intel Core i9 Processor 14900K" specifications (ARK) — intel.com
  6. M. Horowitz, "1.1 Computing's energy problem (and what we can do about it)", ISSCC… — gwern.net
  7. R. Hameed, W. Qadeer, M. Wachs, O. Azizi, A. Solomatnikov, B. C. Lee, S. Richardson,… — seas.upenn.edu
  8. B. Bailey, "What Will That Chip Cost?", Semiconductor Engineering, 30 Oct 2023 (IBS… — semiengineering.com
  9. IEEE Spectrum: "The FPGA Chip Is an IEEE Milestone" (Xilinx XC2064, 1985; Freeman; uses) — spectrum.ieee.org
  10. M. B. Taylor, "Bitcoin and the Age of Bespoke Silicon", CASES 2013 — cseweb.ucsd.edu
  11. M. B. Taylor, "The Evolution of Bitcoin Hardware", IEEE Computer 50(9), 2017 — cseweb.ucsd.edu
  12. K. Shirriff, "Reverse-engineering the first FPGA chip, the XC2064", Righto, Sept 2020 — righto.com
  13. I. Kuon and J. Rose, "Measuring the Gap Between FPGAs and ASICs", IEEE Trans. CAD… — eecg.utoronto.ca
  14. Microsoft Research: "Project Catapult" — microsoft.com
  15. Nvidia blog: "What's the Difference Between a CPU and a GPU?" (K. Krewell, Dec 2009;… — blogs.nvidia.com
  16. Nvidia Technical Blog: "NVIDIA Hopper Architecture In-Depth" (H100: SMs, FP32 cores,… — developer.nvidia.com
  17. Nvidia: Corporate timeline (1999 GPU; 2006 CUDA; 2012 AlexNet) — nvidia.com
  18. Nvidia, CUDA C Programming Guide (v12.4), §4.1 "SIMT Architecture" (warps of 32… — docs.nvidia.com
  19. Nvidia press release via Computer Graphics World, "Nvidia Introduces CUDA Architecture… — cgw.com
  20. Nvidia Developer Forums: "CUDA 1.0 Released", 26 June 2007 — forums.developer.nvidia.com
  21. R. Raina, A. Madhavan, A. Y. Ng, "Large-scale Deep Unsupervised Learning using… — mlanthology.org
  22. P. Warden, "Why GEMM is at the heart of deep learning", 20 April 2015 — petewarden.com
  23. N. P. Jouppi et al., "In-Datacenter Performance Analysis of a Tensor Processing Unit",… — arxiv.org
  24. A. Krizhevsky, I. Sutskever, G. E. Hinton, "ImageNet Classification with Deep… — papers.nips.cc
  25. N. Jouppi, "Google supercharges machine learning tasks with TPU custom chip", Google… — cloud.google.com
  26. Apple Newsroom: "The future is here: iPhone X", 12 Sept 2017 (A11 Bionic neural engine) — apple.com

Researched and scripted with AI assistance, fact-checked claim by claim, with synthetic narration and diagrams drawn in code. How we make episodes.

More from Silicon