Thinking Freely with Nita Farahany

Thinking Freely with Nita Farahany

Why China Quit US Chips (Inside my AI Law and Policy Class #7)

When controlling compute failed to control AI

Nita Farahany's avatar
Nita Farahany
Sep 17, 2025
∙ Paid

January 27, 2025. My eldest daughter’s 10th birthday.

It was also the day Nvidia’s stock crashed 18%, wiping out $600 billion in market value. That’s about Sweden’s entire GDP, lost in a single day. Microsoft, Oracle—every major tech company—watched billions evaporate.

Why?

Seven days earlier, a Chinese startup called DeepSeek had done something that wasn’t supposed to be possible. They released an AI model more efficient than GPT-4, which they claimed to have built it for just $6 million. OpenAI spent over $100 million on GPT-4.

DeepSeek did this while being banned from buying the advanced chips everyone said were essential for AI. They couldn’t buy Nvidia's H100s. They couldn’t buy the weaker alternatives. They couldn’t even buy the machines that make the chips. (Breaking news we’ll get to later: China has just instructed its top tech companies to stop buying NVIDIA’s AI chips).

Welcome back to class. Every Monday and Wednesday this Fall you’re attending my AI Law & Policy class alongside my Duke Law students. They’re taking notes. You should be, too. And remember that this class is 85 minutes long. Take your time working through this material.

Want a live AI governance class? Check out Luiza Jarovsky, PhD‘s AI academy, still enrolling for November! Want a hot take on compute? Check out Michael Spencer’s post from today on AI Compute Warehouses.

Still here? Let’s dive into our fourth module on compute power.

By the end of class, you’re going to understand not just what’s happening in these compute fights, but why it matters for AI governance, your electricity bill, your access to AI, and whether we’re heading toward technological abundance or scarcity.

Share

Let’s Start with the Basics: What is “Compute”?

What is compute? Take a guess. Just shout it out (or if you’re reading this at work, whisper it quietly to yourself so people don’t think you’ve gone mad).

If you said something like “processing power” or “calculations,” you’re on the right track. But let’s be more precise. In the AI world, compute actually means three different things:

Loading...

If you said it was the math (the calculations which we will dive into below), the calculator (the physical hardware/chips doing the math), and the full infrastructure stack (data centers, cooling systems, and power plants), you got it right.

DeepSeek needed all of this, but they didn’t have access to the good calculators (the best chips), so they had to get creative with the math (the calculations) instead.

Understanding FLOPs and Why AI Needs So Much Math

Right now, I need all 8,200+ of you reading this to pull out your phones. Yes, I’m that professor who tells you to use your phone in class. Now open your calculator app.

Type in 3.14 times 2.5. But don’t hit equals yet!

Ready? On the count of three, everyone hit equals together.

One... two... three!

Congratulations! You just performed one FLOP—one Floating-point Operation. That’s how we measure compute, as FLOPs per second.

Your live counterparts in class today all did this at the same time—doing many FLOPs in parallel. The 8200+ of you at home? You just demonstrated the same principle. One calculation by one of you was one FLOP. And if you were all doing it at the same time, 8200+ FLOPs together.

Now, here’s where it gets fun. Set a timer for 10 seconds. See how many multiplication problems you can solve on your calculator app.

Keep it simple and just put in numbers—2 times 3, 5 times 7, whatever.

Go!

How many calculations did you do?

If you did 3-4, that means you are running at about 0.3 to 0.4 FLOPs.

Which means you’re a very slow computer, my friend.

An Nvidia H100 chip—the one DeepSeek desperately wanted but couldn’t buy—runs at about 67 trillion FLOPs. That’s 67,000,000,000,000 calculations per second. To put this in perspective, if you did one calculation per second, 24 hours a day, with no breaks and didn’t sleep, it would take you over 2 million years to do what an H100 does in one second

Here’s What’s Actually Happening When You Ask AI a Question

When you type “Write me a poem about compute,” into a LLM, the AI doesn’t understand these words—it converts them to math. Here’s the staggering amount of calculation required:

  • Step 1: Words become numbers Each word gets converted into about 1,000 numbers representing its meaning—like GPS coordinates, but instead of 3 numbers for location, you need 1,000 to capture all aspects of meaning. “Rose” might be [0.342, -0.891, 0.127...] encoding that it’s a flower, romantic, red, has thorns, appears in poems.

  • Step 2: To check if “poem” and “compute” relate, the AI multiplies their 1,000 numbers together and sums them up—2,000 calculations for one relationship. Your six-word prompt requires checking every word pair, which is already 30,000 calculations before generating anything.

  • Step 3: DeepSeek has 671 billion parameters—like 671 billion opinions about what word comes next. To pick even the first word, it must:

    • Consider all 170,000 possible English words

    • Check millions of parameters per word

    • Calculate scores for each

That’s 100 billion calculations for one word. A 100-word response? 10 trillion calculations.

Your laptop would take hours to do what AI chips do in seconds.

CPUs, GPUs, and TPUs

Think of a CPU—Central Processing Unit—like a brilliant senior law partner. This partner can handle any case, understands every nuance of the law, works through incredibly complex legal theories... but works alone and handles cases one at a time. That’s your laptop’s processor.

A GPU—Graphics Processing Unit—is like having 10,000 first-year law students. Each one is less legally sophisticated than the senior partner. They don’t know all the nuances of law and legal arguments. But they can simultaneously highlight documents.

For AI’s 170,000 possible next words, which approach wins? The GPUs.

But GPUs weren’t originally designed for AI! They were designed for video games. Really!

In the 1990s, Nvidia was making chips so teenagers could play Doom and Quake. Your screen has about 2 million pixels that need to be recalculated 60 times per second when you’re playing Call of Duty. That’s 120 million calculations per second just to render your game.

Around 2012, some researcher (probably engaged in useful mind wandering where big ideas are born) thought, “Wait a minute. The math needed for AI is basically the same as the math needed for graphics. Both need massive numbers of simple calculations done in parallel.”

And Nvidia pivoted. They started modifying their gaming chips for AI, and suddenly the company making toys for gamers became the most valuable company in the world.

Loading...

Google looked at this whole situation and said, “Wait, why would we use gaming chips someone else owns? That seems super inefficient. We’re going to build our own chips.” They created TPUs—Tensor Processing Units. (In April 2025, they released their 7th generation TPU, named Ironwood). These chips can ONLY do AI math. They don’t run Windows, they’re not designed to let you mine Bitcoin, nor do they render your graphics for Fortnite. Like hiring the best lawyer on the planet who does tax law—TPUs are great at that one thing, but not useful for your other legal work.

Nvidia also controls CUDA, the software that tells GPUs how to do AI calculations. It only works with Nvidia chips, which some see as a problematic monopoly. Even DeepSeek, despite being banned from buying Nvidia’s best hardware, still had to use CUDA software.

Why Didn’t China Just Build Its Own Chips for DeepSeek?

Maybe you’re thinking, “China makes iPhones, laptops, everything electronic. Why don’t they just make their own chips?”

Well, let’s take a walk along the supply chain from hell. You know, the one that could make this entire “AI bubble” burst.

Map Your Own Supply Chain

As we go through each layer of the supply chain, write it down. Then, draw arrows between them. By the end of these layers, you’ll see why the supply chain is impossible to replicate quickly.

  • Layer 1: Chip design. Nvidia designs the H100. This takes years and costs about $80 million just for the blueprint. (That was then, now they are developing even more sophisticated chips).

  • Layer 2: Software. That’s CUDA we just talked about.

  • Layer 3: Fabrication. Here’s where it starts to get crazy. Only one company on Earth can actually manufacture the most advanced chips. TSMC, the Taiwan Semiconductor Manufacturing Company. They have 90% market share. In other words, the entire global AI industry depends on factories on an island 100 miles from mainland China, which China claims as its territory, in one of the most geopolitically tense regions on Earth. If that doesn’t make you nervous, perhaps this will. TSMC’s most advanced fabs are concentrated in three science parks in Taiwan—Hsinchu, Taichung, and Tainan. All are on the western coast facing China. A single earthquake, blockade, or conflict could halt global chip production.

  • Layer 4: The machines that make the machines. Even TSMC can’t make chips without equipment from ASML, a company in the Netherlands. ASML makes these machines called EUV lithography systems. These machines cost $200 million each, are the size of a school bus, and use lasers to carve circuits onto silicon wafers. The circuits are so small—we’re talking 5 nanometers—that you could fit 15,000 of them across the width of a human hair. Under US pressure the Netherlands banned selling these machines to China.

  • Layer 5: Data center infrastructure. Even with chips, you need massive warehouses with cooling systems (chips generate enormous heat), reliable power (data centers use as much electricity as small cities), network infrastructure (to connect everything), physical security and redundancy, building a single data center takes 2-3 years and costs billions. China has data centers, but they’re optimized for different chips and different scales.

Loading...

DeepSeek was boxed in. They had no access to top chips, no access to weaker chips (the US banned exports of those too), no way to make the chips (TSMC can’t freely ship to China, either), and no way to buy chip-making equipment from the Netherlands. So… they could give up or get creative.

They chose creative.

The Mixture of Experts Revolution

Unable to match American compute power, DeepSeek had to completely reimagine how AI models work.

In the Duke classroom, I hypothetically assigned six students to serve as my “lawyers” on a question of how to best structure my estate planning for tax purposes. I want you to envision this, too, so grab six objects from your desk. Each one represents a lawyer you’ve hired for the same task.

In the traditional approach—the way GPT-4 works—when a client asks a tax question, all six lawyers research it. Criminal law question? All six lawyers work on that too. Corporate merger? Everyone’s on it.

But I assigned each of the students to a different area of expertise. And it immediately became clear – why would I pay six lawyers to work on the problem, when only ONE of them had an expertise in tax? Why not just have that one work on the problem and save my hard-earned money for something else?

That’s exactly what DeepSeek realized in making its AI more efficient.

Their approach, called Mixture of Experts, has specialized “experts” in their model, but only a handful active for any given query. Tax question? Only the tax expert wakes up. Criminal law? Only the criminal expert. The other experts remain dormant, saving compute.

DeepSeek’s V3 architecture, for example, has 256 “routed experts per layer” and only 8 are activated per token, along with one or two shared experts that handle general tasks.

One of your Duke counterparts—always the skeptic—asked, “Okay, but how does the model know which expert to activate?”

Great question! There’s a “router”—think of it as the world’s best legal secretary who reads the question and instantly knows which lawyers to call. This routing decision takes about 1% of the compute that activating all the experts at the same time would take.

Which means even with all the limitations on its access along the AI supply chain, DeepSeek achieved massive efficiency gains. So why isn’t America just copying this approach?

The Trillion-Dollar Bet

The truth is, we are copying it. GPT-4 reportedly uses Mixture of Experts. So do Google’s models. But we’re also throwing unprecedented amounts of money at infrastructure. And there’s a good reason for that.

On January 21, 2025, President Trump, alongside Sam Altman from OpenAI, Larry Ellison from Oracle, and Masayoshi Son from SoftBank announced Project Stargate, which is meant to be a $500 billion investment over four years to build AI infrastructure. That’s more than the entire Apollo moon program, adjusted for inflation.

These aren’t just server farms. The Stargate facility in Abilene, Texas, will run 2 million chips and require 4.5 gigawatts of power, which is more than twice what the entire Hoover Dam generates, and is enough energy to power 3.5 million homes.

But our electrical grid can’t provide as much power as these data centers demand. Grid connections take 5-7 years to build. These data centers need power by 2027. The math doesn’t add up.

So tech companies are getting desperate:

  • Microsoft is restarting the Three Mile Island nuclear plant

  • Google is drilling for geothermal energy (fun fact, our house in Durham was built with Geothermal)

  • Oracle is building small modular nuclear reactors

  • Amazon is buying up entire wind farms

We’re literally building power plants just to run AI models. Meanwhile, China just instructed its companies to stop buying Nvidia chips. They haven’t officially said why, but the timing is hard to ignore. Are they betting algorithmic efficiency beats raw hardware? Building their own chips? Playing geopolitical chess? We don’t know yet, but the symbolism is unmistakable.

Loading...

Why Inference Changes Everything

Everyone focuses on the one-time computing cost of training these models. But the ongoing compute expense is in inference—running the model millions of times per day for actual users. Because inference never stops.

ChatGPT has about 100 million daily users, but what does that actually mean in terms of compute and energy? A basic ChatGPT query requires about 100 billion calculations and uses 0.003 kWh of electricity—enough to run an LED bulb for 20 minutes. That seems tiny until you multiply it out. 100 million users making five queries each means 1.5 million kWh daily, the equivalent of powering 50,000 American homes.

But the real shock comes with advanced queries. When you use OpenAI’s reasoning model for complex problems, a single query can require 10 trillion calculations and burn through 3 kWh—your laptop’s entire day of power. The infamous “$1,000 query”? That’s 10 quadrillion calculations consuming 300 kWh, enough electricity to power an average home for 10 days.

This is why Sam Altman admitted OpenAI loses money on $200/month Pro subscriptions. One power user running 50 complex queries daily consumes as much electricity as five average homes while paying the price of a Netflix subscription. They pay $200 monthly but burn through $7,500 worth of compute.

DeepSeek’s Mixture of Experts and fewer experts activating per query seems like it should offer a massive energy savings, right?

But MIT Tech review showed that DeepSeek might not be so efficient, after all, at least when it comes to inferences. DeepSeek uses “chain of thought” reasoning, breaking problems into steps and working through them methodically. This produces better answers but much longer responses.

Early testing by researchers found DeepSeek actually used MORE energy per query than comparable models—not less—because it generates such detailed responses. A simple ethics question generated 1,000 words from DeepSeek, using about as much energy as streaming a 10-minute YouTube video.

Which means the efficiency breakthrough that briefly crashed the stock market on my daughter’s birthday might actually increase energy consumption if widely adopted.

Bringing us full circle on the inference trap. Training a model happens once (or infrequently through updates), while inference happens billions of times daily, forever. And unlike training costs which are dropping, inference scales with users—meaning the more successful AI becomes, the more unsustainable it gets.

Loading...

The DeepSeek Paradox

The US tried to slow or even cripple China’s AI development by denying them chips and access to the AI supply chain. Instead, we forced them to innovate. DeepSeek exists because of our restrictions, not despite them.

So export controls as a governance strategy may well backfire, depending on their objective. Infrastructure races accelerate regardless of efficiency. Open source makes control impossible. And we’ve built the entire AI industry on three buildings in Taiwan that could disappear in an earthquake or conflict.

Which is why on Monday, we grapple with even more difficult governance questions. If controlling hardware doesn’t control AI development, what governance tools could possibly work? If innovations in model efficiency leads to MORE energy use, not less, how do we prevent AI from consuming everything? We’ll dive into those issues, and more, on Monday.

Your Homework: The AI Power Check

  1. Share this post with ONE person you think should know more about AI Law and policy.

    Share

  2. Try this exercise:

  • Google: “[Your city] daily electricity use”

  • Then ask yourself: If everyone in your city used ChatGPT 10 times a day, that would need about 8,400 kWh for every 280,000 people.

  • Could your city handle it?

Post what you find in the comments. Which cities would struggle most?

The entire class lecture is above, but for those of you who want to support my work (thank you!) or go deeper in the class, the class readings, video assignments, and virtual chat-based office-hours details are below.

User's avatar

Continue reading this post for free, courtesy of Nita Farahany.

Or purchase a paid subscription.
© 2026 Nita Farahany · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture