半导体股暴跌40%后,重新审视算力需求的基本面
- 核心观点:当前AI与半导体领域30-40%的股价下跌并非泡沫破裂,而是建立在对算力需求有限的误判之上。文章认为,从政府军事、科学研究到企业和个人,各群体对算力的需求本质上是无限的,且AI具备递归需求特性,即算力本身会催生更多算力需求,而非传统基础设施的线性增长。
- 关键要素:
- 需求主体多元且支付意愿分层:政府(军事)、科学家(研究)、企业(降本扩张)、个人(赋能)均为算力买家,不同群体愿意支付的价格天花板不同,其中主权政府和云厂商的支出不可商议。
- 算力需求具有递归性:与铁路或互联网不同,AI算力能创造自身的需求增长。找到正ROI用例的用户会持续增加算力投入,使其成为“赚钱机器”,而非像传统基建那样依赖外部扩散。
- 超大规模云厂商定价能力极强:公有云算力成本被标高10-20倍,并通过锁定效应(如数据出口费)控制客户。作者自建机器的成本分析显示,云GPU按需定价在4.6个月内即可回本,体现了其高利润率。
- 算力订单积压显示强现实需求:全球算力订单积压从2025年初的5000亿美元增长至超2万亿美元,且主要由客户(非实验室)驱动,表明市场存在结构性转变,而非单纯泡沫。
- 硬件资产可能升值而非贬值:因内存成本随时间上涨,5年前的GPU租赁成本不降反升。硬件被概念化为“计算(贬值)+内存(升值)”的双组件游戏,整体投资回报率可能远超预期。
- 实验室与开源模型竞争不直接威胁硬件需求:即便开源模型(如Kimi K3)出现,其巨大内存需求(1.5-2TB)仍需大量硬件支持。推理服务利润率高达50-70%,且主要工作负载为推理,需求远超供应。
Original Author: Kerman Kohli
Original Translation: Deep Tide TechFlow
Foreword: After semiconductor stocks plunged 30-40%, many declared the AI bubble had burst. But this judgment rests on a fatal assumption: that demand for computing power is finite. From government military and scientific research to enterprise products and personal applications, all groups are competing for computing power, and each group has a different price ceiling they are willing to pay. More critically, AI has a characteristic that other infrastructure lacks—recursive demand: computing power itself generates more demand for computing power. When the backlog of computing power orders for hyperscale cloud vendors rises from $500 billion to $2 trillion, calling it a bubble requires stronger evidence.

Core Question: Infinite Demand or Finite Demand
As of the publication of this article, semiconductor and other AI/momentum-related stocks have fallen 30-40% from their all-time highs.
Many are willing to call this the top for semiconductors/AI/memory and celebrate their victory for not participating.
They are likely celebrating too early.
In my view, the semiconductor/AI investment thesis boils down to this single question:
"Do you believe the demand for computing power is finite or infinite?"
In discussions, I see too many people trapped in their local experience of enterprise usage/adoption and extrapolating it to the broader market. I think the argument that enterprise adoption will take longer may indeed hold true.
However, this also creates a perverse incentive for smaller, more AI-native companies to beat them with fewer employees, because AI, if used correctly, can be much cheaper than scaling with human labor.
Regardless, let's zoom out and stop viewing AI capital expenditure merely as enterprise demand. From a high-level perspective, AI construction is about bringing computing power online.
Computing power, in turn, can and will be used by the following groups:
- Governments for military and defense purposes
- Scientists for medical and other frontier research
- Enterprises to build new products and expand without labor constraints
- Individuals empowered to do more (coding, designing, creating, asking questions)
These use cases and buyers each have different willingness to pay for computing power. While some may be priced out, it is unwise to think that no one else has a higher budget. For some, computing expenditure is non-negotiable because it's a brutal race. Examples include sovereign governments and hyperscale cloud vendors. For others, computing power is a substitute for labor costs and is still much cheaper (no legal overhead, management time, etc.).
Believing we are "overbuilding" or "building beyond capacity" implies that any of the above groups have reached a final state and are satisfied with the status quo. Simply put, it believes that:
- Governments think they don't need smarter weapons and defense capabilities
- Scientists are satisfied with the amount of research already completed
- Enterprises believe they have done enough with their product lines and don't want to grow further
- Individuals have reached a final state of curiosity and don't want to do more
If you genuinely believe any of the above, then you are right that AI is a bubble with overinvestment.
Each group's willingness/ability to pay for computing power differs, but the market will organize around demand to ensure quality/price can be met. It's a misconception to think intelligence is too expensive for all groups, because those who can profitably orchestrate intelligence will continue to drive demand for it.
But, if you believe that humanity and the aforementioned groups will never be satisfied, then you must believe the demand for computing power is infinite. We are currently in a great race for computing power, and few realize it.
Follow-up: What is the ROI?
After spending more time in the market, this is the biggest concern for investors regarding the sustainability of computing power construction. Is the trend of hyperscale cloud vendors spending all their free cash flow excessive? Are they betting the farm?
Beyond the capital expenditure of hyperscale cloud vendors, there is also intense focus on what revenues and profits the large labs are generating. The threat of new open-source models to the value labs can extract from frontier models complicates matters further.
I'll take the time to discuss each, but let's start with the capital expenditure of hyperscale cloud vendors.
For those who think hyperscale cloud vendors have miscalculated, I need you to understand this is far from the truth: public cloud services are outrageously expensive, and they know how to squeeze every penny out of you. They convinced an entire generation of companies that they couldn't scale without them.
In return, they mark up the cost of standard non-CPU computing power by 10-20 times. Additionally, you pay for logging, data transfer (egress), and 5 other services just to do basic work.
This game works because they lock you into their system. Bandwidth within the GCP/AWS kingdom is cheap, but it spikes dramatically once you try to move out. For many hyperscale cloud vendors, they need clients to stay within their ecosystem, or they risk losing business. Not having enough computing power is fatal to their survival. When your client's data and computing are with you, saying your GPUs are maxed out is completely unacceptable and forces them to slowly migrate to your competitors. Hyperscale cloud vendors have created an interesting dynamic where they can force clients to pay whatever price they want, and clients are largely powerless. Their tenants are so decoupled from bare metal, with massive lock-in effects, that migrating is a multi-year effort (if possible at all).
A simple example to illustrate how crazy they are about this. I wrote a few months ago about how I built this $15,000 machine:
Building an AI Inference Machine

Figure: Author's self-built AI inference machine. Source: Kerman Kohli / Substack
You can find the same GPUs on GCP with lower-spec machines:
https://cloud.google.com/products/compute/pricing/accelerator-optimized
On-demand cost is approximately $3,248 per month, and with a 3-year commitment, it's $1,444 per month.

Figure: Example of Google Cloud GPU instance pricing. Source: Google Cloud
My machine only has 128GB of DDR5, but the Google Cloud one has 180GB of some type of memory (they don't tell you if it's DDR4 or DDR5 lol).
Quick calculation:
- At on-demand pricing, my machine pays for itself in 4.6 months
- The payback period with a three-year commitment is 10 months
The math is similar for other machines. An H200 cluster (GPUs released in late 2024) pays for itself in under 2 years. No matter where you look, you see very similar math. Of course, this doesn't account for: land costs, ongoing electricity bills, financing costs, and on-site staff, but it should serve as an illustrative example of how hyperscale cloud vendors know how to price at a huge premium. Of course, spot prices vs. committed prices again create variance.
Making this math even crazier is that 5-year-old GPUs are: a) holding their value b) their rental costs are rising!
It is highly likely that hardware will not depreciate but will appreciate from here. While new chips are more compute-efficient, the efficiency they lack is compensated by rising memory costs.
I conceptualize hardware as a two-component game, where one component (compute) technically becomes less valuable but is offset by another component (memory) which becomes more valuable over time.
If this is the case, the ROI on their capital expenditure is much higher than anyone remotely expects. Regarding the credit risk of hyperscale cloud vendors, this tweet from Gavin Baker summarizes it well:

Figure: Image accompanying Gavin Baker's tweet. Source: X / @GavinSBaker
Now you might say, how do we know this demand is sufficient? I mean, I can't model every scenario for every client, but at some point, you need the humility to say that people queuing up to pay is the strongest signal, and you believe they are rational actors spending on positive ROI efforts.
If we look at it from this perspective, the backlog of orders has grown from $500 billion in early 2025 to well over $2 trillion in just 1.5 years. When this is driven by customers, it becomes a stretch to say it's all fake/not positive ROI. Now, the counterpoint is that labs constitute a significant portion of this, but this view is wrong. According to EpochAI, frontier labs constitute a portion, but not the entirety of global computing demand.

Figure: Computing order backlog growing from $500 billion to $2 trillion. Source: EpochAI

Figure: Composition of global computing demand. Source: EpochAI
Regardless of what you believe, the fact of a $2 trillion backlog should indicate something. Calling trillions in spending a reflection of an excessive bubble rather than a structural shift is an interesting opinion.
Many investors like to reason by analogy, comparing it to the internet buildout, railroads, or past infrastructure projects. I understand the logic here, but it misses a key feature of AI: recursive demand.
For railroads or the internet, you need more people to adopt the technology, and then ensure a ceiling on each person's usage to guarantee enough diffusion in the economy. AI doesn't have these dynamics. In this race, computing power can generate its own demand; the computing limit for an individual or organization is effectively infinite. If you find a useful, positive ROI use case, you can keep investing in computing power, and it becomes a money-making machine.
Where this gets very tricky is that different people have very different experiences using AI. Most people in the world use it as a single-prompt Q&A machine. For people like me, as agentic engineering becomes more capable, it's becoming indispensable, enabling me to do more.
My computing expenditure continues to rise and will continue to rise as I discover more positive ROI use cases. Regardless of how much revenue AI generates, the cost savings it creates are undeniable, which drives the case for most end customers.
Uncertainty Factor: OpenAI / Anthropic
Building on the previous point, we can see the demand backlog is insane. But how real is the demand from labs (a fairly significant portion of computing demand)? Now this is where I think the answer is less clear but still reasonable to reason about. I want to break this answer down into inference and training.
If we assume a new SOTA model costs hundreds of millions of dollars, we can say it's an investment asset that generates some useful lifecycle value over time through inference (despite a steep depreciation curve).
As a counterforce, you have open-source models diffusing into the market and competing with frontier labs for computing at cheaper costs. These open-source models may or may not be distilled, which is not critical for understanding the dynamics.
So the dynamic we must question is what happens when a SOTA model comes out? The reality is, not everyone will use them for every problem all the time. However, given their SOTA capabilities, they can solve problems that current model classes cannot, and you are willing to pay a premium for this.
You could say models like Kimi K3 change this dynamic because they are open-source, but people forget an important fact: SOTA models are very large, and the hardware required to run them far exceeds what any home model can do. Kimi K3 itself requires close to 1.5TB - 2TB of memory. Good luck finding that.
What makes the model more interesting is that Kimi ran out of capacity immediately after opening the gates to K3. The model exists out there, but someone still needs to serve it. It must still run on capable hardware. Labs will indeed be forced to become more competitive over time, but this won't threaten their business because, ultimately, the premium for certain workloads will persist. Furthermore, having the capacity to serve that model for your workloads is equally important. It is disingenuous to think frontier models aren't worth any premium. How large this premium is remains to be seen.
If open-source models were banned or made illegal, then labs would win big at the expense of innovation.
Inference has been proven profitable, with service provider margins between 50% - 70%. Even if people leave managed providers, this demand must flow to them buying their own hardware. Given that inference is the dominant workload, demand far exceeds supply.
Now, as for our large labs, how are they performing in this world? I think the answer is probably fine, but margins might not be sky-high.
It would be wrong to think they will fail and collapse. Although I wish I could offer an opinion backed by more data, we lack clear data on their specific profit margins; however, it's reasonable to infer that their inference service optimization has reached industry standards. Furthermore, companies with tens or hundreds of millions of monthly active users are not outright Ponzi schemes or frauds, and their revenues are growing rapidly.

Figure: Competitive landscape of large labs vs. open-source models
Return on Investment for Lab Models
Labs may still not generate substantial returns on SOTA models, but this would mean the market does not reward new models with stronger capabilities nor is willing to pay for this premium. Considering frontier models represent the next class of problems AI can solve, betting against the frontier seems like a bad idea. Exiting this race carries the cost of being harder to catch up later (unless through distillation).
Regardless of the SOTA


