半导体股暴跌40%后,重新审视算力需求的基本面
- 核心观点:当前AI与半导体领域30-40%的股价下跌并非泡沫破裂,而是建立在对算力需求有限的误判之上。文章认为,从政府军事、科学研究到企业和个人,各群体对算力的需求本质上是无限的,且AI具备递归需求特性,即算力本身会催生更多算力需求,而非传统基础设施的线性增长。
- 关键要素:
- 需求主体多元且支付意愿分层:政府(军事)、科学家(研究)、企业(降本扩张)、个人(赋能)均为算力买家,不同群体愿意支付的价格天花板不同,其中主权政府和云厂商的支出不可商议。
- 算力需求具有递归性:与铁路或互联网不同,AI算力能创造自身的需求增长。找到正ROI用例的用户会持续增加算力投入,使其成为“赚钱机器”,而非像传统基建那样依赖外部扩散。
- 超大规模云厂商定价能力极强:公有云算力成本被标高10-20倍,并通过锁定效应(如数据出口费)控制客户。作者自建机器的成本分析显示,云GPU按需定价在4.6个月内即可回本,体现了其高利润率。
- 算力订单积压显示强现实需求:全球算力订单积压从2025年初的5000亿美元增长至超2万亿美元,且主要由客户(非实验室)驱动,表明市场存在结构性转变,而非单纯泡沫。
- 硬件资产可能升值而非贬值:因内存成本随时间上涨,5年前的GPU租赁成本不降反升。硬件被概念化为“计算(贬值)+内存(升值)”的双组件游戏,整体投资回报率可能远超预期。
- 实验室与开源模型竞争不直接威胁硬件需求:即便开源模型(如Kimi K3)出现,其巨大内存需求(1.5-2TB)仍需大量硬件支持。推理服务利润率高达50-70%,且主要工作负载为推理,需求远超供应。
Original Author: Kerman Kohli
Original Translation: Deep Tide TechFlow
Introduction: After semiconductor stocks plummeted 30-40%, many declared the AI bubble had burst. However, this judgment rests on a fatal assumption: that demand for compute is finite. From government military applications and scientific research to enterprise products and personal use, every group is vying for compute power, and each is willing to pay a different price ceiling. More critically, AI possesses a characteristic that other infrastructures lack—recursive demand: compute power itself begets more demand for compute. When the order backlog for compute from hyperscale cloud providers surges from $500 billion to $2 trillion, calling this a bubble requires stronger evidence.

Core Question: Finite or Infinite Demand?
As of the publication date of this article, semiconductor and other AI/momentum-related stocks have fallen 30-40% from their all-time highs.
Many are quick to call this the top for semiconductors/AI/memory and celebrate their lack of participation.
They may be celebrating far too early.
In my view, the semiconductor/AI investment thesis boils down to this single question:
"Do you believe demand for compute is finite or infinite?"
In conversations, I see too many people trapped in their local experiences of enterprise usage/adoption and extrapolating that to the broader market. I think the argument that enterprise adoption will take longer could indeed hold true.
However, this also creates a reverse incentive for smaller, more AI-native companies to defeat incumbents with fewer employees, because AI, if used correctly, can be far cheaper than scaling with human labor.
Regardless, let's zoom out and stop viewing AI CapEx solely as enterprise demand. From a high-level perspective, building AI is about bringing compute online.
Compute, in turn, can and will be used by the following groups:
- Governments for military and defense purposes
- Scientists for medical and other cutting-edge research
- Enterprises to build new products and expand without labor constraints
- Individuals to empower themselves to do more (coding, designing, creating, asking questions)
These use cases and buyers are each willing to pay different prices for compute. While some may be priced out, it is unwise to assume that others with higher budgets won't step in. For some, compute expenditure is non-negotiable because it's a cutthroat race. Examples include sovereign states and hyperscale cloud providers. For others, compute is a substitute for labor costs, and it remains significantly cheaper (no legal overhead, management time, etc.).
Believing we are "overbuilding" or "building excess capacity" implies that any of the aforementioned groups have reached a final state and are satisfied with the status quo. Simply put, it assumes:
- Governments believe they don't need smarter weapons and defense capabilities
- Scientists are satisfied with the volume of research completed
- Enterprises think they've done enough with their product lines and don't want to grow further
- Individuals have reached a final state of curiosity and don't want to do more
If you truly believe any of the above, then you are correct in saying AI is a bubble and that there will be overinvestment.
Each group is willing/able to pay a different price for compute, but the market will organize around demand to ensure quality/pricing can be met. The misconception that intelligence is too expensive for all groups fails to recognize that those who can profitably orchestrate intelligence will continue to drive demand for it.
However, if you believe that humanity and the aforementioned groups will never be satisfied, then you must believe that demand for compute is infinite. We are currently in a great race for compute power, and few realize it.
Deeping the Inquiry: What is the ROI?
After spending more time in the market, this is the biggest concern for investors regarding sustained compute buildout. Is the trend of hyperscale cloud providers spending all their free cash flow excessive, and are they betting their futures on it?
Beyond the CapEx of hyperscalers, there is also significant focus on the revenues and profits generated by large labs. The threat of open-source models diminishing the value that labs can extract from frontier models complicates the picture further.
I'll take the time to discuss each point, but let's start with hyperscaler CapEx.
For those who think hyperscalers have miscalculated, I need you to understand this is far from the truth: public cloud services are outrageously expensive, and they know exactly how to extract every penny from you. They have convinced an entire generation of companies that they cannot scale without them.
In return, they mark up the cost of standard non-CPU compute by 10-20x. Additionally, you pay for logs, data transfer (egress), and five other services just to get basic tasks done.
This game works because they lock you into their system. Bandwidth within the GCP/AWS kingdom is cheap, but it skyrockets as soon as you try to move out. For many hyperscalers, they need customers to stay within their ecosystem, or they risk losing business. Running out of compute capacity is fatal for their survival. When your customer's data and compute are with you, saying your GPUs are exhausted is completely unacceptable and forces them to slowly migrate to your competitors. Hyperscalers have created an interesting dynamic where they can charge customers whatever price they want, and customers are largely powerless. Their tenants are so abstracted from bare metal that there are massive lock-in effects, making migration a multi-year effort (if possible at all).
A simple example illustrates how crazy this is. I wrote a few months ago about how I built this $15,000 machine:
Building an AI Inference Machine

Figure: Author's self-built AI inference machine. Source: Kerman Kohli / Substack
You can find a similar machine on GCP with the same GPU but lower specs:
https://cloud.google.com/products/compute/pricing/accelerator-optimized
The on-demand cost is roughly $3,248 per month, and with a 3-year commitment, it's $1,444 per month.

Figure: Example of Google Cloud GPU instance pricing. Source: Google Cloud
My machine only has 128GB of DDR5, but the Google Cloud instance has 180GB of some memory type (they don't tell you if it's DDR4 or DDR5, haha).
Quick calculation:
- At on-demand pricing, my machine would pay for itself in 4.6 months
- With a three-year commitment, the payback period is 10 months
The math for other machines is similar. An H200 cluster (GPU released in late 2024) pays for itself in under 2 years. No matter where you look, you see very similar math. This doesn't account for land costs, ongoing electricity bills, financing costs, or on-site staff, but it should serve as an illustrative example of how hyperscalers know how to price at a massive premium. Of course, spot pricing vs. committed pricing diverges again.
What makes this math even crazier is that GPUs released 5 years ago are: a) Holding their value, b) Seeing rental costs increase!
Hardware is very likely not going to depreciate but will instead appreciate from here on out. While new chips are being released that are more compute-efficient, the efficiency they lack is compensated by the rising cost of memory.
I conceptualize hardware as a two-component game where one component (compute) technically becomes less valuable over time, but this is offset by the other component (memory), which becomes more valuable over time.
If that's the case, the ROI on their CapEx is far higher than anyone remotely expects. Regarding the credit risk of hyperscalers, this tweet from Gavin Baker sums it up well:

Figure: Image accompanying Gavin Baker's tweet. Source: X / @GavinSBaker
Now you might say, how do we know this demand is sufficient? I mean, I can't model every scenario for every customer, but at some point, you need the humility to say that people lining up to pay is the strongest signal, and you should trust that they are rational actors spending on positive ROI endeavors.
If we look at it from this perspective, the order backlog grew from $500 billion in early 2025 to well over $2 trillion in just 1.5 years. When this is customer-driven, saying it's all fake or non-ROI becomes a stretch. Now, the counterargument is that labs make up a large portion of this. But this view is incorrect. According to EpochAI, frontier labs constitute a part but not the whole story of global compute demand.

Figure: Compute order backlog growing from $500 billion to $2 trillion. Source: EpochAI

Figure: Composition of global compute demand. Source: EpochAI
Regardless of what you believe, the fact of a $2 trillion order backlog should indicate something. Arguing that trillions in spending do not reflect a structural shift but rather an excessive bubble is an interesting take.
Many investors like to reason by analogy, comparing it to the internet buildout, railroads, or past infrastructure projects. I understand the logic here, but it misses a key feature of AI: recursive demand.
For railroads or the internet, you need more people to adopt the technology, and then ensure a ceiling on each person's usage to guarantee sufficient diffusion in the economy. AI doesn't have these dynamics. In this race, compute can generate its own demand for compute; the compute ceiling for an individual or organization is effectively infinite. If you find a useful, positive ROI use case, you can continue investing in compute, and it becomes a money-making machine.
The tricky part is that different people have vastly different experiences using AI. Most of the world uses it as a single-prompt Q&A machine. For people like me, as agentic engineering becomes more capable, they are becoming indispensable, enabling me to do more.
My compute spending continues to rise and will keep rising because I discover more positive ROI use cases. Regardless of the revenue AI generates, the cost savings it creates are undeniable, and this drives the case for most end customers.
The Uncertainty Factor: OpenAI / Anthropic
Continuing from the previous point, we can see the demand backlog is insane. But how real is the demand from labs (a significant portion of compute demand)? This is where I think the answer is less clear but still inferable. I want to break this answer down into inference and training.
If we consider that a new SOTA (State Of The Art) model costs hundreds of millions of dollars, we can say it is an investment asset that generates some useful lifetime value over time through inference (despite a steep depreciation curve).
As a counterforce, open-source models diffuse into the market and compete with frontier labs for compute at cheaper costs. These open-source models may or may not be distilled, but that's not critical for understanding the dynamics.
So the dynamic we must question is: what happens when a SOTA model is released? The reality is that not everyone will use them for every problem. However, given their SOTA capabilities, they can solve problems that current model classes cannot, and people are willing to pay a premium for this.
You could argue that models like Kimi K3 change this dynamic because they are open-source, but people forget an important fact: SOTA models are massive, and the hardware required to run them far exceeds what any home model can do. Kimi K3 itself requires close to 1.5TB - 2TB of memory. Good luck finding that.
What makes the model situation even more interesting is that Kimi ran out of capacity shortly after opening the gates to K3. Sure, the model exists, but someone still needs to serve it. It still needs to run on capable hardware. Labs will indeed be forced to become more competitive over time, but this won't threaten their business because the premium for certain workloads will persist. Furthermore, having the capacity to serve that model for your workload is equally important. It's disingenuous to say that frontier models are not worth any premium. How large that premium is remains to be seen.
If open-source models were banned or made illegal, labs would win big at the expense of innovation.
Inference has proven to be profitable, with service provider margins ranging between 50% - 70%. Even if people leave hosted providers, this demand must flow to them buying their own hardware. Given that inference is the dominant workload, demand far exceeds supply.
So how are our large labs performing in this world? I think the answer is probably "okay," but profit margins might not be that high.
It would be wrong to think they will fail and collapse. While I would love to offer a more data-supported view, we don't have clear data on their specific profit margins; however, it's reasonable to infer that their ability to optimize inference services has reached industry standards. Furthermore, companies with tens or even hundreds of millions of monthly active users are not blatant Ponzi schemes or fraudulent projects, and their revenues are growing rapidly.

Figure: Competition landscape between large labs and open-source models
Return on Investment for Lab Models
Labs might still not generate substantial returns on SOTA models, but this would imply that the market does not reward new models with stronger capabilities or is unwilling to pay a premium for


