半导体股暴跌40%后,重新审视算力需求的基本面
- 核心观点:当前AI与半导体领域30-40%的股价下跌并非泡沫破裂,而是建立在对算力需求有限的误判之上。文章认为,从政府军事、科学研究到企业和个人,各群体对算力的需求本质上是无限的,且AI具备递归需求特性,即算力本身会催生更多算力需求,而非传统基础设施的线性增长。
- 关键要素:
- 需求主体多元且支付意愿分层:政府(军事)、科学家(研究)、企业(降本扩张)、个人(赋能)均为算力买家,不同群体愿意支付的价格天花板不同,其中主权政府和云厂商的支出不可商议。
- 算力需求具有递归性:与铁路或互联网不同,AI算力能创造自身的需求增长。找到正ROI用例的用户会持续增加算力投入,使其成为“赚钱机器”,而非像传统基建那样依赖外部扩散。
- 超大规模云厂商定价能力极强:公有云算力成本被标高10-20倍,并通过锁定效应(如数据出口费)控制客户。作者自建机器的成本分析显示,云GPU按需定价在4.6个月内即可回本,体现了其高利润率。
- 算力订单积压显示强现实需求:全球算力订单积压从2025年初的5000亿美元增长至超2万亿美元,且主要由客户(非实验室)驱动,表明市场存在结构性转变,而非单纯泡沫。
- 硬件资产可能升值而非贬值:因内存成本随时间上涨,5年前的GPU租赁成本不降反升。硬件被概念化为“计算(贬值)+内存(升值)”的双组件游戏,整体投资回报率可能远超预期。
- 实验室与开源模型竞争不直接威胁硬件需求:即便开源模型(如Kimi K3)出现,其巨大内存需求(1.5-2TB)仍需大量硬件支持。推理服务利润率高达50-70%,且主要工作负载为推理,需求远超供应。
Original Author: Kerman Kohli
Original Compilation: Deep Tide TechFlow
Introduction: After semiconductor stocks plunged 30-40%, many declared the AI bubble had burst. However, this judgment is based on a fatal assumption: that demand for computing power is limited. From government military and scientific research to enterprise products and personal applications, every group is competing for computing power, and each has a different price ceiling they are willing to pay. More critically, AI has a characteristic that other infrastructure lacks—recursive demand: computing power itself creates more demand for computing power. When the backlog of computing power orders from hyperscale cloud vendors grows from $500 billion to $2 trillion, calling it a bubble requires stronger evidence.

Core Question: Infinite Demand or Finite Demand?
As of the time this article was published, semiconductor and other AI/momentum-related stocks had declined 30-40% from their all-time highs.
Many are quick to call this the top for semiconductors/AI/memory and are celebrating their non-participation.
They are likely celebrating far too early.
In my view, the semiconductor/AI investment thesis boils down to this single question:
"Do you believe the demand for computing power is finite or infinite?"
In discussions, I see too many people trapped in their local experience of enterprise usage/adoption, extrapolating it to the broader market. I think the argument that enterprise adoption takes longer might indeed hold true.
However, this also creates a perverse incentive for smaller, more AI-native companies to defeat them with fewer employees, because AI, if used correctly, can be far cheaper than scaling with human labor.
Regardless, let's zoom out and stop viewing AI capex solely as enterprise demand. At a high level, AI construction is about bringing computing power online.
Computing power, in turn, can and will be used by the following groups:
- Governments for military and defense purposes
- Scientists for medical and other frontier research
- Enterprises building new products and scaling without labor constraints
- Individuals empowered to do more (coding, designing, creating, asking questions)
These use cases and buyers each have different price points they are willing to pay for compute. While some may be priced out, it is unwise to think others won't have higher budgets. For some, compute spending is non-negotiable because it's a brutal race. Examples include sovereign governments and hyperscalers. For others, compute is a substitute for labor costs, and it's still far cheaper (no legal overhead, management time, etc.).
Believing we are "overbuilding" or "building beyond capacity" implies that any of the groups above have reached their final state and are satisfied with the status quo. Simply put, it believes:
- Governments think they don't need smarter weapons and defense capabilities.
- Scientists are satisfied with the amount of research completed.
- Enterprises believe they have done enough with their product lines and don't want to grow further.
- Individuals have reached a final state of curiosity and don't want to do more.
If you truly believe any of the above, then you are right to say AI is a bubble and there will be overinvestment.
The price each group is willing/able to pay for compute differs, but the market will organize around demand to ensure quality/price can be met. The misconception that intelligence is too expensive for everyone ignores that those who can orchestrate intelligence profitably will continue to drive demand for it.
But, if you believe that humanity and the groups above will never be satisfied, then you must believe that demand for computing power is infinite. We are currently in a massive race for compute power, and few realize it.
Deeping the Question: What is the ROI?
After spending more time in the market, this is the biggest concern investors have about sustained compute construction. Is the trend of hyperscalers burning through all their free cash flow excessive? Are they betting the farm?
Beyond the capex of hyperscalers, there is also significant focus on what revenue and profits the large labs are generating. The situation is complicated by new open-source models threatening the value large labs can extract from frontier models.
I'll spend time discussing these one by one, but let's start with the capex of hyperscalers.
To those who think hyperscalers have miscalculated, I need you to understand this is far from the truth: public cloud services are outrageously expensive, and they know exactly how to squeeze every penny out of you. They convinced an entire generation of companies that they couldn't scale without them.
In return, they mark up the cost of normal non-CPU computing power by 10-20 times. Additionally, you pay for logs, data transfer (egress), and 5 other services just to get basic things done.
This game works because they lock you into their system. Bandwidth within the GCP/AWS kingdom is cheap, but it skyrockets the moment you try to move out. For many hyperscalers, they need customers to stay within their ecosystem, or risk losing business. Not having enough compute is existential for them. When your customer's data and compute are with you, running out of GPUs is completely unacceptable and forces them to slowly move to your competitors. Hyperscalers have created an interesting dynamic where they can force customers to pay whatever price they want, and customers are largely helpless. Their tenants are so disconnected from bare metal that there's immense lock-in, making migration a multi-year effort (if possible).
Here's a simple example to illustrate how crazy they are with pricing. I wrote a few months ago about how I built this $15,000 machine:
Building an AI Inference Machine

Figure: Self-built AI inference machine by the author. Source: Kerman Kohli / Substack
You can find the same GPU, lower-spec machine on GCP:
https://cloud.google.com/products/compute/pricing/accelerator-optimized
On-demand cost is about $3,248 per month, and a 3-year reserved instance is about $1,444 per month.

Figure: Example Google Cloud GPU instance pricing. Source: Google Cloud
My machine only has 128GB of DDR5, but the Google Cloud one has 180GB of some type of memory (they don't tell you if it's DDR4 or DDR5, lol).
Quick math:
- At on-demand pricing, my machine pays for itself in 4.6 months.
- With a 3-year commitment, the payback period is 10 months.
The math for other machines is similar. An H200 cluster (a GPU released in late 2024) pays for itself in under 2 years. Everywhere you look, you see very similar math. This doesn't account for: land costs, ongoing electricity, financing costs, and on-site staff, but it should serve as an illustrative example of how hyperscalers know how to price at a massive premium. Of course, spot prices vs. committed prices diverge again.
What makes this math even crazier is that 5-year-old GPUs are: a) holding their value, b) *rising in rental cost*!
It is highly likely that hardware will not depreciate but will instead appreciate from here on out. While newer, more compute-efficient chips are being released, their efficiency gains are offset by rising memory costs.
I conceptualize hardware as a two-component game where one component (compute) technically becomes less valuable, but is offset by another (memory), which becomes more valuable over time.
If this is the case, the ROI on their capex is far higher than anyone is remotely expecting. Regarding the credit risk of hyperscalers, Gavin Baker's tweet sums it up well:

Figure: Image accompanying Gavin Baker's tweet. Source: X / @GavinSBaker
Now you might say, how do we know this demand is sufficient? I mean, I can't model every scenario for every customer, but at some point, you need the humility to say that people queuing up to pay is the strongest signal, and you believe they are rational actors spending money on positive ROI efforts.
If we look at it from this angle, the backlog of orders grew from $500 billion in early 2025 to well over $2 trillion in just 1.5 years. When this is customer-driven, it becomes a stretch to say it's all fake/not positive ROI. Now, the counterargument is that labs make up a huge portion of this, but that view is incorrect. According to EpochAI, frontier labs form a part, but not the entirety of global compute demand.

Figure: Compute order backlog growing from $500 billion to $2 trillion. Source: EpochAI

Figure: Composition of global compute demand. Source: EpochAI
Regardless of what you believe, the fact that there is a $2 trillion backlog should indicate something. Calling trillions in spending a reflection of an overheated bubble rather than a structural shift is an interesting opinion.
Many investors like to reason by analogy, comparing it to the building of the internet, railways, or past infrastructure projects. I understand the logic here, but it ignores a key feature of AI: recursive demand.
For railways or the internet, you need more people to adopt the technology, and then ensure a cap on how much each person uses it to have enough diffusion in the economy. AI doesn't have these dynamics. In this race, computing power can generate its own demand for computing power. The compute limit for an individual or organization is practically infinite. If you find a useful, positive-ROI use case, you can keep investing in compute, and it becomes a money-making machine.
The tricky part is that different people have very different experiences using AI. Most of the world uses it as a single-prompt Q&A machine. For people like me, as agentic engineering becomes more capable, these agents are becoming indispensable, enabling me to do more.
My compute spend continues to rise, and will continue to rise, as I discover more positive-ROI use cases. Regardless of the revenue AI generates, the cost savings it creates are undeniable, and this drives the case for most end customers.
Uncertain Factor: OpenAI / Anthropic
Continuing from the previous point, we can see that the demand backlog is insane. But how real is the demand from labs (a significant portion of compute demand)? This is where I think the answer is less clear but still analyzable. I want to break this answer down into inference and training.
If we consider that a new SOTA model costs hundreds of millions of dollars, we can say it is an investment asset that generates some useful lifetime value over time through inference (despite a steep depreciation curve).
As a counterforce, you have open-source models diffusing into the market and competing with frontier labs for compute at cheaper costs. These open-source models may or may not be distilled, but that detail is not crucial for understanding the dynamics.
So the dynamic we must question is: what happens when a SOTA model comes out? The reality is that not everyone will use them for every problem all the time. However, given their SOTA capabilities, they can solve problems that current model classes cannot, and you are willing to pay a premium for this.
You could argue that models like Kimi K3 change this dynamic because they are open-source, but people forget an important fact: SOTA models are massive, and the hardware required to run them far exceeds what any home-based model can do. Kimi K3 itself requires close to 1.5TB - 2TB of memory. Good luck finding that.
What makes the model situation more interesting is that Kimi ran out of capacity the moment it opened the gates for K3. Sure, the model exists out there, but someone still needs to serve it. It must run on capable hardware. Labs will indeed be forced to become more competitive over time, but this won't threaten their business because a premium for certain workloads will persist. Furthermore, the capacity to serve that model for your workload is equally important. It would be disingenuous to say frontier models are not worth any premium. How large this premium is remains to be seen.
If open-source models were banned or made illegal, then labs would win big at the expense of innovation.
Inference has proven profitable, with service provider margins ranging between 50% - 70%. Even as people move away from managed providers, this demand must flow to them buying their own hardware. Given that inference is the dominant workload, demand far exceeds supply.
So, how are our large labs performing in this world? I think the answer is probably okay, but margins might not be that high.
It is wrong to think they will fail and collapse. Though I would love to provide a view backed by more data, we lack clear data on their specific margins; however, it is reasonable to infer that their inference service optimization capabilities have reached industry standards. Furthermore, companies with tens of millions or even hundreds of millions of monthly active users are not outright Ponzi schemes or fraudulent projects, and their revenue is growing rapidly.

Figure: Competitive landscape between large labs and open-source models
Return on Investment for Lab Models
Labs may still fail to generate substantial returns on SOTA models, but this implies the market does not reward new models with greater capability and is unwilling to pay for that premium. Considering that frontier models represent the next class of problems AI can solve, betting against the frontier seems unwise. The cost of exiting this race is being harder to catch up later (unless through distillation).


