半导体股暴跌40%后,重新审视算力需求的基本面
- 核心观点:当前AI与半导体领域30-40%的股价下跌并非泡沫破裂,而是建立在对算力需求有限的误判之上。文章认为,从政府军事、科学研究到企业和个人,各群体对算力的需求本质上是无限的,且AI具备递归需求特性,即算力本身会催生更多算力需求,而非传统基础设施的线性增长。
- 关键要素:
- 需求主体多元且支付意愿分层:政府(军事)、科学家(研究)、企业(降本扩张)、个人(赋能)均为算力买家,不同群体愿意支付的价格天花板不同,其中主权政府和云厂商的支出不可商议。
- 算力需求具有递归性:与铁路或互联网不同,AI算力能创造自身的需求增长。找到正ROI用例的用户会持续增加算力投入,使其成为“赚钱机器”,而非像传统基建那样依赖外部扩散。
- 超大规模云厂商定价能力极强:公有云算力成本被标高10-20倍,并通过锁定效应(如数据出口费)控制客户。作者自建机器的成本分析显示,云GPU按需定价在4.6个月内即可回本,体现了其高利润率。
- 算力订单积压显示强现实需求:全球算力订单积压从2025年初的5000亿美元增长至超2万亿美元,且主要由客户(非实验室)驱动,表明市场存在结构性转变,而非单纯泡沫。
- 硬件资产可能升值而非贬值:因内存成本随时间上涨,5年前的GPU租赁成本不降反升。硬件被概念化为“计算(贬值)+内存(升值)”的双组件游戏,整体投资回报率可能远超预期。
- 实验室与开源模型竞争不直接威胁硬件需求:即便开源模型(如Kimi K3)出现,其巨大内存需求(1.5-2TB)仍需大量硬件支持。推理服务利润率高达50-70%,且主要工作负载为推理,需求远超供应。
Original Author: Kerman Kohli
Original Translation: TechFlow
Introduction: After semiconductor stocks plunged 30-40%, many declared the AI bubble had burst. However, this judgment rests on a fatal assumption: that demand for computing power is finite. From government military applications and scientific research to enterprise products and personal use, every group is competing for computing power, and each has a different price ceiling they are willing to pay. More critically, AI possesses a characteristic absent in other infrastructure—recursive demand: computing power itself creates more demand for computing power. When the backlog of computing power orders from hyperscale cloud providers surges from $500 billion to $2 trillion, calling it a bubble requires stronger evidence.

Core Question: Infinite Demand or Finite Demand
As of the time of this article's publication, semiconductor and other AI/momentum-related stocks have fallen 30-40% from their all-time highs.
Many are quick to call this the top for semiconductors/AI/memory and celebrate their lack of participation.
They are likely celebrating far too early.
In my view, the semiconductor/AI investment thesis boils down to this single question:
"Do you believe the demand for computing power is finite or infinite?"
In discussions, I see too many people trapped in their local experience of enterprise usage/adoption, generalizing it to the broader market. I agree that the argument for longer enterprise adoption timelines might hold true.
However, this also creates a perverse incentive for smaller, more AI-native companies to beat them with fewer employees, because AI, if used correctly, is far cheaper than scaling with human labor.
Regardless, let's zoom out and stop viewing AI capital expenditure solely as enterprise demand. From a high-level perspective, AI construction is about bringing computing power online.
This computing power, in turn, can and will be utilized by:
- Governments for military and defense purposes.
- Scientists for medicine and other cutting-edge research.
- Enterprises building new products and expanding without labor constraints.
- Individuals empowered to do more (coding, designing, creating, questioning).
These use cases and buyers each have different willingness to pay for computing power. While some may be priced out, it's unwise to assume others won't have higher budgets. For some, computing power expenditure is non-negotiable because it's a ruthless race. Examples include sovereign states and hyperscalers. For others, computing power is a substitute for labor costs and remains significantly cheaper (without legal overhead, management time, etc.).
Believing we are "overbuilding" or "building beyond capacity" implies any of the aforementioned groups have reached their final state and are satisfied with the status quo. Simply put, it believes:
- Governments believe they don't need smarter weapons and defenses.
- Scientists are satisfied with the amount of research already completed.
- Enterprises think they've done enough with their product lines and don't want to grow further.
- Individuals have reached the final state of curiosity and don't want to do more.
If you truly believe any of the above, then you are correct that AI is a bubble and there will be overinvestment.
Each group's willingness/ability to pay for computing power differs, but the market will organize around demand to ensure quality/price meets expectations. Believing intelligence is too expensive for all groups is a misunderstanding, as those who can profitably orchestrate intelligence will continue to drive demand for it.
But if you believe humanity and these groups will never be satisfied, you must believe the demand for computing power is infinite. We are currently in a great race for computing power, and few realize it.
Probing Deeper: What is the ROI?
Having spent more time in the market, this is the biggest concern for investors regarding sustained computing power construction. Is the trend of hyperscalers burning through all their free cash flow excessive? Are they betting the farm?
Beyond hyperscaler CapEx, there's also significant focus on the revenue and profits large labs generate. The threat of new open-source models eroding the value that labs can extract from frontier models complicates matters further.
I'll address these one by one, but let's start with hyperscaler capital expenditure.
For those who think hyperscalers miscalculated their numbers, I need you to understand this is far from the truth: public cloud services are exorbitantly priced, and they know exactly how to extract every penny from you. They have convinced an entire generation of companies that they cannot scale without them.
In return, they mark up the cost of standard non-CPU computing power by 10-20 times. Additionally, you pay for logs, data transfer (egress), and five other services just to get basic tasks done.
This game works because they lock you into their system. Bandwidth within the GCP/AWS kingdom is cheap, but it skyrockets as soon as you try to move out. For many hyperscalers, keeping customers within their ecosystem is critical, or they risk losing business. Not having enough computing power is fatal for their survival. Saying your GPUs are fully utilized is completely unacceptable when your customers' data and compute are with you; it would gradually force them to move to your competitors. Hyperscalers have created an interesting dynamic where they can charge customers virtually any price, and customers are largely powerless. Their tenants are so disconnected from bare metal, with enormous lock-in effects, that migration becomes a multi-year effort (if possible at all).
A simple example illustrates how insane this is. I wrote months ago about building this $15,000 machine:
Building an AI Inference Machine

Figure: Author's self-built AI inference machine. Source: Kerman Kohli / Substack
You can find the same GPU, but lower spec machine on GCP:
https://cloud.google.com/products/compute/pricing/accelerator-optimized
On-demand cost is approximately $3,248 per month, and a 3-year reserved instance costs $1,444 per month.

Figure: Example of Google Cloud GPU instance pricing. Source: Google Cloud
My machine has only 128GB DDR5, but the Google Cloud one has 180GB of some type of memory (they don't tell you if it's DDR4 or DDR5 lol).
Quick calculation:
- At on-demand pricing, my machine pays for itself in 4.6 months.
- With a three-year commitment, the payback period is 10 months.
The math is similar for other machines. An H200 cluster (a GPU released in late 2024) pays for itself in under 2 years. Wherever you look, you see very similar math. This excludes land costs, ongoing electricity, financing costs, and on-site staff, but it should serve as an illustrative example of how hyperscalers know how to price at a massive premium. Of course, spot prices vs. committed prices create further divergence.
What makes this math even more insane is that GPUs from 5 years ago are: a) holding their value, and b) their rental costs are increasing!
It is highly likely that hardware will not depreciate but will appreciate from this point onward. While new chips with better compute efficiency are being released, their lack of efficiency in other areas is compensated by rising memory costs.
I conceptualize hardware as a two-component game, where one component (compute) technically becomes less valuable over time, but is offset by another component (memory) that becomes more valuable over time.
If this is the case, the ROI on their capital expenditure is much higher than anyone remotely expects. Regarding the credit risk of hyperscalers, this tweet from Gavin Baker summarizes it well:

Figure: Accompanying image to Gavin Baker's tweet. Source: X / @GavinSBaker
Now you might say, how do we know this demand is sufficient? I can't model every scenario for every customer, but at some point, you must be humble enough to see that people lining up to pay is the strongest signal, and you should believe they are rational actors spending on positive-ROI endeavors.
If we look at it from this perspective, the backlog of orders has grown from $500 billion in early 2025 to well over $2 trillion in just 1.5 years. When this is driven by customer demand, calling it all fake or non-positive-ROI becomes a stretch. Now, the counterargument is that labs comprise a large portion of this, but that view is incorrect. According to EpochAI, frontier labs constitute a part, but not the entirety of global computing power demand.

Figure: Computing power order backlog grows from $500 billion to $2 trillion. Source: EpochAI

Figure: Composition of global computing power demand. Source: EpochAI
Regardless of what you believe, the fact of a $2 trillion order backlog should indicate something. Believing trillions in spending do not reflect a structural shift but rather an excessive bubble is an interesting viewpoint.
Many investors like to reason by analogy, comparing it to the internet buildout, railroads, or past infrastructure projects. I understand the logic here, but it misses a key feature of AI: recursive demand.
For railroads or the internet, you need more people to adopt the technology and then ensure a ceiling on each person's usage to guarantee sufficient diffusion in the economy. AI doesn't have these dynamics. In this race, computing power can generate its own demand. For an individual or organization, the constraint on computing power is effectively infinite. If you find a useful, positive-ROI use case, you can continue to invest in computing power, and it will become a money-making machine.
Where this gets very tricky is that different people have very different experiences using AI. Most people in the world use it as a single-turn Q&A machine. For people like me, as agentic engineering becomes more capable, tools are becoming indispensable, enabling me to do more.
My computing power expenditure continues to rise and will keep rising as I discover more positive-ROI use cases. Regardless of the revenue AI generates, the cost savings it creates are undeniable, and this is the primary driver of the case for most end customers.
Uncertain Factor: OpenAI / Anthropic
Continuing from the previous point, we can see that demand backlog is insane. But how real is the demand from labs (a substantial portion of computing power demand)? This is where I think the answer is less clear but still debatable. I want to break this answer down into inference and training.
If we assume a new state-of-the-art (SOTA) model costs hundreds of millions of dollars, we can say it's an investment asset that generates some useful lifetime value over time through inference (despite a steep depreciation curve).
As a counterforce, you have open-source models diffusing into the market and competing with frontier labs for computing power at cheaper costs. Whether these open-source models are distilled or not isn't crucial for understanding the dynamics.
So, the dynamic we must question is: what happens when a SOTA model is released? The reality is not everyone will use them for every problem all the time. However, given their SOTA capabilities, they can solve problems that current model classes cannot, and you are willing to pay a premium for that.
You might say models like Kimi K3 change this dynamic because they are open-source, but people forget an important fact: SOTA models are enormous, and the hardware required to run them far exceeds what any consumer model can handle. Kimi K3 itself requires close to 1.5TB - 2TB of memory. Good luck finding that.
What makes models more interesting: Kimi ran out of capacity right after opening the gates for K3. Of course, the model exists, but someone still needs to serve it. It still must run on capable hardware. Labs will indeed be forced to become more competitive over time, but this won't threaten their business because eventually, premiums for certain workloads will persist. Furthermore, having the capacity to serve that model for your workload is equally important. It is disingenuous to claim that frontier models aren't worth any premium. The size of this premium remains to be seen.
If open-source models were banned or made illegal, then labs would win big at the cost of innovation.
Inference has proven to be profitable, with margins for service providers ranging between 50% and 70%. Even if people leave hosted providers, this demand must flow to them buying their own hardware. Given that inference is the dominant workload, demand far exceeds supply.
So, how are our large labs performing in this world? I think the answer is probably okay, but margins might not be as high.
It would be wrong to assume they will fail and fall. Although I would love to offer an opinion backed by more data, we don't have clear data on their specific margins; however, it's reasonable to assume their inference service optimization capabilities have reached industry standards. Furthermore, companies with tens or even hundreds of millions of monthly active users are not outright Ponzi schemes or fraudulent projects, and their revenue is growing rapidly.

Figure: Competitive landscape of large labs vs. open-source models.
Return on Investment for Lab Models
Labs may still not generate significant returns on SOTA models, but this doesn't mean the market won't reward new models with stronger capabilities or be unwilling to pay a premium for them. Given that frontier models represent the next class of problems AI can solve, betting against the frontier seems like a bad idea. Exiting this race makes it harder to catch up later (unless through distillation).
Regardless of the SOTA race, inference still requires hardware, and


