BTC
ETH
HTX
SOL
BNB
시장 동향 보기
简中
繁中
English
日本語
한국어
ภาษาไทย
Tiếng Việt

반도체주 40% 급락 후, 연산 수요 펀더멘털 재조명

深潮TechFlow
特邀专栏作者
2026-07-30 06:57
이 기사는 약 5409자로, 전체를 읽는 데 약 8분이 소요됩니다
하이퍼스케일 클라우드 업체들의 연산 주문 backlog이 5,000억 달러에서 2조 달러로 증가했을 때, 이를 단순한 거품이라고 단정하기에는 더 강력한 증거가 필요하다.
AI 요약
펼치기
  • 핵심 관점: 현재 AI 및 반도체 분야의 30~40% 주가 하락은 거품 붕괴가 아니라 연산 수요가 제한적이라는 오판에 기인한다. 본 글은 정부, 군사, 과학 연구부터 기업 및 개인에 이르기까지 각 주체의 연산 수요는 본질적으로 무한하며, AI는 전통적인 인프라의 선형 성장과 달리 연산 자체가 더 많은 연산 수요를 창출하는 '재귀적 수요' 특성을 지닌다고 주장한다.
  • 핵심 요소:
    1. 수요 주체 다양화 및 지불 의사 계층화: 정부(군사), 과학자(연구), 기업(비용 절감 및 확장), 개인(역량 강화) 모두 연산 구매자이며, 각 그룹이 지불할 의사가 있는 가격 상한선은 다르다. 특히 주권 정부와 클라우드 업체의 지출은 협상 대상이 아니다.
    2. 연산 수요의 재귀성: 철도나 인터넷과 달리 AI 연산은 자체적인 수요 성장을 창출한다. 긍정적인 ROI 사례를 발견한 사용자는 연산 투입을 지속적으로 늘려, 이는 전통적인 인프라처럼 외부 확산에 의존하는 것이 아니라 '수익 창출 기계'로서 기능한다.
    3. 하이퍼스케일 클라우드 업체의 강력한 가격 결정력: 퍼블릭 클라우드 연산 비용은 실제보다 10~20배 높게 책정되며, 데이터 반출 수수료 등 잠금 효과를 통해 고객을 통제한다. 저자가 자체 구축 머신의 비용을 분석한 결과, 클라우드 GPU 종량제 가격 기준으로 4.6개월 만에 원금 회수가 가능할 정도로 높은 이윤율을 보여준다.
    4. 연산 주문 backlog이 보여주는 강력한 실수요: 전 세계 연산 주문 backlog은 2025년 초 5,000억 달러에서 2조 달러 이상으로 증가했으며, 주로 실험실이 아닌 고객에 의해 주도된다. 이는 단순한 거품이 아닌 구조적 전환이 진행 중임을 시사한다.
    5. 하드웨어 자산의 가치 상승 가능성: 메모리 비용이 시간이 지남에 따라 상승하기 때문에, 5년 전 GPU 임대 비용은 오히려 하락하지 않고 상승했다. 하드웨어는 '연산(감가상각) + 메모리(가치 상승)'의 이중 구성 요소 게임으로 개념화될 수 있으며, 전반적인 투자 수익률은 예상을 크게 웃돌 가능성이 있다.
    6. 연구소 및 오픈소스 모델 경쟁이 하드웨어 수요에 직접적인 위협이 되지 않음: Kimi K3와 같은 오픈소스 모델이 등장하더라도, 이들의 막대한 메모리 요구량(1.5~2TB)은 여전히 대규모 하드웨어를 필요로 한다. 추론 서비스의 이윤율은 50~70%에 달하며, 주요 워크로드는 추론이고 수요가 공급을 훨씬 초과한다.

Original Author: Kerman Kohli

Translation: TechFlow

Introduction: After semiconductor stocks plunged 30-40%, many declared the AI bubble had burst. But this judgment rests on a fatal assumption: that demand for computing power is finite. From government military use and scientific research to enterprise products and personal applications, all groups are vying for computing power, and each group has a different price ceiling they are willing to pay. More critically, AI possesses a characteristic other infrastructure lacks—recursive demand: computing power itself creates more demand for computing power. When the order backlog for computing power at hyperscale cloud providers grows from $500 billion to $2 trillion, calling this a bubble requires stronger evidence.

Core Question: Infinite Demand or Finite Demand

As of the publication of this article, semiconductor and other AI/momentum-related stocks have fallen 30-40% from their all-time highs.

Many are quick to call this the top for semiconductors/AI/memory and are celebrating their non-participation.

They are likely celebrating too early.

In my view, the semiconductor/AI investment thesis boils down to one question:

"Do you believe the demand for computing power is finite or infinite?"

In discussions, I see too many people stuck in their local experiences of enterprise usage/adoption, extrapolating that to the broader market. I agree that the argument for slower enterprise adoption might hold.

However, this also creates a perverse incentive for smaller, more AI-native companies to defeat incumbents with fewer employees, because AI, if used correctly, can be significantly cheaper than scaling with human labor.

Regardless, let's zoom out and stop viewing AI CapEx solely as enterprise demand. At a high level, AI construction is about bringing computing power online.

This computing power, in turn, can and will be used by:

  • Governments for military and defense purposes
  • Scientists for medical and other cutting-edge research
  • Enterprises to build new products and expand without labor constraints
  • Individuals to empower themselves to do more (code, design, create, ask questions)

These use cases and buyers are each willing to pay different prices for computing power. While some may be priced out, it's unwise to assume no one else has a higher budget. For some, computing expenditure is non-negotiable due to the nature of the brutal contest. Examples include sovereign states and hyperscale cloud providers. For others, computing is a substitute for labor costs, and still much cheaper (no legal overhead, management time, etc.).

Believing we are "overbuilding" or "building beyond capacity" implies that any of the above groups have reached an end state and are satisfied with the status quo. Simply put, it believes:

  • Governments think they don't need smarter weapons and defense capabilities
  • Scientists are satisfied with the amount of research already done
  • Enterprises believe they've done enough work on their product lines and don't want to grow further
  • Individuals have reached a final state of curiosity and don't want to do more

If you truly believe any of the above, then you are right to say AI is a bubble and there is overinvestment.

The price each group is willing/able to pay for computing power differs, but the market will organize around demand to ensure quality/price can be met. It's a misconception to think intelligence is too expensive for all, because those who can profitably orchestrate intelligence will continue to drive demand for it.

But, if you believe that humanity and the groups above will never be satisfied, then you must believe that demand for computing power is infinite. We are currently in a massive race for computing power, and few realize it.

Further Inquiry: What is the ROI?

After spending more time in the market, this is the biggest concern investors have regarding continued computing buildout. Is the trend of hyperscale cloud providers spending all their free cash flow excessive? Are they betting the farm?

Beyond the CapEx of hyperscalers, there's significant focus on what revenue and profits the large labs are generating. The threat of new open-source models devaluing the value labs can extract from frontier models complicates the situation further.

I'll take the time to discuss these one by one, but let's start with the CapEx of hyperscalers.

For those who think hyperscalers are miscalculating, I need you to understand this is far from the truth: public cloud services are outrageously expensive, and they know how to squeeze every penny out of you. They convinced an entire generation of companies that they couldn't scale without them.

In return, they mark up the cost of normal non-CPU compute by 10-20x. On top of that, you pay for logging, data transfer (egress), and 5 other services just to get basic work done.

This game works because they trap you in their system. Bandwidth within the GCP/AWS kingdom is cheap, but skyrockets the moment you try to move out. For many hyperscalers, they need customers to stay in their ecosystem, or they risk losing the business. Not having enough computing power is existential for them. When your customer's data and compute are with you, running out of GPUs is completely unacceptable and forces them to slowly migrate to your competitors. Hyperscalers have created an interesting dynamic where they can force customers to pay whatever they want, and customers are largely powerless. Their tenants are so decoupled from bare metal that there's massive lock-in, making migration a multi-year effort (if possible at all).

Here's a simple example of just how crazy they are on this front. I wrote a few months ago about how I built this $15,000 machine:

Building an AI Inference Machine

Figure: Author's self-built AI inference machine. Source: Kerman Kohli / Substack

You can find the same GPU, lower-spec machine on GCP:

https://cloud.google.com/products/compute/pricing/accelerator-optimized

The on-demand cost is about $3,248 per month, and a 3-year reserved instance is about $1,444 per month.

Figure: Example of Google Cloud GPU instance pricing. Source: Google Cloud

My machine only has 128GB DDR5, but the Google Cloud one has 180GB of some type of memory (they don't tell you if it's DDR4 or DDR5, lol).

Quick math:

  • At on-demand pricing, my machine would break even in 4.6 months
  • With a 3-year commitment, the breakeven is 10 months

The math is similar for other machines. An H200 cluster (GPU released in late 2024) would break even in less than 2 years. Wherever you look, you see very similar math. This doesn't account for: land costs, ongoing electricity, financing costs, and on-site staff, but it should serve as an illustrative example of how hyperscalers know how to price at a massive premium. Of course, spot pricing vs. committed pricing diverges again.

What makes this math even crazier is that GPUs from 5 years ago are: a) holding their value, b) seeing rental costs rise!

It's highly likely that hardware will not depreciate but will appreciate from here onward. While new chips with better compute efficiency are being released, the efficiency they lack is compensated by rising memory costs.

I conceptualize hardware as a two-component game, where one component (compute) technically becomes less valuable, but is offset by another component (memory), which becomes more valuable over time.

If this is the case, the ROI on their CapEx is higher than anyone remotely expects. Regarding the credit risk of hyperscalers, this tweet from Gavin Baker summarizes it well:

Figure: Image accompanying Gavin Baker's tweet. Source: X / @GavinSBaker

Now you might say, how do we know this demand is sufficient? I mean, I can't model every scenario for every customer, but at some point, you need the humility to say that people queuing up to pay is the strongest signal, and you believe they are rational actors spending on positive ROI efforts.

If we look at it from this angle, the order backlog has grown from $500 billion in early 2025 to well over $2 trillion in just 1.5 years. When this is driven by customers, claiming it's all fake/not positive ROI becomes a stretch. Now, the counter-argument is that labs account for a large part of this, but this view is incorrect. According to EpochAI, frontier labs account for a portion, but not all of the global computing demand.

Figure: Computing order backlog grows from $500 billion to $2 trillion. Source: EpochAI

Figure: Composition of global computing demand. Source: EpochAI

Whatever you believe, the fact of a $2 trillion order backlog should indicate something. Thinking that trillions in spending don't reflect a structural shift but rather an excessive bubble is an interesting viewpoint.

Many investors like to reason by analogy, comparing it to the internet buildout, railroads, or past infrastructure projects. I understand the logic here, but it misses a key feature of AI: recursive demand.

For railroads or the internet, you need more people to adopt the technology, and then limit the ceiling of how each person uses the technology to ensure sufficient diffusion in the economy. AI doesn't have these dynamics. In this race, computing power can generate its own demand for computing power. The computing constraint for an individual or organization is effectively infinite. If you find a useful, positive ROI use case, you can keep investing in computing, and it becomes a money-making machine.

Where this gets tricky is that different people have very different experiences using AI. Most of the world uses it as a single-prompt Q&A machine. For people like me, as agentic engineering becomes more capable, they are becoming indispensable, enabling me to do more.

My computing expenditure continues to rise, and will continue to rise because I find more positive ROI use cases. Regardless of the revenue AI generates, the cost savings it creates are undeniable, and this drives the use case for most end customers.

Uncertainty Factor: OpenAI / Anthropic

Following the previous point, we can see the demand backlog is crazy. But how real is the demand from labs (a significant portion of computing demand)? This is where I think the answer is less clear but still reasonable to reason about. I want to break this answer down into inference and training.

If we think a new SOTA model costs up to hundreds of millions of dollars, then we can say it's an investment asset that generates some useful lifecycle value over time through inference (albeit with a steep depreciation curve).

As a counterforce, you have open-source models diffusing into the market and competing with frontier labs for compute at a cheaper cost. These open-source models may or may not be distilled; that's not crucial for understanding the dynamics.

So the dynamic we must question is what happens when a SOTA model comes out? The reality is not everyone will use them for every problem all the time. However, given their SOTA capabilities, they can solve problems that current model classes cannot, and you are willing to pay a premium for this.

You could say models like Kimi K3 change this dynamic because they are open-source, but people forget an important fact: SOTA models are very large, and the hardware required to run them far exceeds what any home model can do. Kimi K3 itself requires close to 1.5TB - 2TB of memory. Good luck finding that.

What makes the model even more interesting is that Kimi ran out of capacity shortly after opening the gates for K3. Sure, a model exists, but someone still has to serve it. It still has to run on capable hardware. Labs will be forced to become more competitive over time, but this doesn't threaten their business because the premium for certain workloads will persist. Additionally, the capacity to serve that model for your workload is equally important. It's disingenuous to say that frontier models aren't worth any premium. How large this premium is remains to be seen.

If open-source models were banned or illegal, then labs would win big at the expense of innovation.

Inference has proven to be profitable, with margins for service providers ranging between 50% - 70%. Even if people leave managed providers, this demand must flow to them buying their own hardware. Given that inference is the dominant workload, demand far exceeds supply.

So how are our large labs performing in this world? I think the answer is probably okay, but margins might not be that high.

It would be wrong to think they will fail and collapse. Although I would like to offer an opinion with more supporting data, we don't have clear data on their specific profit margins; however, it's reasonable to infer that their optimization capabilities for inference services have reached industry standards. Furthermore, companies with tens or hundreds of millions of monthly active users are not blatant Ponzi schemes or fraudulent projects, and their revenue is growing rapidly.

Figure: Competitive landscape of large labs vs. open-source models

Returns on Lab Models

Labs might still not generate substantial returns on SOTA models, but this would imply that the market does not reward new models with greater capabilities and is unwilling to pay for this premium. Considering that frontier models

투자하다
기술
AI
Odaily 공식 커뮤니티에 가입하세요