BTC
ETH
HTX
SOL
BNB
View Market
简中
繁中
English
日本語
한국어
ภาษาไทย
Tiếng Việt

Will Open-Source Models Kill the "Compute Demand"? Morgan Stanley Scenarios Three AI Futures

MSX 研究院
特邀专栏作者
@MSX_CN
2026-08-10 08:49
This article is about 6022 words, reading the full article takes about 9 minutes
When the cost per inference drops and more businesses start calling on AI, the total volume of calls may actually increase instead.
AI Summary
Expand
  • Core Takeaway: Open-weight models lower the cost per call, but a surge in call volume could trigger the "Jevons Paradox," pushing total compute demand higher. AI compute will not disappear; instead, it will spread from centralized training clusters to distributed nodes such as enterprise private clouds and edge devices, shifting the value chain focus from the base model layer to routing, orchestration, security, and local infrastructure.
  • Key Elements:
    1. The average enterprise AI task generates about $55 in value, while direct costs are only $2–5; lower prices bring more workflows across the economic feasibility threshold, expanding call volume.
    2. 63% of surveyed enterprises already use open models, but mostly in combination with closed-source models; the latter still dominate complex reasoning and agent workloads.
    3. MIT research estimates that shifting to open models can cut costs by about 70%, but CMU research shows that the payback period for self-hosted deployments could be as long as 6 years, making the economics scenario-dependent.
    4. The proliferation of open weights will increase demand for middleware such as model gateways, task routing, agent orchestration, observability, and security software—with security becoming more complex due to decentralized deployment.
    5. Morgan Stanley outlines three scenarios (closed-source wins, hybrid coexistence, open-source wins), with NVIDIA, on-site power, and security software as common beneficiaries across all scenarios.

Original report: Morgan Stanley Research, "Weighing In: Open-Weights Models & 3 States of the World," August 3, 2026

Compiled and edited by: DaiDai, Frank, MSX Maitong Research Institute

Key Takeaways

  • Open-weight models lower the cost per call and deployment barriers, but do not necessarily reduce total compute demand. As AI enters more enterprises, workflows, and devices, the growth in call volume may outpace efficiency gains, triggering the classic "Jevons Paradox."
  • Enterprises have already entered a multi-model era, where open-weight models primarily handle coding, document processing, high-frequency calls, and domain-specific tasks, while complex reasoning and agent workloads still rely more heavily on frontier closed-source models.
  • Open weights do not mean free. Enterprises may save on model licensing fees or some API costs, but they still bear the expenses of GPUs, cloud services, on-premises data centers, fine-tuning, talent, security, and operations.
  • Regardless of whether closed-source, hybrid, or open-weight models ultimately dominate, NVIDIA, on-site power generation, and security software represent relatively clear beneficiaries across all scenarios.
  • The more prevalent open-weight models become, the more the AI value chain is likely to shift from the foundation model layer toward inference, routing, orchestration, observability, on-premises infrastructure, edge devices, and vertical applications.

Every so often, the AI market seems to experience a bout of "efficiency panic."

When model parameters shrink, inference costs decline, or a model company achieves near-frontier performance with fewer chips, the market quickly develops an intuition: if the compute required to accomplish the same task is shrinking, shouldn't demand for GPUs, data centers, and electricity also peak?

The emergence of a new generation of open-weight models, such as Kimi K3 and the official release of DeepSeek V4 Flash, has once again brought this question to the forefront. These models not only attempt to narrow the capability gap with frontier closed-source models but also allow enterprises to download, modify, and self-deploy models, spreading capabilities that were once highly concentrated among a few American model labs into the broader developer and enterprise technology stack.

But Morgan Stanley's answer in its latest report runs counter to this intuition:

Increased model efficiency does not necessarily kill compute demand. It is more likely to make AI cheap enough to enter use cases that previously weren't worth the cost, ultimately driving up overall call volumes.

On this basis, the real question that needs to be revisited shifts to: where will compute occur, who will provide it, and which companies can capture revenue from this diffusion?

1. Why Could Cheaper Models Lead to Higher Compute Consumption?

The most common misconception about open-weight models is equating "lower model costs" with "lower infrastructure demand."

But for most enterprises, adopting AI is not simply about how much compute a model needs for a single task; it's about whether the value generated outweighs the combined costs of the model, engineering, and infrastructure.

Citing its own estimates, Morgan Stanley notes that an average enterprise AI task can generate approximately $55 in value, while the direct cost is only about $2–$5. Even after accounting for data, engineering, security, and management costs, this value-to-cost ratio still suggests that a large number of enterprise workflows have yet to be fully AI-enabled.

In this context, lower model prices may not lead to reduced enterprise AI spending, but rather to more tasks crossing the threshold of economic viability.

In the past, enterprises might only assign their most important, highest-value tasks to AI. As inference prices continue to fall, tasks such as customer service record classification, contract review, code testing, product descriptions, enterprise search, marketing materials, data cleaning, and even internal approvals could all be incorporated into model workflows.

The compute required per task may decline, but the number of tasks, frequency of execution, and user scale will expand simultaneously.

This is the "Jevons Paradox" that the report repeatedly highlights: When the efficiency of a resource improves and its unit cost falls, the rapid expansion in its application scope can actually lead to an increase in total consumption.

For example, more fuel-efficient cars did not make global oil demand disappear; cheaper network bandwidth did not reduce data traffic. Similarly, more efficient models may not reduce GPU usage—they may instead transform AI from a few high-value tasks into an omnipresent foundational capability within enterprises.

This diffusion is already underway.

A McKinsey survey cited in the report shows that 63% of responding enterprises already use open models in their technology stacks, but most have not abandoned closed-source models entirely. Instead, they use a combination: open-weight models are primarily used for coding, document parsing, high-frequency calls, and domain-specific tasks, while frontier closed-source models continue to handle complex reasoning, high-reliability requirements, and work that is harder to standardize.

Between February and July 2026, the proportion of tokens routed weekly by U.S. companies via OpenRouter to Chinese open models at one point exceeded 30%. This data may skew toward developers and startups and cannot directly represent large enterprise spending, but it at least demonstrates that open-weight models have moved from lab concepts to real-world call and deployment environments.

Of course, open-weight models also do not mean "free models."

When self-hosting, enterprises can avoid paying per-token API fees to model providers, but they still need to purchase or lease GPUs and bear the costs of data centers, cloud services, fine-tuning, engineering teams, security, and day-to-day operations. This means that when using open-weight models through managed APIs offered by model vendors or cloud platforms, enterprises may still pay based on tokens or compute usage.

Therefore, what open weights change is not whether compute costs exist, but how enterprises pay for compute and whether the value is captured by model providers, cloud platforms, or the enterprises' own infrastructure.

Morgan Stanley cites an MIT study estimating that switching from closed-source to open models could reduce average prices by around 70%, saving consumers approximately $25 billion annually. However, the report also cautions that this study was completed earlier, that different models are not fully comparable in capability, and that it may not adequately account for hidden costs such as engineering, fine-tuning, and operations.

Another Carnegie Mellon study shows that the payback period for self-hosting open models could range from roughly 3 months to as long as 6 years. Models with smaller parameter counts, fixed tasks, and high call frequencies can amortize hardware costs more quickly, while larger models and complex enterprise applications may struggle for a long time to prove that self-hosting is cheaper than APIs due to underutilization, frequent updates, and high fine-tuning costs.

So, there is no one-size-fits-all answer to the economics of open-weight models.

It depends on model size, usage frequency, infrastructure utilization, enterprise engineering capability, and whether data must remain on-premises. The more open the model, the more choices enterprises have—but also the more technical responsibility they assume.

2. Where Will Compute Occur in the Next Phase?

If the future of AI were dominated solely by a few closed-source models, training and inference would continue to concentrate in large cloud platforms and hyperscale data centers.

But if open-weight models gain broader adoption, AI compute will not disappear. Rather, it will diffuse from a few central hubs further into private clouds, enterprise data centers, sovereign data centers, edge servers, and personal devices.

This means the market's focus will shift from "how many GPUs do we need" to "where are GPUs deployed, who manages them, and how are they invoked."

In the closed-source era, enterprises could outsource most of the complexity to model companies: plug into an API, pay per token, and let the provider handle model updates, infrastructure, safety alignment, and some legal liability.

In the multi-model era, this simple structure begins to break down.

Enterprises may assign their most complex tasks to frontier closed-source models and high-frequency, cost-sensitive tasks to open-weight models; sensitive data stays on-premises while general workloads run on public clouds; some requests are processed in data centers, while others are handled directly on computers, phones, or other edge devices.

The more models there are and the more dispersed their deployment, the more complex the enterprise AI stack becomes:

  • First, the gateway layer: Enterprises need a unified entry point to manage authentication, access permissions, rate limits, logging, and failover across different models.
  • Second, the routing layer: Systems need to assign each request to the most appropriate model based on accuracy, latency, cost, data sensitivity, and task difficulty.
  • Third, the orchestration layer: Complex agent workflows typically require coordination across multiple models, databases, and external tools, with a single user request potentially decomposed into dozens of steps and calls.
  • Finally, the observability and evaluation layer: Enterprises must continuously track model output quality, response latency, token usage, operational costs, failure causes, and security risks.

This is why the diffusion of open-weight models benefits more than just model vendors.

When foundation models themselves become more accessible, enterprises become more willing to pay for "how to reliably put models into production." Model gateways, task routing, agent orchestration, data governance, observability, and security software may become the harder-to-compress segments of the value chain.

Security stands out in particular. When closed-source models run in a centralized manner, some security responsibilities are borne by model labs and cloud platforms. But when open models enter enterprise on-premises environments, private clouds, and edge devices, identity, data, endpoints, model weights, and runtime environments all need independent protection.

Enterprises not only need to prevent employees from sending sensitive data to the wrong models, but also need to manage which databases and tools different models can access, and whether they are vulnerable to prompt injection, model distillation, weight tampering, or permission abuse.

The more dispersed the model deployment, the larger the attack surface; the more models there are, the higher the governance costs.

Therefore, what open-weight models truly bring is not the disappearance of AI infrastructure value, but the diffusion of value from a single foundation model and API layer across the entire system stack.

Future AI spending may no longer be reflected solely in a few tech giants building massive training clusters. It will also manifest as enterprises purchasing servers, upgrading networks and storage, deploying security systems, building private AI platforms, and equipping phones and PCs with more powerful on-device compute.

The shift of compute demand from centralized to distributed does not mean the total volume shrinks. It simply means the beneficiaries are no longer concentrated solely among model labs and major cloud providers.

3. The Three States of AI: No Matter Who Wins, Compute, Power, and Security Are Unavoidable

Morgan Stanley does not assign explicit probabilities to the three future scenarios. Instead, it separately models how value distribution across the AI industry chain might play out if closed-source models win, if a hybrid architecture prevails, or if open-weight models dominate.

State One: Closed-Source Models Maintain Frontier Capabilities

In this scenario, the performance, reliability, and safety capabilities of frontier models remain difficult to replicate, and a few well-capitalized model labs continue to lead.

Enterprises prioritize accuracy, deployment convenience, IP indemnification, and brand trust over full control of model weights, and thus remain willing to pay for closed-source APIs and enterprise subscriptions.

Training and inference workloads become even more concentrated on hyperscaler platforms like AWS and Google Cloud, with super-training clusters continuing to drive demand for GPUs, high-speed networking, optical communications, and custom ASICs.

Platforms with both cloud infrastructure and model capabilities, such as Google and Amazon, are in a stronger position, while chip and networking vendors like Broadcom, Arista Networks, Lumentum, and Coherent may also benefit.

The report estimates that if Google can run the Gemini API on its own infrastructure backed by a leading model, its illustrative return on invested capital could reach approximately 45%. Even if its models are not in an absolute leadership position, serving purely as an infrastructure provider could still yield returns close to 30%.

This shows that in the closed-source world, the most important asset is not just the model itself, but also the ownership of the compute on which it runs. Whoever owns the chips, data centers, and customer entry points is better positioned to retain the profits generated from model calls.

State Two: Open and Closed Models Coexist Long-Term

This is the scenario that most closely mirrors the current real-world state of enterprise usage.

Frontier closed-source models handle complex reasoning, long-horizon agents, and high-reliability tasks, while open-weight models and smaller models process high-frequency, cost-sensitive, low-latency, or highly specialized work.

Enterprises do not choose a single model; they continuously switch and route based on the task at hand. This also means AI deployment exists simultaneously across public clouds, private clouds, on-premises infrastructure, and edge devices.

In this scenario, no single model vendor can easily establish monopolistic pricing power, but the entire AI software and infrastructure market sees the broadest demand.

  • Microsoft, Amazon, and Google still benefit from cloud workloads;
  • Infrastructure and workflow software vendors such as Datadog, Palantir, and Appian may benefit from model orchestration, data connectivity, and observability demand;
  • Security vendors including Palo Alto Networks, CrowdStrike, Fortinet, Zscaler, Netskope, and Okta benefit from the continued expansion of enterprise attack surfaces and identity boundaries.

Meanwhile, networking equipment vendors like Cisco and F5, as well as enterprise infrastructure companies such as Dell, HPE, and NetApp, may also see incremental demand from on-premises and hybrid deployments.

The biggest investment implication of the hybrid architecture is that AI spending is not concentrated solely in training clusters but diffuses layer by layer along the enterprise technology stack. Moreover, the more intense the competition at the model layer, the more enterprises need neutral software and infrastructure to switch between different models and deployment environments.

State Three: Open-Weight Models Become Mainstream

In this scenario, open-weight models progressively approach frontier closed-source models, foundation model intelligence becomes broadly accessible, and API prices drop significantly.

Enterprises are no longer willing to pay a high premium for generic model capabilities. Instead, they fine-tune with proprietary data and direct the bulk of their budgets toward inference optimization, agents, industry tools, and application deployment.

The center of innovation shifts from pre-training to post-training and application layers, and AI infrastructure becomes more decentralized. Governments and large enterprises may deploy more models in sovereign clouds, private data centers, and on-premises environments for reasons including data sovereignty, privacy, latency, and avoiding vendor lock-in.

  • Microsoft may strengthen its position through Azure, GitHub, enterprise software, and the open-model ecosystem;
  • The strategic value of open-model providers such as MiniMax, Zhipu/Z.ai, Alibaba, and Tencent rises;
  • On-premises infrastructure vendors like Dell, HPE, and NetApp, endpoint device companies like HP and Apple, as well as system integrators and IT distribution channels, also see more pronounced demand.

But the open world does not mean cloud platforms lose their value.

Most enterprises will still not build infrastructure entirely on their own. Open models are still likely to run on Azure, AWS, or Google Cloud. The hyperscaler role shifts from being the sole model gateway to providing open-model hosting, compute resources, data services, and enterprise AI platforms.

Thus, cloud providers face not a simple win-or-lose, but a change in profit structure. For example, foundation model rents may decline, but revenue from compute, storage, databases, security, and enterprise services may still grow.

Common Winners Across All Three States

In Morgan Stanley's asset matrix, what stands out most is not the company unique to each scenario, but the names that recur across all three columns.

First, NVIDIA.

If closed-source models win, more compute concentrates in hyperscale training and inference clusters. If open-weight models win, inference nodes spread to enterprise, on-premises, and edge environments.

The chip form factors, customer structures, and cluster sizes may differ between these two paths, but both require continuously increasing compute capacity. Ultimately, open models lower the barrier to using models—they do not eliminate the compute itself.

Second, power.

The closed-source world needs stable gigawatt-scale power for super data centers, while the open world increases power demand from enterprise data centers, sovereign clouds, and on-site inference facilities. More efficient models may reduce power consumption per task, but they may also enable more tasks and devices to run AI continuously.

Companies such as Bloom Energy, Williams, and Liberty Energy—providers of on-site power, natural gas, and energy infrastructure—are repeatedly listed by the report as beneficiaries.

Third, security software.

Whether enterprises use closed-source or open models, they all need to protect identity,

AI
Welcome to Join Odaily Official Community