BTC
ETH
HTX
SOL
BNB
ดูตลาด
简中
繁中
English
日本語
한국어
ภาษาไทย
Tiếng Việt

AI 算力金融化:开源模型正在把算力推向资本市场(下)

欧易OKX
特邀专栏作者
2026-08-17 11:45
บทความนี้มีประมาณ 7444 คำ การอ่านทั้งหมดใช้เวลาประมาณ 11 นาที
算力金融化正在从产品叙事走向真实需求,债务融资和公开采购正在催生真实的套保需求。
สรุปโดย AI
ขยาย
  • 核心观点:算力金融化正从硬件租赁向标准化金融基础设施演进,市场定价权将属于率先积累真实订单流和合同标准的一方;推理经济催生四类资产化方向,但链上机会在于标准形成前的产权设计窗口期。
  • 关键要素:
    1. 算力市场金融基础设施分四层:指数建立基准、交易所提供标准化合约、Dealer 管理基差、容量平台连接物理交割,各层竞争取决于真实订单的积累。
    2. 算力期货冷启动困难,GPU 规格、地区、租期等高度非标导致流动性难以集中;远期定价依赖芯片交付、电力等物理信息,传统做市商难以深报。
    3. OTC 市场由 Dealer 承接期限与规格错配,赚取 bid-ask spread 和 basis premium;FalconX、Wintermute 已开始试水 H100 远期报价。
    4. 推理服务金融化催生四类资产:GPU 信用、预付服务权、产能激励、模型收益权;其中 GPU 信用最易被传统资本吸收,链上留存的细分市场在中小型贷款和次级层。
    5. Router(路由层)被低估,因其掌握真实成交账单,可形成质量调整后的 workload cost index,并通过调度替代性降低对衍生品的依赖。
    6. 链上窗口期在标准形成前:合同统一规格后大额资金将流向传统机构,链上协议应在此期间积累订单流、合同标准和履约数据。

This is an in-depth research report produced by OKX Ventures. Due to its length, it is published in two parts: the first part focuses on the assetization of compute power, risk exposure, and derivatives pricing logic; the second part will analyze the financial infrastructure for compute power, inference assetization opportunities, and industry trends. This is the second part.

4. Financial Infrastructure of the Compute Market: Pricing, Risk, and Settlement

The foundational layers of the compute financial market are beginning to take shape. Indices compress fragmented GPU rental prices into a unified benchmark; exchanges organize standardized risk around these benchmarks; Dealers absorb the basis arising from specific data centers, tenors, and SLAs; and capacity platforms connect financial positions back to physical compute power. The layer that accumulates the most real orders and contract references first will be closest to establishing pricing power in this early market.

4.1 Compute Indices: Who Can Become the Default Reference Price

Transaction prices for the same GPU model can still diverge significantly based on cluster size, interconnect type, region, lease duration, prepayment percentage, SLAs, and supplier creditworthiness. For index providers, the challenge lies in determining whether they can process these variations well enough for the market to use their index for contract settlement, loan valuation, and derivatives delivery.

Currently, three main approaches are visible. Silicon Data and Ornn are closer to independent benchmark providers, primarily collecting market quotes and contract data, then standardizing it to form a benchmark. Compute Desk also acts as a compute broker and settlement layer, allowing it to directly see executed spot and forward contracts, including prices, tenors, and actual settlement terms. SemiAnalysis, on the other hand, builds TCO models based on electricity, chip procurement, depreciation, and data center costs, making it more suitable for asset valuation and research, further removed from being a financial settlement benchmark.

4.2 Exchange-Traded/Standardized Futures: Traditional and Emerging Compliant Exchanges Enter the Compute Market

The core value offered by exchanges is standardization, clearing, and distribution. For Neoclouds, AI Labs, and private credit institutions, if hedging can be directly integrated into existing FCM, margin, and clearing systems, the adoption cost for institutions will decrease significantly. Consequently, both traditional exchanges and emerging compliant platforms are beginning to test compute products: CME and ICE are extending their mature futures infrastructure, while Architect AX/AIX and Pluto/PMEX are exploring different paths with perpetual contracts and fixed-expiry futures.

The cold start for compute futures is far more complex than listing a contract. Compute contracts can easily become fragmented. Different GPU models, regions, and cluster configurations correspond to different delivery values, making it difficult for liquidity to naturally concentrate in a single contract. Furthermore, far-dated prices are highly dependent on chip delivery, electricity, and model efficiency. Standard financial market makers lack access to physical order information, making it difficult to continuously provide deep quotes on six-month or one-year tenors.

Industrial hedging flows may not be sufficiently large in the short term either. Hyperscalers can absorb significant risk internally, and top-tier Neoclouds have already covered a substantial portion of their cash flows with long-term contracts. The early exchange market is likely to first accumulate trading and speculative liquidity. Whether industrial hedging can catch up will depend on whether mid-sized operators and AI teams procuring compute on-demand begin to trade consistently and repeatedly.

4.3 OTC Trading, Market Makers: OTC Customization and Basis Management

Before sufficient on-exchange depth develops, a large volume of risk is likely to be transferred through OTC markets first. The specifics of physical compute—region, network, cluster size, SLA, and delivery time—are highly detailed and difficult to fit directly into a standardized exchange contract.

Dealers therefore occupy a valuable position. On one side, Neoclouds want to lock in rental income in advance; on the other, AI Labs want to lock in procurement costs. A Dealer can offer customized quotes to both parties and manage the overall price risk using standardized futures or other positions. When buy and sell demand doesn't align in time or specification, the Dealer must also use its own balance sheet to take on the positions.

This business earns more than just the bid-ask spread. Data center costs, electricity, network conditions, delivery constraints, and client creditworthiness all contribute to the basis. The ability to continuously price these non-standard differences and use capital to take on the associated risk is a capability that is harder for others to replicate. FalconX's H100 swap with Robert Leshner and Wintermute's introduction of H100 forwards can be seen as early attempts by market makers to test the waters in compute forward pricing and risk transfer. Financial intermediaries are beginning to offer customized tenor quotes for physical compute, keeping the non-standardizable parts on their own books.

Real order flow is therefore crucial. Indices tell the market the average price, but the quotes and transactions within a Dealer's book are closer to the actual premium paid for a specific region, configuration, and tenor. This information directly determines how the basis is priced and where profits reside.

4.4 Capacity Platforms and Physical Delivery: Fulfilling Paper Positions with Physical Assets

For AI Labs, a financial position locks in a price but does not guarantee that a compliant set of GPUs will be available at the delivery date. This distinction is particularly important during times of compute scarcity. Without physical fulfillment capability, futures prices and real data center rental rates can easily diverge.

Several attempts at different levels are already emerging. Spot platforms like Vast.ai, RunPod, and Hyperbolic provide richer physical pricing and available capacity data. SF Compute allows for the resale or buyback of reserved capacity, giving long-term capacity contracts a degree of liquidity. The EFP network attempts to connect standardized futures with specific physical contracts, using futures for broad market price exposure and OTC contracts for region, network, and SLA specifics.

Models like SF Compute are particularly noteworthy. In traditional Take-or-Pay contracts, unused capacity is often close to a sunk cost. By allowing resale, future capacity begins to acquire a tradable rights attribute. This increases the liquidity of long-term contracts and allows a clearer forward price to emerge in the physical market.

EFP addresses a different problem. An AI Lab can first use standardized futures to lock in market beta, and then, when it comes time for actual deployment, use the fulfillment network to find a specific data center, paying the basis corresponding to region, network, and SLA. This way, the financial market doesn't need to shoehorn all physical differences into a single futures contract to serve real industrial delivery.

4.5 Inference Capital Markets: Financing, Service Rights, and Revenue Rights in the Inference Economy

Once GPUs are converted into Tokens by inference service providers, financial needs extend from hardware leasing to the working capital and contractual rights of the inference business. If an inference service provider wants to lock in its full gross margin, it theoretically needs to manage both its GPU input costs and its Token output prices.

An inference service provider's profit is roughly equal to: Token Realized Price × Token Output per GPU-hour × GPU Utilization − GPU Costs − Other Operating Costs

4.5.1 What are the Demand Points for Assets in the Inference Economy?

GPU and capacity need to be procured in advance, while API revenue accrues subsequently as client requests come in. Overestimating demand leaves idle machines and ties up cash. Large model companies and Hyperscalers can absorb this volatility with their own balance sheets, but smaller inference platforms often rely on customer prepayments or external financing.

The time lag between capacity deployment and demand realization creates four types of assets: GPU credit finances upfront investment; prepaid service rights collect future revenue early; capacity incentives subsidize supply before demand forms; and model deployment assets attempt to distribute residual profit after operating costs.

The traditional cloud market has long used prepaid credits and reserved capacity to handle similar supply-demand mismatches. For example, OpenAI's Scale Tier allows enterprises to pre-purchase throughput units for input and output Tokens of specific models. This is already a form of standardized prepaid inference service, but the entitlement is tied to the enterprise account and cannot be freely transferred.

Crypto adds a new layer of property rights design on-chain. Service entitlements previously recorded in a provider's account can now be held by a wallet, and can be further transferred and used for financing. In the future, if Agents begin procuring inference services for themselves, such rights could directly enter a machine's budget system. As programmatic payments become a ubiquitous infrastructure capability in the future, the more distinctive position for blockchain lies in the rights layer. Whether a future service right can be independently held, transferred, or used for financing will directly impact its ability to evolve from a simple credit into a financial asset.

Currently, the share of on-chain inference at the execution layer remains very small. Over the past three months, crypto-native inference service providers have accounted for only about 0.5%–1% of OpenRouter's daily Token traffic. This also implies that the more realistic opportunity for blockchain at this stage may lie in the financial layer, where inference services continue to be produced by the most efficient platforms, while the blockchain handles capital flows and rights transfer.

4.5.2 Four Attempts at Inference Assetization

Currently, there are four directions for equity assetization within Inference Capital Markets.

GPU Credit

GPU credit is the first to reach scale, partly because it already has an external payer and its financial product structure closely resembles traditional credit assets, making it easier for institutional capital to understand.

A typical GPU loan still follows traditional asset financing structures: the GPUs are held by a Delaware SPV, and the lender establishes a security interest through a Loan and Security Agreement, UCC-1 filings, and data center lien waivers. Loan shares, repayments, and profit distributions are recorded on-chain; however, in the event of default, debt enforcement still relies on off-chain contracts, physical control of the equipment, and the court system.

After the borrower deploys the GPUs, downstream customers continuously pay for compute services, and this cash flow supports the loan's principal and interest. As long as the offtake contracts continue to be honored, the loan has a clear source of repayment.

However, GPU credit is also the direction most likely to be absorbed by traditional capital. As senior loans acquire standard tenors, interest rates, LTVs, and legal recourse, the cost of capital for banks, insurance funds, and private credit will be far lower than on-chain capital. Large, investment-grade, structurally clear loans will likely eventually flow into traditional syndicated loans, ABS, or private credit funds.

The long-term on-chain niche will likely remain where TradFi service costs are higher, such as: loan origination for smaller operators, cross-border stablecoin funding, distribution of loan participations, real-time equipment and utilization monitoring, and holding the first-loss junior or mezzanine tranches. A more likely future structure involves on-chain protocols handling borrower discovery, asset monitoring, and organizing the first round of capital; once loan performance stabilizes, the senior portion is sold to traditional capital, while the blockchain retains the servicing fees, subordinated risk, and data relationships.

Prepaid Service Rights

Prepaid service rights address the issuer's current funding needs. Platforms sell future API capacity in advance, receive cash upfront, and fulfill the service over time. As the duration lengthens, the platform risk borne by the buyer becomes increasingly apparent. Will the model of a few years from now still be competitive? How low could service prices go? Can the purchasing power of the credits be maintained? All of these affect today's valuation.

Therefore, these products are likely to gradually move closer to standardized capacity contracts. Tenors will be shorter, service standards clearer, and model substitutions can follow agreed-upon rules. The clearer the contract, the easier it is for prepaid service rights to develop a relatively independent market price.

Furthermore, there is an adverse selection issue here. Platforms with the best models, strongest customer demand, and most stable pricing power typically have little incentive to issue long-term, transferable inference rights. They prefer to retain the ability to price discriminate among customers, recognize revenue from unused credits, and adjust prices at will. Those most in need of selling long-term service rights in advance are platforms that require financing and user acquisition but whose future competitiveness has yet to be proven.

Capacity Incentives

Capacity incentives address the early-stage supply organization problem for inference networks. Nodes need to configure GPUs and models in advance and bear the operational costs. When demand hasn't stabilized on the demand side, nodes lack sufficient incentive to come online early. Network token emissions can provide a period of income security for the supply side, allowing the network to acquire usable capacity first.

Its long-term viability still depends on whether external customer payments can gradually cover node revenue. A protocol can verify that an inference task was indeed executed or that a node used a specified model, but the mere occurrence of computation does not create economic value. Token emissions can buy supply for a period, but customer demand still needs to be proven through real orders.

A more rational incentive mechanism might gradually become tied to real revenue. Customers first pay for tasks in stablecoins, nodes complete inference and receive income after verification, and the protocol then provides additional rewards based on external revenue, gradually reducing subsidies as the network matures. This allows emissions to bear the cost of cold-starting the supply side while sending a clearer signal to the market: how much of node revenue comes from customers versus how much still comes from capital subsidies. In the long run, the real task for capacity incentives is transitioning from using tokens to buy supply to using customer revenue to sustain supply. Only after this transition occurs does the network begin to establish an economic foundation independent of its token market cycle.

Single-Model Revenue Rights

Model revenue rights are closest to the crack spread mentioned earlier:

Model Deployment Residual Profit = API Revenue − GPU Costs − Network, Storage, and Engineering Costs − Customer Acquisition and Subsidies

Many Model Tokens offer holders mainly buybacks and staking incentives, which may not correspond to a legally enforceable claim on income. If there is no clear contractual relationship between API revenue and token value, valuation lacks a stable anchor.

Model lifespan further amplifies this problem. Model iteration is rapid; by the time an asset is financed and deployed, the underlying model may already be overtaken by newer alternatives. A single model thus struggles to support a long-term financial asset.

A more likely development path is to gradually expand the underlying unit to a workload strategy. For example, a coding inference strategy can continuously switch underlying models, routing based on quality and cost. In this way, the holder is acquiring a going-concern inference strategy, and the asset's lifespan follows customer demand and workload characteristics.

Additionally, if these products are to enter the institutional market, they will require clear legal entities, audited API revenues, and a clearly defined cash flow waterfall. Without these fundamentals, Model Token valuations will remain highly dependent on future usage and the issuer's buyback policy.

4.5.3 Routers Provide Demand-Side Risk Management Before Token Futures

Routers are the underestimated layer in inference financialization.

What enterprises ultimately care about is the cost of a task. Which specific model is used is, in many scenarios, just the implementation path. A Router can continuously compare different models and providers, reallocating traffic as prices change. The underlying services are highly heterogeneous, but through Router orchestration, their substitutability gradually increases.

This directly impacts the scale of demand-side derivatives. Much cost volatility can be absorbed first through model switching and software optimization. It is the remaining exposure that cannot be addressed this way which requires financial instruments.

Routers also hold a significant data advantage: they see real invoices. Public API official prices rarely represent the actual procurement costs for enterprises. Routers see what customers actually pay and the quality of service associated with that price. As order volumes accumulate, this data could potentially form a quality-adjusted workload cost index.

Going a step further, Routers could directly offer enterprises cost ceilings. For example, an enterprise locks in a budget for its coding workload for the next three months. The Router manages model selection and capacity procurement in the background, using GPU derivatives to handle residual risk when needed. The customer ultimately gains budget certainty without having to manage a portfolio of models, futures positions, or basis risk themselves. This type of product may align better with enterprise procurement habits than generic Token futures. Routers can keep the underlying complexity internal and offer customers a more stable and predictable price.

While Tokens haven't formed a unified standard yet, a Router can continuously seek alternative supply for a specific, well-defined workload. Once trading accumulates to a certain scale, the workload itself could become a new pricing unit.

If this path proves viable, Routers will gradually extend from traffic distribution tools into procurement and risk management. Their most valuable long-term assets will also become increasingly clear: real order flow, real transaction prices, and the capability to manage cost and quality across different models. The price standard for workloads is likely to form

AI
ยินดีต้อนรับเข้าร่วมชุมชนทางการของ Odaily
กลุ่มสมาชิก
https://t.me/Odaily_News
กลุ่มสนทนา
https://t.me/Odaily_GoldenApe
บัญชีทางการ
https://twitter.com/OdailyChina
กลุ่มสนทนา
https://t.me/Odaily_CryptoPunk
ค้นหา
สารบัญบทความ
ดาวน์โหลดแอพ Odaily พลาเน็ตเดลี่
ให้คนบางกลุ่มเข้าใจ Web3.0 ก่อน
IOS
Android