BTC
ETH
HTX
SOL
BNB
View Market
简中
繁中
English
日本語
한국어
ภาษาไทย
Tiếng Việt

AI Computing Power Financialization: Open-Source Models Are Pushing Compute Power Toward Capital Markets (Part 2)

欧易OKX
特邀专栏作者
2026-08-17 11:45
This article is about 7444 words, reading the full article takes about 11 minutes
The financialization of computing power is moving from product narratives to real demand, with debt financing and public procurement driving genuine hedging needs.
AI Summary
Expand
  • Core Takeaways: The financialization of compute power is evolving from hardware leasing into standardized financial infrastructure, and market pricing power will belong to those who first accumulate real order flow and contract standards; the inference economy is giving rise to four types of assetization directions, but the on-chain opportunity lies in the window for property rights design before standards are established.
  • Key Elements:
    1. Compute market financial infrastructure is divided into four layers: indices establish benchmarks, exchanges provide standardized contracts, Dealers manage basis risk, and capacity platforms connect physical settlement — competition at each layer depends on the accumulation of real order flow.
    2. Compute futures face cold-start difficulties, as GPU specifications, regions, and lease terms are highly non-standardized, making liquidity difficult to concentrate; forward pricing depends on physical information such as chip delivery and electricity, making it hard for traditional market makers to quote deeply.
    3. The OTC market sees Dealers absorbing term and specification mismatches, earning bid-ask spreads and basis premium; FalconX and Wintermute have already begun testing H100 forward quotes.
    4. The financialization of inference services gives rise to four types of assets: GPU credit, prepaid service rights, capacity incentives, and model revenue rights; GPU credit is the easiest for traditional capital to absorb, with the on-chain niche market remaining in small-to-medium lending and subordinated layers.
    5. The Router (routing layer) is undervalued because it holds real executed invoices, enabling the formation of quality-adjusted workload cost indices, and can reduce reliance on derivatives through scheduling alternatives.
    6. The on-chain window exists before standards are formed: once contracts standardize specifications, large capital flows will move to traditional institutions, so on-chain protocols should accumulate order flow, contract standards, and performance data during this period.

This article is an in-depth research report produced by OKX Ventures. Due to its length, it is published in two parts: the first part focuses on the assetization of compute power, risk exposure, and derivative pricing logic; the second part will focus on analyzing compute financial infrastructure, inference assetization opportunities, and industry trends. This is the second part.

IV. The Financial Infrastructure of the Compute Market: Prices, Risks, and Settlement

The foundational layers of the compute financial market are beginning to take shape. Indices compress disparate GPU rental prices into a unified benchmark; exchanges organize standardized risk around the benchmark; Dealers absorb the basis arising from specific data centers, tenures, and SLAs; and capacity platforms connect financial positions back to physical compute power. The layer that accumulates enough real order flow and contract references first will be the one closest to pricing power in the early market.

4.1 Compute Indices: Who Can Become the Default Reference Price

Transaction prices for the same GPU model are still differentiated by cluster size, interconnect type, region, lease term, prepayment ratio, SLA, and supplier credit. For index providers, the challenge lies in standardizing these differences to a degree where the market is willing to use the index for contract settlement, loan valuation, and derivative delivery.

Three approaches are currently visible. Silicon Data and Ornn are closer to independent benchmark providers, primarily collecting market quotes and contract data before standardizing it into benchmarks; Compute Desk also acts as a compute broker and settlement layer, allowing it to directly observe executed spot and forward contracts, including prices, tenures, and actual settlement terms; SemiAnalysis, meanwhile, takes a TCO modeling approach based on electricity, chip procurement, depreciation, and data center costs, making it more suitable for asset valuation and research, but further from a financial settlement benchmark.

4.2 Exchange-Traded / Standardized Futures: Traditional and Emerging Compliant Exchanges Enter the Compute Market

The core value exchanges provide is standardization, clearing, and distribution. For Neoclouds, AI Labs, and private credit institutions, if hedging can be directly integrated into existing FCM, margin, and clearing systems, the cost of institutional adoption would decrease significantly. Consequently, both traditional exchanges and emerging compliant platforms are starting to test compute products: CME and ICE are building on mature futures infrastructure, while Architect AX/AIX and Pluto/PMEX are exploring different paths with perpetual contracts and fixed-expiration futures.

The cold start for compute futures is far more complex than listing a contract. Compute contracts fragment easily. Different GPU models, regions, and cluster configurations correspond to different delivery values, making it difficult for liquidity to naturally concentrate in a single contract. Far-dated prices are also highly dependent on chip delivery, electricity, and model efficiency. General financial market makers lack physical order information, making it difficult for them to sustain deep quotes on six-month or one-year tenors.

Industry hedging flow may not be substantial enough in the short term either. Hyperscalers can absorb significant risk internally, and top-tier Neoclouds have already covered much of their cash flow with long-term contracts. The early exchange-traded market is likely to first accumulate trading and speculative liquidity; whether industry hedging follows depends on whether small and medium-sized operators and AI teams procuring compute on demand begin trading consistently and repeatedly.

4.3 OTC Trading and Market Makers: OTC Customization and Basis Management

Before sufficient on-exchange depth develops, a significant amount of risk is likely to be transferred over-the-counter (OTC) first. The region, network, cluster size, SLA, and delivery time of physical compute are all highly specific, making them difficult to fit directly into a standardized exchange contract.

Dealers, therefore, occupy a valuable position. On one side, Neoclouds want to lock in rental income in advance; on the other, AI Labs want to lock in procurement costs. A Dealer can provide custom quotes to both parties and manage overall price risk using standard futures or other positions. When buying and selling demand doesn't align in time and specification, the Dealer must also step in, using its own balance sheet to take on the position.

This business earns more than just the bid-ask spread. Data centers, electricity costs, network conditions, delivery constraints, and client credit all contribute to basis. Continuously pricing these non-standard differences, and being willing to deploy capital to absorb the risk, is a far more difficult capability for Dealers to replicate. FalconX's H100 swap with Robert Leshner and Wintermute's H100 forwards can be seen as early attempts by market makers to test compute forward pricing and risk transfer. Financial intermediaries are beginning to offer customized tenor quotes for physical compute, keeping the non-standardizable portion on their own books.

Real order flow is therefore crucial. Indices tell the market the average price, while the quotes and executed trades a Dealer sees are closer to what a specific region, configuration, and tenor actually costs in terms of premium; this information directly determines how basis is priced and where profits reside.

4.4 Capacity Platforms and Physical Delivery: From Paper Positions to Physical Settlement

For AI Labs, locking in a price via a financial position does not guarantee that a compliant set of GPUs will be available at the time of delivery. This distinction is especially important during periods of compute shortage. Without physical delivery capabilities, futures prices and real data center rental rates can easily diverge.

Several attempts at different levels have emerged. Spot platforms like Vast.ai, RunPod, and Hyperbolic offer a richer set of physical quotes and available capacity data; SF Compute allows for the resale or buyback of reserved capacity, giving long-term capacity contracts a degree of liquidity; and the EFP network attempts to connect standard futures with specific physical contracts, using futures for broad market pricing and off-exchange contracts for region, network, and SLA specifics.

The SF Compute model is particularly noteworthy. In traditional Take-or-Pay contracts, unused capacity is often close to a sunk cost. By allowing resale, future capacity begins to take on the attributes of a tradable right. This increases the liquidity of long-term contracts and allows for a clearer forward price curve to emerge in the physical market itself.

EFP solves a different problem. An AI Lab can first use standard futures to lock in market beta, and then, when it's time for actual deployment, use the fulfillment network to find a specific data center, paying the basis corresponding to region, network, and SLA. This way, the financial market doesn't need to cram all physical differences into a single futures contract to serve genuine industrial delivery.

4.5 Inference Capital Markets: Financing, Service Rights, and Revenue Rights in the Inference Economy

Once GPUs are converted into Tokens by inference service providers, financial needs extend from hardware leasing to the working capital and contract rights of the inference business. If an inference service provider wishes to lock in full gross margin, it theoretically needs to manage both the input cost of GPUs and the output price of Tokens.

An inference service provider's profit is roughly equal to: Token Realized Price × Token Output per GPU-hour × GPU Utilization Rate − GPU Cost − Other Operating Costs

4.5.1 What Are the Demand Points for Inference Economy Assets?

GPU and capacity need to be procured in advance, while API revenue accrues gradually as client requests come in. Overestimating demand leaves idle machines and ties up cash. Large model companies and Hyperscalers can absorb this fluctuation with their own balance sheets, but small and medium-sized inference platforms often rely on customer prepayments or external financing.

The time gap between capacity being deployed and demand materializing gives rise to four types of assets: GPU Credit finances upfront investment; Prepaid Service Rights collect future revenue early; Capacity Incentives subsidize supply before demand forms; and Model Deployment Assets attempt to distribute residual profit after operating costs.

Traditional cloud markets have long used prepaid credits and reserved capacity to handle similar supply-demand mismatches. For instance, OpenAI's Scale Tier allows enterprises to pre-purchase input and output Token throughput units for specific models. This is already a standardized prepaid inference service, but the credits are tied to the enterprise account and are not freely transferable.

Crypto adds a new layer of property rights design on-chain. Service credits previously recorded in a supplier's account can now be held by a wallet and further transferred or used for financing. In the future, if Agents begin procuring inference services themselves, these rights could also enter machine budget systems. As programmatic payments become a common infrastructural capability industry-wide, the most distinctive position on-chain will actually be in the rights layer. Whether a future service can be held independently, transferred, or used for financing will directly affect whether it can evolve from ordinary credits into a financial asset.

Currently, the share of on-chain inference at the execution layer remains very small. Crypto-native inference service providers have accounted for only about 0.5%–1% of OpenRouter's daily Token traffic over the past three months. This also implies that the more realistic opportunity on-chain at this stage may appear in the financial layer, where inference services continue to be produced by the most efficient platforms, and the chain handles capital and rights transfer.

4.5.2 Four Attempts at Inference Assetization

Currently, Inference Capital Markets are exploring four directions for the assetization of rights.

GPU Credit

GPU Credit was the first to achieve scale, partly because it already has external payers and partly because its financial product structure is very similar to traditional credit assets, making it easier for institutional capital to understand.

A typical GPU loan still follows traditional asset financing structures: the GPU is held by a Delaware SPV, and the lender establishes a security interest through a Loan and Security Agreement, UCC-1 filings, and data center lien waivers. Loan shares, repayments, and profit distributions are recorded on-chain; however, in the event of default, debt enforcement still relies on off-chain contracts, equipment control, and the court system.

After the borrower deploys the GPUs, downstream clients continuously pay compute fees, and this cash flow supports the loan's principal and interest. As long as the offtake contracts continue to perform, the loan has a clear source of repayment.

However, GPU Credit is also the direction most easily absorbed by traditional capital. As senior loans with standard tenors, interest rates, LTVs, and legal recourse develop, the cost of capital for banks, insurance funds, and private credit will be significantly lower than on-chain capital. Large, investment-grade, clearly structured loans will likely eventually flow into traditional syndication, ABS, or private credit funds.

The long-term niche for on-chain lies in the areas where TradFi service costs are high, such as loan origination for small and medium-sized operators, cross-border stablecoin funding, loan share distribution, real-time equipment and utilization monitoring, and junior or mezzanine tranches bearing initial losses. A more likely future structure is on-chain protocols handling borrower discovery, asset monitoring, and organizing the first round of capital; once loan performance stabilizes, the senior portion is sold to traditional capital, while the chain retains service fees, subordinated risk, and data relationships.

Prepaid Service Rights

Prepaid Service Rights address the issuer's current funding needs. The platform sells future API capacity in advance, receives cash upfront, and then fulfills the service over time. As the tenor extends, the platform risk borne by the buyer becomes increasingly pronounced. Will the model still be competitive in a few years? To what level will service prices fall? Can the purchasing power of credits be maintained? All of these factors will affect today's valuation.

Therefore, these products are likely to gradually resemble standardized capacity contracts. Tenors will be shorter, service standards will be clearer, and models can be substituted according to agreed-upon rules. The clearer the contract, the easier it is for Prepaid Service Rights to develop a relatively independent market price.

Furthermore, there is an adverse selection problem here. Platforms with the best models, strongest customer demand, and most stable pricing power typically have no incentive to issue long-term, transferable inference rights. They prefer to retain the ability to price discriminate among customers, recognize revenue from unused credits, and adjust prices at will. The platforms most in need of pre-selling long-term service rights are often those that need financing and user acquisition, but whose future competitiveness has not yet been proven.

Capacity Incentives

Capacity Incentives address the early-stage supply organization problem in inference networks. Nodes need to configure GPUs and models in advance and bear operational costs. When demand-side hasn't stably formed, nodes lack sufficient incentive to come online early. Network token emissions can provide a period of revenue certainty for the supply side, allowing the network to first acquire usable capacity.

Long-term viability still depends on whether external client payments can gradually cover node revenue. A protocol can verify that an inference task was executed or that a node used a specified model, but the mere occurrence of computation doesn't create economic value. Token emissions can buy supply for a period of time, but customer demand still needs to be proven by real orders.

A more sensible incentive mechanism may gradually become tied to real revenue. Clients pay for tasks in stablecoins first; nodes receive revenue after completing inference and passing verification; the protocol then provides additional rewards based on external revenue, gradually reducing subsidies as the network matures. This allows emissions to cover cold-start costs while giving the market a clearer signal: how much of node revenue comes from clients versus capital subsidies. In the long run, the real task for Capacity Incentives is to transition from buying supply with tokens to sustaining supply with client revenue. Once this transition occurs, the network begins to have an economic foundation independent of token market cycles.

Single Model Revenue Rights

Model Revenue Rights are closest to the earlier mentioned crack spread:

Model Deployment Residual Profit = API Revenue − GPU Cost − Network, Storage, and Engineering Costs − Customer Acquisition and Subsidies

What many Model Tokens offer holders is primarily buybacks and staking incentives, which may not correspond to a legally enforceable claim on income. If there is no clear contractual link between API revenue and Token value, it becomes very difficult to establish a stable anchor for valuation.

Model lifespan further amplifies this problem. Models are updated rapidly; by the time an asset is financed and deployed, the underlying model may already be surpassed by newer alternatives. A single model therefore struggles to support a long-term financial asset.

A more likely development path is to gradually expand the underlying unit to a workload strategy. For example, a coding inference strategy could continuously switch underlying models, routing based on quality and cost. This way, the holder is exposed to a going-concern inference strategy, with the asset's lifespan tracking customer demand and workloads.

Additionally, for these products to enter the institutional market, they will need clear legal entities, audited API revenues, and a well-defined cash flow waterfall. Without these fundamentals, Model Token valuations will remain highly dependent on future usage and the issuer's buyback policies.

4.5.3 Routers Precede Token Futures in Providing Demand-Side Risk Management

Routers are an underappreciated layer in the financialization of inference.

What enterprises ultimately care about is the cost of a task. The specific model used is, in many scenarios, merely an implementation detail. A Router can continuously compare different models and providers, reallocating traffic as prices change. Underlying services are highly heterogeneous, but through Router orchestration, their substitutability gradually increases.

This directly impacts the scale of demand-side derivatives. Much cost volatility can first be absorbed through model switching and software optimization. Only the residual exposure that cannot be managed this way truly requires financial instruments.

Routers also have a significant data advantage: they see real invoices. Public API list prices rarely reflect an enterprise's actual procurement costs. A Router sees what clients actually pay and knows what quality of service corresponds to that price. As order flow accumulates, this data has the potential to form a quality-adjusted workload cost index.

Going a step further, a Router can directly offer enterprises cost caps. For example, an enterprise locks in a budget for its next three months of coding workload; the Router manages model selection and capacity procurement in the background, using GPU derivatives for residual risk when needed. The client ultimately gains budget certainty without having to manage a basket of models, futures positions, and basis themselves. This type of product may align better with enterprise procurement habits than generic Token futures. The Router can keep the underlying complexity internal while offering clients a more stable and predictable price.

Tokens themselves haven't formed a unified standard yet, but a Router can continuously seek substitute supply around a specific, well-defined workload. Once trading volume reaches a certain scale, the workload itself could become a new unit of price.

If this path holds, Routers will gradually extend from traffic allocation tools to procurement and risk management. Their most valuable long-term asset will also become increasingly clear: real orders, real transaction prices, and the ability to manage cost and quality across different models. The price standard for workloads is likely to emerge at this layer first, before entering indices and derivatives markets.

AI
Welcome to Join Odaily Official Community