Compute Exchange
TOKEN FORWARDS

BUY TOKENS FORWARD.
SECURE YOUR CAPACITY.

Procure inference tokens in advance and tap them over a term of up to six months. Lock unit economics, commit to capacity terms, and secure supply ahead of demand across leading open-weight models.

HOW IT WORKS

COMMIT. LOCK. TAP.

01
COMMIT

Specify the open model, token volume (input / cached / output), term up to six months, and whether batch processing is acceptable. We aggregate committed-use quotes across the provider network.

02
LOCK

Lock a committed per-token rate for the full term. Settlement (upfront, milestone, or monthly), priority allocation, and rollover terms surface per quote so you can compare bilaterally.

03
TAP

Tap tokens against your committed balance over the term, against your real demand curve. Usage reconciles per the agreed settlement schedule.

CONTRACT SPECIFICATION

STANDARDIZED TOKEN UNITS

Compute Exchange standardizes inference commitments as Standardized Token Units (STUs). Buyers commit to a chosen STU volume; providers fulfill against the published methodology.

INPUT TOKEN
UNCACHED
1.00
STU
CACHED INPUT
REUSED CONTEXT
0.20
STU
OUTPUT TOKEN
GENERATED
4.21
STU
BATCH MODE
NON-REALTIME
0.50
× DRAW

Per-STU pricing varies by model family. Providers absorb the basis between the published STU methodology and their underlying per-token economics — buyers hold a single fungible commitment denominated in STUs.

PRINCIPLES

FORWARD VS ON-DEMAND

BUDGET CERTAINTY

Lock per-token unit economics for the full term. Forecast inference spend with confidence across the commitment.

PRIORITY ALLOCATION

Quotes surface each provider's priority and reservation terms for committed balances during demand spikes.

SUPPLY SECURITY

Secure token supply ahead of anticipated demand growth or open-model availability constraints.

FLEXIBLE TAP

Tap your committed balance against real usage over the term, with rollover terms surfaced per quote.

COVERAGE

OPEN MODELS ONLY

LARGE OPEN-WEIGHT

Flagship open models — Llama, DeepSeek, Qwen class — served across the provider network at committed volume.

SMALL & EFFICIENT

Distilled and small open models for high-volume, latency-sensitive, or cost-optimized inference.

MULTIMODAL & VISION

Open vision, speech, and multimodal models for document, image, and audio inference workloads.

EMBEDDING & SPECIALIZED

Open embedding, reranking, and classification models priced per standard inference billing units.

CALIBRATION REFERENCE

CONVERSIONS BY MODEL

The published STU methodology (1.0 · 0.2 · 4.21 · 0.5×) is one calibration — anchored on the Kimi K2 line. Other open models have different native input:output economics. Below: how 1 input, cached, and output token convert to STUs under each model's native ratio.

HOW STU PRICING WORKS · Conversion to STU is fixed per model; the per-STU price floats by provider based on hardware and underlying economics. Providers absorb the basis between native ratios and the published index at quote time.

IMPORTANT · Per-STU pricing varies materially by model family and provider. Quotes are indicative; final terms are confirmed bilaterally with the matched provider.

MODELPROVIDERINPUTCACHEDOUTPUT
Kimi-K2.6
VISION
Moonshot AI1.000.204.21
gpt-oss-120b
TEXT
OpenAI1.00varies4.00
Nemotron-3-Nano-Omni
TEXT
NVIDIA1.00varies4.00
MiniMax-M2.5
TEXT
Minimax1.00varies4.00
GLM-5.2
TEXT
Z.ai1.00varies3.14
GLM-5.1
TEXT
Z.ai1.00varies3.14
Hermes-4-70B
TEXT
Nous Research1.00varies3.08
Nemotron-3-Ultra-550b-a55b
TEXT
NVIDIA1.00varies3.00
Cosmos3-Super-Reasoner
VISION
NVIDIA1.00varies3.00
Nemotron-3-Super-120b-a12b
TEXT
NVIDIA1.00varies3.00
Hermes-4-405B
TEXT
Nous Research1.00varies3.00
Qwen3-235B-A22B-Instruct-2507
TEXT
Qwen1.00varies3.00
DeepSeek-V4-Pro
TEXT
DeepSeek1.00varies2.00
MiniCPM-V-4.5
VISION
OpenBMB1.00varies1.69
Qwen3.5-397B-A17B
TEXT
Qwen1.00varies6.00

Native ratios derived from active open-model provider catalogs. Cached-input ratio shown only where the provider publishes one — most non-Kimi endpoints don't currently expose cached pricing publicly, so the cell reads "varies." Quotes price against the published 4.21× index; providers absorb the basis between native ratios and the index.

Frequently Asked Questions

TOKEN FORWARDS, EXPLAINED

What is a token forward?+
A token forward is a committed-use agreement to procure a fixed volume of inference tokens in advance, tapped over an agreed term of up to six months. Companies with predictable inference demand lock in unit economics and guaranteed capacity; providers gain committed volume. Compute Exchange aggregates token forward agreements across the open-model provider network.
What term lengths are available?+
Token forwards run for terms of up to six months. Shorter terms of one to three months suit pilot programs and seasonal demand; the full six-month term suits production workloads with stable, forecastable inference volume. Compute Exchange structures the term and tap schedule around your demand curve.
What happens if I do not use all my committed tokens?+
Tap terms vary by provider. Some agreements allow unused balance to roll over within the term; others treat the commitment as use-it-or-lose-it at term end. Compute Exchange surfaces the tap and rollover terms for each quote so you can compare before committing. Structure the commitment conservatively against your forecast inference demand.
Which models can I procure tokens for?+
Token forwards cover open-weight models only — large flagship open models such as Llama, DeepSeek, and Qwen class, small and efficient open models, multimodal and vision models, and embedding and specialized models. Available model coverage depends on the providers in the network; Compute Exchange confirms model availability during the quote process.
How is this different from Reserved GPU Rental?+
Reserved GPU Rental commits you to raw GPU capacity that you operate yourself — you run the models, manage the infrastructure, and own utilization risk. Token forwards commit you to inference output at the model layer, with the provider operating the infrastructure. Choose GPU rental for custom models and full control; choose token forwards for managed open-model inference with predictable per-token economics.
How are token forwards settled?+
Pricing is quoted as a committed per-token rate, with input and output tokens priced separately per standard inference billing. Settlement terms — upfront, milestone, or monthly — are negotiated bilaterally with the provider. Compute Exchange facilitates the introduction and normalizes quotes across providers; the contract is directly between you and the provider.

SECURE YOUR TOKEN SUPPLY

Submit a commitment request and Compute Exchange returns a token forward quote across the verified open-model provider network.

DISCLAIMER

DISCLAIMER: Token forward quotes are aggregated from verified third-party inference providers serving open-weight models. Compute Exchange facilitates introductions but does not operate inference infrastructure, take ownership of token balances, or guarantee model availability or SLA. All commitment terms — including pricing, tap, rollover, and settlement — are negotiated directly between buyer and provider.