BService
AI model tokens at volume
An AI model token is the unit that LLM usage is measured and billed in: providers charge per million input and output tokens. DaoWorks sources that usage for teams running inference at scale — per-token APIs, dedicated endpoints and committed-spend agreements — and negotiates the terms. The buyer pays us a commission only when a deal closes.
Updated
A note on the word “token”
Not crypto.
On this site, a token always means a unit of AI model usage. DaoWorks doesn't offer crypto or other digital-asset services.
How usage is billed
- Unit
- 1M tokens
- Input
- Your prompt and context
- Output
- The model's reply
- Cached input
- Repeated context, often cheaper
01What we source
Three ways to buy tokens.
Per-token APIs
Pay-as-you-go or discounted access to open-weight models such as Llama, Qwen, DeepSeek, Mistral and gpt-oss, hosted by inference providers — and to proprietary models through their authorized channels.
Good for
Variable workloads, many models, a fast start.
Dedicated endpoints
A model deployed on GPUs reserved for you alone, priced by the hour or by throughput. Predictable latency and no shared rate limits.
Good for
Steady high volume, latency-sensitive products, fine-tuned models.
Committed-spend agreements
A lower per-token price in exchange for a monthly or annual commitment, with agreed rules for overage and unused volume.
Good for
Teams whose volume is predictable.
02Your brief
What to send us.
Last month's usage export from your current provider is the best brief there is.
- The models you use now, or the task and the quality bar
- Monthly volume, split into input and output tokens
- Peak load: requests and tokens per minute
- Latency targets: time to first token, output speed
- The context length you actually use
- Data handling: retention, training opt-out, processing region
- The API you need to stay compatible with, for example OpenAI-compatible
- Budget, and how long you can commit
03Comparison
How we compare offers.
A lower price per token is worth nothing if the model is quantized harder, the context is shorter or the rate limit caps you at peak. We compare on all of it.
| Dimension | What we pin down |
|---|---|
| Price | Per million input and output tokens, cached-input pricing and any minimums — on your real input/output mix, not a list price. |
| Model fidelity | Exact model version, quantization and maximum context length. The same model name can be served very differently. |
| Throughput | Rate limits in requests and tokens per minute, and how bursts are handled. |
| Latency | Time to first token and output speed on prompts like yours. |
| Reliability | Uptime commitment, service credits, and what happens when a model is overloaded. |
| Data | Retention, whether your data is used for training, processing region, and certifications such as SOC 2. |
| Commitment | Minimum spend, term, overage pricing, and what happens to unused volume. |
04Questions
AI model tokens FAQ
What is an AI model token?
A token is a chunk of text, often part of a word, that a language model reads or writes. Providers meter usage in tokens and price it per million input tokens and per million output tokens, with output usually costing more.
Are these crypto tokens?
No. They are units of AI model usage. DaoWorks doesn't offer crypto or other digital-asset services.
Do you resell API access?
No. You contract with the provider and get your API keys and invoices from them. DaoWorks negotiates the terms, and the buyer pays us a commission when the deal closes.
Can you get better pricing on proprietary models?
Where a model's owner or its authorized resellers offer committed-use pricing, we can negotiate it for you. We don't source proprietary model access through unauthorized channels.
Sources Google Gemini API: Understand and count tokens · Hugging Face: Summary of the tokenizers
→Next step
Send us a token brief.
GPU type and count, or model and monthly tokens. Term, region, budget. We come back with what is actually available and on what terms.