# Models & Pricing

All prices are per million tokens. Checkpoint storage is charged at $0.10 per GB per month.

### What's changing

We now provide an 80% discount on cached prefill tokens.

Due to rising compute costs, we are also increasing our prefill and sample prices by ~50% and our train prices by ~10% starting July 17.

| Model | Tinker ID | Type | Arch | Size | Context | PrefillCached: 80% discount | Sample | Train |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Inkling](https://huggingface.co/thinkingmachines/Inkling) Limited-time 50% discount | thinkingmachines/Inkling | Hybrid + Audio + Vision | MoE | Large | 64K | $1.87$0.374 (cached)$1.87$0.374 (cached) | $4.68$4.68 | $5.61$5.61 |
| [Inkling (256K)](https://huggingface.co/thinkingmachines/Inkling) Limited-time 50% discount | thinkingmachines/Inkling:peft:262144 | Hybrid + Audio + Vision | MoE | Large | 256K | $3.74$0.748 (cached)$3.74$0.748 (cached) | $9.36$9.36 | $11.23$11.23 |
| [Nemotron-3-Ultra-550B-A55B](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16) Limited-time 50% discount | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | Hybrid | MoE | Large | 64K | $1.66$0.332 (cached)$2.49$0.498 (cached) | $4.15$6.225 | $4.98$5.478 |
| [Nemotron-3-Ultra-550B-A55B (256K)](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16) Limited-time 50% discount | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16:peft:262144 | Hybrid | MoE | Large | 256K | $3.32$0.664 (cached)$3.32$0.664 (cached) | $8.30$8.30 | $9.96$9.96 |
| [Nemotron-3-Super-120B-A12B](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16) Limited-time 50% discount | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | Hybrid | MoE | Large | 64K | $0.38$0.076 (cached)$0.57$0.114 (cached) | $0.96$1.44 | $1.16$1.276 |
| [Nemotron-3-Super-120B-A12B (256K)](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16) Limited-time 50% discount | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16:peft:262144 | Hybrid | MoE | Large | 256K | $0.76$0.152 (cached)$0.76$0.152 (cached) | $1.92$1.92 | $2.32$2.32 |
| [Nemotron-3-Nano-30B-A3B](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16) Limited-time 50% discount | nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 | Hybrid | MoE | Medium | 64K | $0.13$0.026 (cached)$0.195$0.039 (cached) | $0.33$0.495 | $0.40$0.44 |
| [Kimi-K2.6](https://huggingface.co/moonshotai/Kimi-K2.6) | moonshotai/Kimi-K2.6 | Hybrid + Vision | MoE | Large | 32K | $1.47$0.294 (cached)$2.205$0.441 (cached) | $3.66$5.49 | $4.40$4.84 |
| [Kimi-K2.6 (128K)](https://huggingface.co/moonshotai/Kimi-K2.6) | moonshotai/Kimi-K2.6:peft:131072 | Hybrid + Vision | MoE | Large | 128K | $5.15$1.03 (cached)$5.15$1.03 (cached) | $12.81$12.81 | $15.40$15.40 |
| [Kimi-K2.5](https://huggingface.co/moonshotai/Kimi-K2.5) [Retiring July 12](https://tinker-docs.thinkingmachines.ai/tinker/model-deprecations/) | moonshotai/Kimi-K2.5 | Hybrid + Vision | MoE | Large | 32K | $1.47$0.294 (cached)$2.205$0.441 (cached) | $3.66$5.49 | $4.40$4.84 |
| [Kimi-K2.5 (128K)](https://huggingface.co/moonshotai/Kimi-K2.5) [Retiring July 12](https://tinker-docs.thinkingmachines.ai/tinker/model-deprecations/) | moonshotai/Kimi-K2.5:peft:131072 | Hybrid + Vision | MoE | Large | 128K | $5.15$1.03 (cached)$5.15$1.03 (cached) | $12.81$12.81 | $15.40$15.40 |
| [Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) | Qwen/Qwen3.6-35B-A3B | Hybrid + Vision | MoE | Medium | 64K | $0.36$0.072 (cached)$0.54$0.108 (cached) | $0.89$1.335 | $1.07$1.177 |
| [Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B) | Qwen/Qwen3.6-27B | Hybrid + Vision | Dense | Medium | 64K | $1.24$0.248 (cached)$1.86$0.372 (cached) | $3.73$5.595 | $3.73$4.103 |
| [Qwen3.5-397B-A17B](https://huggingface.co/Qwen/Qwen3.5-397B-A17B) | Qwen/Qwen3.5-397B-A17B | Hybrid + Vision | MoE | Large | 64K | $2.00$0.40 (cached)$3.00$0.60 (cached) | $5.00$7.50 | $6.00$6.60 |
| [Qwen3.5-397B-A17B (256K)](https://huggingface.co/Qwen/Qwen3.5-397B-A17B) | Qwen/Qwen3.5-397B-A17B:peft:262144 | Hybrid + Vision | MoE | Large | 256K | $4.00$0.80 (cached)$4.00$0.80 (cached) | $10.00$10.00 | $12.00$12.00 |
| [Qwen3.5-35B-A3B-Base](https://huggingface.co/Qwen/Qwen3.5-35B-A3B-Base) | Qwen/Qwen3.5-35B-A3B-Base | Base | MoE | Medium | 64K | $0.36$0.072 (cached)$0.54$0.108 (cached) | $0.89$1.335 | $1.07$1.177 |
| [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) | Qwen/Qwen3.5-9B | Hybrid + Vision | Dense | Small | 64K | $0.44$0.088 (cached)$0.66$0.132 (cached) | $1.33$1.995 | $1.33$1.463 |
| [Qwen3.5-9B-Base](https://huggingface.co/Qwen/Qwen3.5-9B-Base) | Qwen/Qwen3.5-9B-Base | Base | Dense | Small | 64K | $0.44$0.088 (cached)$0.66$0.132 | $1.33$1.995 | $1.33$1.463 |
| [Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) | Qwen/Qwen3.5-4B | Hybrid + Vision | Dense | Compact | 64K | $0.22$0.044 (cached)$0.33$0.066 | $0.67$1.005 | $0.67$0.737 |
| [Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) | Qwen/Qwen3-8B | Hybrid | Dense | Small | 32K | $0.13$0.026 (cached)$0.195$0.039 | $0.40$0.60 | $0.40$0.44 |
| [GPT-OSS-120B](https://huggingface.co/openai/gpt-oss-120b) | openai/gpt-oss-120b | Reasoning | MoE | Medium | 32K | $0.18$0.036 (cached)$0.33$0.066 | $0.44$0.84 | $0.52$0.737 |
| [GPT-OSS-120B (128K)](https://huggingface.co/openai/gpt-oss-120b) | openai/gpt-oss-120b:peft:131072 | Reasoning | MoE | Medium | 128K | $0.63$0.126 (cached)$0.78$0.156 | $1.54$1.94 | $1.82$2.33 |
| [GPT-OSS-20B](https://huggingface.co/openai/gpt-oss-20b) | openai/gpt-oss-20b | Reasoning | MoE | Small | 32K | $0.12$0.024 (cached)$0.18$0.036 | $0.30$0.45 | $0.36$0.396 |
| [DeepSeek-V3.1](https://huggingface.co/deepseek-ai/DeepSeek-V3.1) | deepseek-ai/DeepSeek-V3.1 | Hybrid | MoE | Large | 32K | $1.13$0.226 (cached)$1.695$0.339 | $2.81$4.215 | $3.38$3.718 |

## Pricing Terms

- **Prefill**: Processing input/prompt tokens (forward pass only)
- **Cached prefill**: The smaller price under each prefill price; applies to input tokens that hit the prompt cache (80% off)
- **Sample**: Generating output tokens (forward pass + sampling)
- **Train**: Forward and backward pass for gradient computation
- **Context**: Maximum sequence length. Models with `:peft:` suffix support extended context at higher prices.
- **Tinker ID**: The exact string to pass to `create_lora_training_client(base_model=...)` or `create_sampling_client(base_model=...)`

MoE models are priced by active parameters, making them significantly more cost-effective than dense models of similar quality.

## Model Types

- **Base**: Raw pretrained models with no chat or instruction tuning. Best for post-training research or running the full post-training pipeline yourself.
- **Reasoning**: Always produce chain-of-thought before their answer. Highest intelligence, higher latency and token cost.
- **Hybrid**: Run in both thinking and non-thinking modes. They reason by default, but chain-of-thought can be disabled via a renderer or argument for faster, cheaper direct answers.
- **Vision**: Vision-language models that accept images alongside text. Shown as a `+ Vision` suffix on the underlying type (for example, `Hybrid + Vision`).
- **Audio**: Models that accept audio alongside text. Shown as a `+ Audio` suffix on the underlying type.

**Architecture** is either **Dense** (all parameters active per token) or **MoE** (mixture-of-experts, only a subset of parameters active per token). MoE models are highlighted in amber.

## Choosing a Model

- **Cost-effective**: Use MoE models (highlighted in amber)
- **Research/post-training**: Use Base models
- **Task-specific fine-tuning**: Start with a Hybrid model
- **Low latency**: Use a Hybrid model with chain-of-thought disabled
- **High intelligence**: Use Reasoning or Hybrid models (chain-of-thought)
- **Vision tasks**: Use models with Vision in the type

## Retired Models

These models have been retired and can no longer be used for training or inference, grouped by retirement date.

### June 12, 2026

- **Qwen:**`Qwen3-235B-A22B-Instruct-2507`, `Qwen3-VL-235B-A22B-Instruct`, `Qwen3.5-35B-A3B`, `Qwen3.5-27B`, `Qwen3-32B`, `Qwen3-30B-A3B`, `Qwen3-30B-A3B-Instruct-2507`, `Qwen3-VL-30B-A3B-Instruct`, `Qwen3-30B-A3B-Base`, `Qwen3-8B-Base`, `Qwen3-4B-Instruct-2507`
- **Llama:**`Llama-3.3-70B-Instruct`, `Llama-3.1-70B`, `Llama-3.1-8B`, `Llama-3.1-8B-Instruct`, `Llama-3.2-3B`, `Llama-3.2-1B`
- **DeepSeek:**`DeepSeek-V3.1-Base`
- **Kimi:**`Kimi-K2-Thinking`
