Models
Every chat lets you pick from 7 models across multiple providers. 2 are free; the rest unlock on Pro.
Free models
- Mistral Large (default) — balanced & fast, 262K context window, up to 65,536 output tokens
- GPT-OSS 20B — lightweight GPT-OSS, up to 120,000 output tokens
Pro models
- DeepSeek V3.1 — strong reasoning & coding, served on SambaNova
- GLM 4.7 — served on Cerebras for the fastest inference in the lineup
- Llama 3.3 70B — served on Groq for the lowest latency in the lineup
- Nemotron 3 Ultra — NVIDIA, up to a 1M-token context window
- Gemma 4 31B — Google, served on Cerebras for ultra-fast inference