2026-07-02 · Thamodharan Ganesan
TextMate AI isn't locked to a single model. Every chat lets you pick from 7 models across multiple providers, two of which are available on the Free plan.
| Model | Best for |
|---|---|
| Mistral Large (default) | Balanced, general-purpose, largest context of the free models — the right default for most conversations. 262K context window. |
| GPT-OSS 20B | A lightweight, fast model for quick questions where you don't need Mistral Large's full context. |
| Model | Best for |
|---|---|
| DeepSeek V3.1 | Strong at reasoning and coding — reach for it on harder technical questions. |
| GLM 4.7 | Served on Cerebras for the fastest inference in the lineup — great when you want a quick answer from a different model family. |
| Llama 3.3 70B | Served on Groq for the lowest latency in the lineup. |
| Nemotron 3 Ultra | NVIDIA's largest model here — a huge context window for feeding it long documents or conversations. |
| Gemma 4 31B | Google's model, served on Cerebras for ultra-fast inference — another quick second opinion from a different model family. |
Every model has a per-response output cap. Mistral Large, DeepSeek V3.1, and Nemotron 3 Ultra support up to 65,536 output tokens — enough for long rewrites or documents. GPT-OSS 20B supports up to 120,000. GLM 4.7 caps at 40,000, Llama 3.3 70B at 4,096 (Groq's free tier is the tightest of the lineup), and Gemma 4 31B uses a conservative 8,192-token cap. If you're generating long-form content in one shot, Mistral Large, DeepSeek V3.1, or Nemotron 3 Ultra are the models built for that.
See the Models page for the live picker with plan requirements, or Pricing for what Free vs Pro unlocks.