Hardware presets
Card = tokens/sec + $/tok/s. Tap to load. The slider stays yours — presets only set the starting point.
Green $/tok/s = cheapest speed on the list. tok/s are typical published community ranges (±30%) for the model class shown — a 120B dense model will be far slower than an 8B one on identical hardware. Prices are Sep 2026 street; verify before buying.
Realism
Jitter Instantaneous rate wobbles ±50%, average stays on the slider
Cold start First message also loads weights from disk
Show live counters Running elapsed / tokens while it streams
Reset
Reset to preset defaults
How to read it
Faster than you can read (100+ tok/s) — feels like a hosted API.
Snappy (60–100) — no waiting perceptible on short answers.
Comfortable (30–60) — fine unless you're generating 2,000 tokens.
Noticeable (15–30) — you wait, but it reads as "thinking".
Slow (8–15) — you'll browse another tab.
Painful (<8) — unusable for chat, fine for batch.
The number that actually decides it is the 600-token column: a normal answer. If that's over ~20 s you'll feel it every single turn.