All systems operational
Live status and speed of the Umans Code gateway and its models, refreshed every 30 seconds. For each model we show its output speed and median time to first token.
API endpoint
Models in production
tok/s = output tokens per second · TTFT = time to first token · p50 = median over the last 5 minutes. On each gauge the midpoint is that model's target; the marker sits further right when it's beating target (faster TTFT, higher throughput).
Operational68.8tok/sthroughput · p50 · last 5 min1.96sTTFT · p50 · last 5 min99.93%uptime · 24h
GLM 5.2 is our best model for coding right now, with a 400K context window for large codebases. Vision is available on the Anthropic Messages API (`/v1/messages`) only, through a server-side handoff (GLM 5.2 generates the text, Kimi preprocesses the image); that handoff will be retired soon in favour of more efficient client-side image handling.
90-day speed trends & events →Operational116.8tok/sthroughput · p50 · last 5 min2.24sTTFT · p50 · last 5 min99.58%uptime · 24h
Kimi K2.7-Code via Umans Code - Moonshot's strongest coding model and the successor to Kimi K2.6. Built for complex, tool-heavy agentic coding; it reasons more efficiently than K2.6, so agent sessions run faster at the same depth.
90-day speed trends & events →Umans Flash Fastestalso served as umans-qwen3.6-35b-a3bOperational384.0tok/sthroughput · p50 · last 5 min509msTTFT · p50 · last 5 min99.40%uptime · 24h
Our fastest model: a light workflow complement, not a standalone coder. Think Haiku next to Opus: not everything needs a frontier model, and Flash's speed (200+ tokens per second) compounds on the roles around umans-coder: gathering context, scout subagents, research, summaries, documentation, and quick edits.
90-day speed trends & events →In the playground
Umans Kimi K3 ExperimentalIn testing56.6tok/sthroughput · p50 · last 5 min4.52sTTFT · p50 · last 5 min99.86%in testing
Kimi K3: Moonshot's most capable model and the first open 3T-class release - 2.8T parameters, a 1M-token context window, and native vision, built for repository-scale understanding and long agentic runs at closed-frontier quality (92.4% vs Claude Fable 5's 92.6% across ~1,030 agentic tasks in Fireworks' independent study). It thinks by default at maximum reasoning effort; select none, low, high, or max to trade depth for speed.
90-day speed trends & events →Umans DeepSeek V4 Flash (experimental) ExperimentalIn testing…throughput · p50 · no requests…TTFT · p50 · no requests…in testing
DeepSeek V4 Flash as a Labs experiment, open for a short test window: temporary, not a permanent model. DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release with speculative decoding for speed, on a 1M-token context. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Access is seat-gated through the Labs page while an experiment is live. It runs on a single node with no fallback at limited capacity and availability, so expect it to be flaky and to go down under load: crash it, give it a moment, and try again. For production work we recommend umans-coder or umans-glm-5.2.
90-day speed trends & events →