umans/status/umans-deepseek-v4-flash
Live · refreshes every 30s
← all models
Umans DeepSeek V4 Flash (experimental) Experimental
umans-deepseek-v4-flash · DeepSeek-V4-Flash · DeepSeek
In testing
throughput · p50 · no requests
TTFT · p50 · no requests
uptime · 24h

DeepSeek V4 Flash as a Labs experiment, open for a short test window: temporary, not a permanent model. DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release with speculative decoding for speed, on a 1M-token context. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Access is seat-gated through the Labs page while an experiment is live. It runs on a single node with no fallback at limited capacity and availability, so expect it to be flaky and to go down under load: crash it, give it a moment, and try again. For production work we recommend umans-coder or umans-glm-5.2.

no uptime history yet
90 days agotoday
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/low/high/max
Trends

Speed over the last 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
gathering data
90 days agotoday
TTFT p50 · time to first token, lower is better
gathering data
90 days agotoday
Changelog

Events for Umans DeepSeek V4 Flash (experimental)

incl. gateway-wide announcements
No recent events.