{"umans-kimi-k2.7":{"name":"umans-kimi-k2.7","display_name":"Umans Kimi K2.7 Code","deprecation":{"sunset_date":"2026-08-10","replacement":"umans-kimi-k3"},"description":"Kimi K2.7-Code via Umans Code - Moonshot's strongest coding model and the successor to Kimi K2.6. Built for complex, tool-heavy agentic coding; it reasons more efficiently than K2.6, so agent sessions run faster at the same depth. Deprecated: use `umans-kimi-k3` instead (sunset 2026-08-10).","base_model":{"name":"kimi-k2.7-code","provider":"Moonshot","oss_base":"Kimi K2.7-Code"},"capabilities":{"max_completion_tokens":262144,"recommended_max_tokens":32768,"context_window":262144,"supports_vision":true,"supports_tools":true,"reasoning":{"supported":true,"can_disable":false,"levels":[],"default_level":null}},"benchmarks":{},"weights":{"precision":"full","hf_url":"https://huggingface.co/moonshotai/Kimi-K2.7-Code"}},"umans-glm-5.2":{"name":"umans-glm-5.2","display_name":"Umans GLM 5.2","description":"GLM 5.2 is our best model for coding right now, with a 400K context window for large codebases. Vision is available on the Anthropic Messages API (`/v1/messages`) only, through a server-side handoff (GLM 5.2 generates the text, Kimi preprocesses the image); that handoff will be retired soon in favour of more efficient client-side image handling.","base_model":{"name":"GLM-5.2","family":"GLM"},"capabilities":{"max_completion_tokens":131072,"recommended_max_tokens":131071,"context_window":405504,"supports_vision":"via-handoff","supports_tools":true,"reasoning":{"supported":true,"can_disable":true,"levels":["none","high","max"],"default_level":"high"}},"benchmarks":{},"weights":{"precision":"fp8","hf_url":"https://huggingface.co/zai-org/GLM-5.2-FP8"}},"umans-coder":{"name":"umans-coder","display_name":"Umans Coder","description":"Our recommended model for complex, coding-heavy workloads. Optimized for the best experience with coding agents like Claude Code.","base_model":{"name":"kimi-k2.7-code","provider":"Moonshot","oss_base":"Kimi K2.7-Code"},"capabilities":{"max_completion_tokens":262144,"recommended_max_tokens":32768,"context_window":262144,"supports_vision":true,"supports_tools":true,"reasoning":{"supported":true,"can_disable":false,"levels":[],"default_level":null}},"benchmarks":{},"weights":{"precision":"full","hf_url":"https://huggingface.co/moonshotai/Kimi-K2.7-Code"}},"umans-deepseek-v4-flash-0731":{"name":"umans-deepseek-v4-flash-0731","display_name":"Umans DeepSeek V4 Flash","lifecycle":{"production_start_date":"2026-08-03"},"description":"DeepSeek V4 Flash: DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release on a 1M-token context. The cheapest production model in the lineup for real agentic work. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Served on our own GPU infrastructure with high availability.","base_model":{"name":"DeepSeek-V4-Flash","provider":"DeepSeek","oss_base":"DeepSeek-V4-Flash"},"capabilities":{"max_completion_tokens":393216,"recommended_max_tokens":393215,"context_window":1048576,"supports_vision":false,"supports_tools":true,"reasoning":{"supported":true,"can_disable":true,"levels":["none","low","high","max"],"default_level":"low"}},"benchmarks":{},"weights":{"precision":"full","hf_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731"}},"umans-deepseek-v4-flash-0731-vision-lab":{"name":"umans-deepseek-v4-flash-0731-vision-lab","display_name":"Umans DeepSeek V4 Flash Vision (lab)","stage":"playground","description":"DeepSeek V4 Flash Vision as a Labs experiment, open for a short test window: temporary, not a permanent id. Our own V4 Flash with vision added: the same performance and speed as DeepSeek V4 Flash (284B total, 13B active, 1M-token context), plus the ability to read images. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Access is seat-gated through the Labs page while an experiment is live. It is offered at limited capacity and low availability, and it is served on our own GPU infrastructure, so expect it to be flaky and to go down under load: crash it, give it a moment, and try again.","base_model":{"name":"DeepSeek-V4-Flash-Vision","provider":"DeepSeek","oss_base":"DeepSeek-V4-Flash-Vision"},"capabilities":{"max_completion_tokens":393216,"recommended_max_tokens":393215,"context_window":1048576,"supports_vision":true,"supports_tools":true,"reasoning":{"supported":true,"can_disable":true,"levels":["none","low","high","max"],"default_level":"low"}},"benchmarks":{},"weights":{"precision":"full","hf_url":"https://huggingface.co/umans-ai/DeepSeek-V4-Flash-0731-Vision"}},"umans-kimi-k3":{"name":"umans-kimi-k3","display_name":"Umans Kimi K3","lifecycle":{"production_start_date":"2026-07-31"},"description":"Kimi K3: Moonshot's most capable model and the first open 3T-class release - 2.8T parameters, a 1M-token context window, and native vision, built for repository-scale understanding and long agentic runs at closed-frontier quality (92.4% vs Claude Fable 5's 92.6% across ~1,030 agentic tasks in Fireworks' independent study). It thinks by default at maximum reasoning effort; select none, low, high, or max to trade depth for speed.","base_model":{"name":"Kimi-K3","provider":"Moonshot","oss_base":"Kimi K3"},"capabilities":{"max_completion_tokens":131072,"recommended_max_tokens":131071,"context_window":1048576,"supports_vision":true,"supports_tools":true,"reasoning":{"supported":true,"can_disable":true,"levels":["none","low","high","max"],"default_level":"max"}},"benchmarks":{},"weights":{"precision":"full","hf_url":"https://huggingface.co/moonshotai/Kimi-K3"}},"umans-flash":{"name":"umans-flash","display_name":"Umans Flash","description":"Our fastest model: a light workflow complement, not a standalone coder. Think Haiku next to Opus: not everything needs a frontier model, and Flash's speed (200+ tokens per second) compounds on the roles around umans-coder: gathering context, scout subagents, research, summaries, documentation, and quick edits.","base_model":{"name":"Qwen3.6-35B-A3B","provider":"Qwen","oss_base":"Qwen3.6-35B-A3B"},"capabilities":{"max_completion_tokens":262144,"recommended_max_tokens":32768,"context_window":262144,"supports_vision":true,"supports_tools":true,"reasoning":{"supported":true,"can_disable":true,"levels":["none","low","medium","high"],"default_level":"medium"}},"benchmarks":{},"weights":{"precision":"fp8","hf_url":"https://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8"}},"umans-qwen3.6-35b-a3b":{"name":"umans-qwen3.6-35b-a3b","display_name":"Umans Qwen3.6 35B A3B","description":"Technical alias for `umans-flash`. Use `umans-flash` for the recommended fast coding-agent experience.","base_model":{"name":"Qwen3.6-35B-A3B","provider":"Qwen","oss_base":"Qwen3.6-35B-A3B"},"capabilities":{"max_completion_tokens":262144,"recommended_max_tokens":32768,"context_window":262144,"supports_vision":true,"supports_tools":true,"reasoning":{"supported":true,"can_disable":true,"levels":["none","low","medium","high"],"default_level":"medium"}},"benchmarks":{},"weights":{"precision":"fp8","hf_url":"https://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8"}}}