Grok 4.6: 500K-token context window, $2 input / $6 output per 1M tokens (same as 4.5), AA Intelligence Index 61 (tied with GPT-5.6 Sol Max), GDPVal-AA v2 1753 Elo (highest overall), but DeepSWE and Terminal-Bench still trail GPT-5.6 Sol and Claude Fable 5
Muse Spark 1.2: 1M context window, input $1.25 / output $4.25 per 1M tokens (same as 1.1), AA Intelligence Index 57, GDPval-AA v2 Elo jumps 260 points to 1631 (5th overall), paired with Meta's first code agent Muse Code for long-running multi-agent collaboration
Wan3.0: single-shot length doubles from Wan2.7's 15s to 30s, up to 1080P, supports doc/xls/ppt/pdf/md files and web pages as generation inputs, priced at 480P $0.05 / 720P $0.10 / 1080P $0.20 per second — roughly 50% cheaper than Google Veo 3.1 Standard, but now closed-source API-only, and not yet independently tested by third parties
Qwen3.8-Flash-Next: open-weight preview of the Qwen4 architecture, 125B total parameters with only 6B active (plus a 51B N-gram embedding), 262K native context extensible to 1M, Qwen Community License 1.0 (not Apache 2.0). Official benchmarks show it beating both its own 27B dense model and the 397B Qwen3.7-Plus on agentic coding (DeepSWE 1.1: 58.7) and scoring highest on CoWorkBench long-horizon office tasks (73.9) — but no official API pricing or independent third-party testing exists yet
GLM-5.3-Flash: 320B total / 18B active parameters (MoE), 1M context / 131K max output, natively accepts text + image + video input, MIT-licensed weights on HuggingFace; standard pricing $0.15 input / $0.50 output per 1M tokens (50% launch discount to $0.075/$0.25 through Sept 9), roughly 90% cheaper than sibling model GLM-5.3; Terminal-Bench 2.1 hits 84.3 (just behind Opus 4.8's 85.0), DeepSWE 1.1 jumps from GLM-5.2's 46.2 to 63.4; under its 'Ox Alpha' alias it briefly took the #1 weekly token share spot on OpenRouter
Tencent Hy4 preview: 770B total / 49B active parameters (MoE, 78 layers), 1,048,576-token context window; API pricing $0.834 input / $2.501 output per 1M tokens (cache hit $0.042); Apache 2.0 open weights on HuggingFace; a 163-engineer blind eval scores it 2.99/4.00, just ahead of GLM-5.3 (2.92) and Kimi K3 (2.94); third-party aggregator BenchLM scores it 79.2/100, ranked #7 of 228 models; Tencent discloses for the first time that the model helped optimize its own training pipeline and inference system, lifting throughput 31.8%
Breeze TTS 2: open weights (Apache 2.0 code, research/non-commercial model license), #1 open-weights model on Artificial Analysis Provider Voices (1,215 Elo, +90 over Fish Audio S2 Pro), #1 on both Voice Design (Role Fit 78.02) and Voice Direction (4.25) benchmarks; TTFA p50 133.6ms / p95 163.3ms, RTF 0.32 on H100; hosted API priced at $34 per 1M characters (over 2x Fish Audio S2 Pro); supports 50 languages, commercial use requires a separate license from RESONIA, INC.
DeepSeek-V4-Flash-Vision-Exp: 284B total / 13B active-parameter MoE, 1M context, 384K max output; priced identically to plain V4-Flash (Input $0.44 / Output $1.32 at peak, half that off-peak); wins 6 of 7 text-agent benchmarks against its own predecessor (DeepSWE hits 59.3, edging past Opus-4.8's 58.0); multimodal-agent scores close in on Opus-4.8 (trails by 2.9 on ApexBench, actually leads on ZeroBench); each image is capped at 384 tokens / roughly 800×800 resolution, trading fine detail for near-zero cost
Gemini Omni 1.1 Flash (gemini-omni-1.1-flash): GA since 2026-08-27, replacing the preview that launched 6/30; closed-source, priced per output second: $0.03 at 360p, $0.10 at 720p (default), $0.15 at 1080p, $0.30 at 4K (1080p/4K are upscaled); ranks #1 on the Artificial Analysis Text-to-Video Arena without audio (1322 Elo) and #2 with audio (1237, trailing Wan3.0's 1241); adds scene extension (10s of context, chainable up to 40s total) and first/last-frame interpolation for camera control
Muse Voice Transcribe (muse-voice-transcribe-1.0): Meta Superintelligence Labs' first real-time audio perception model, launched 2026-09-01; closed-source, API-only, $0.18/hour of audio ($3.00 per 1,000 minutes); 3.1% final-transcript WER on streaming (#1 on Artificial Analysis AA-WER Streaming, ahead of Cartesia Ink-2's 3.4%), 0.16s delay from end-of-speech to final transcript; one model does ASR, 20+ speaker diarization, and endpointing together, replacing what used to require three separate systems
K2 Horizon 375B-A23B (IFM/K2-Horizon-375B-A23B): released 2026-09-03 by the Institute of Foundation Models (under MBZUAI, Abu Dhabi), 375B total / 23B active parameters (MoE), 512K (524,288 tokens) native context, fully open under Apache-2.0 (weights, code, training data recipes, and intermediate checkpoints all public), no official API pricing (open weights, self-hosted); Terminal-Bench 2.1 70.2%, SWE Bench Pro 42.6%, SWE-Atlas-QnA 48.4% (highest of any model tested, open or closed); shipped alongside five sibling sizes — 36B-A4B (new MoVA sparse attention), 32B, 7B, 3.7B, 0.9B