Grok 4.6
Keeps xAI’s rapid frontier-model cadence visible while full model-card, benchmark and pricing details remain to be verified.
Frontier model brief · Checked several times daily
A signal-first view of the newest models from the labs shaping AI—what launched, what changed, and why it matters.
No major launch signal is still a result: the refresh ran and the index stayed conservative.
Editorial check queue
These launches are visible on the radar, but at least one important field still needs stronger first-party confirmation before treating it as fully verified.
Keeps xAI’s rapid frontier-model cadence visible while full model-card, benchmark and pricing details remain to be verified.
Marks Meta’s return to frontier APIs with native multi-agent orchestration and active context management.
Shows a compact model moving beyond screens into autonomous navigation across real environments.
Verified release stream
Only first-party announcements. No composite scores, rumours or leaderboard theatre.
High-signal launch with broad capability, access or platform-level implications.
Google’s latest Flash model is generally available in the Gemini API with improved multimodal and coding performance.
Primary model documentation is available and linked from this brief.
Strong source · 2 receiptsMoves the Flash tier forward from 3.6 to 3.7 and gives developers a newer low-latency default with an introductory price window.
Potentially important, but source strength or availability still needs editorial confirmation.
xAI’s latest Grok model release, announced by SpaceXAI on X as the successor to Grok 4.5.
Keeps xAI’s rapid frontier-model cadence visible while full model-card, benchmark and pricing details remain to be verified.
High-signal launch with broad capability, access or platform-level implications.
Google’s newest Flash model, tuned for token-efficient coding, knowledge work and multimodal tasks.
Primary model documentation is available and linked from this brief.
Strong source · 1 receiptPushes a fast model into long-horizon agents and computer-use workflows without Pro-tier pricing.
High-signal launch with broad capability, access or platform-level implications.
A fast frontier model built around coding, agentic execution and professional knowledge work.
Release is backed by the lab’s own announcement or newsroom post.
First-party source · 1 receiptPairs deep software-engineering ability with an advertised serving speed of 80 tokens per second.
High-signal launch with broad capability, access or platform-level implications.
A three-tier family: Sol for frontier work, Terra for balance and Luna for fast, efficient tasks.
Release is backed by the lab’s own announcement or newsroom post.
First-party source · 1 receiptBrings stronger multi-agent work and programmatic tool coordination across ChatGPT, Codex and the API.
Potentially important, but source strength or availability still needs editorial confirmation.
Meta’s multimodal reasoning model for agentic tasks, computer use, coding and rich media understanding.
Tracked on the radar while at least one comparable field awaits stronger source confirmation.
Editorial review · 1 receiptMarks Meta’s return to frontier APIs with native multi-agent orchestration and active context management.
Potentially important, but source strength or availability still needs editorial confirmation.
An embodied navigation model that follows language instructions using a single ordinary RGB camera.
Tracked on the radar while at least one comparable field awaits stronger source confirmation.
Editorial review · 1 receiptShows a compact model moving beyond screens into autonomous navigation across real environments.
High-signal launch with broad capability, access or platform-level implications.
Anthropic’s fifth-generation Sonnet model for autonomous agents, coding, tool use and knowledge work.
Release is backed by the lab’s own announcement or newsroom post.
First-party source · 1 receiptBrings performance close to Opus 4.8 at a lower price, with stronger planning and long-running execution.
High-signal launch with broad capability, access or platform-level implications.
An open flagship model designed to hold context and execute over long software-engineering horizons.
Release is backed by the lab’s own announcement or newsroom post.
First-party source · 1 receiptCombines a stable million-token window with multiple reasoning-effort levels and openly released weights.
High-signal launch with broad capability, access or platform-level implications.
Cohere’s first agentic coding model, built to run locally with a small active compute footprint.
Release is backed by the lab’s own announcement or newsroom post.
First-party source · 1 receiptMakes sovereign coding agents practical: 30B total parameters, but only 3B active at inference time.
High-signal launch with broad capability, access or platform-level implications.
Anthropic’s fifth-generation frontier model for sustained coding, science, vision and demanding knowledge work.
Release is backed by the lab’s own announcement or newsroom post.
First-party source · 1 receiptIntroduces a capability tier above Opus for complex, days-long and asynchronous tasks.
High-signal launch with broad capability, access or platform-level implications.
A natively multimodal frontier model for coding and agents, built on MiniMax Sparse Attention.
Release is backed by the lab’s own announcement or newsroom post.
First-party source · 1 receiptCombines open weights, a million-token context window and long-horizon agentic execution in one model.
Material model update, access expansion, pricing shift or comparable capability change.
A new API generation offered as V4-Pro and V4-Flash, with both thinking and non-thinking modes.
Release is backed by first-party product or API change documentation.
Strong source · 1 receiptKeeps compatibility with both OpenAI- and Anthropic-style interfaces, lowering the cost of switching.
Reading the radar
“Latest” means the newest named model announcement we could verify from each lab’s own newsroom or documentation. Release dates reflect the original announcement, not later availability updates.
Claims are deliberately attributed to the releasing company. Benchmark results are not normalized across labs, so this index does not rank models or declare a winner.