If you are an AI engineer or technical lead waiting on Grok 4.6 to upgrade your agent or coding workflow, Elon Musk's July 28 reply to Vercel CEO Guillermo Rauch on X is the only on-the-record source so far: xAI is targeting Grok 4.6 around August 7 with 1.5 trillion parameters, followed a few weeks later by Grok 4.7 at 2.1 trillion. That is roughly one month after Grok 4.5 shipped — three frontier models in about two months. This article maps the timeline, explains why xAI is emphasizing SFT/RL over raw scale, compares the announced specs against Kimi K3 and rumored Claude Fable 5.1, and flags what remains unverified. Bottom line: no benchmarks, pricing, or model card exist yet — treat "around August 7" as a target, not a guarantee.
01 Four selection risks before Grok 4.6 actually ships
Unlike Grok 4.5, which launched with a full model card and 15 tracked benchmarks, Grok 4.6 and 4.7 have no independent evaluation or official product page. Betting your production stack on social-media leaks creates real exposure:
- Single-source, unconfirmed claims: Everything — "1.5 trillion parameters," "significantly improved SFT & RL," the 2.1T Grok 4.7 follow-up — comes from one X post. No xAI blog post or model card has corroborated it.
- "Musk time" has a track record: Timelines from Musk-led companies have historically slipped by days to weeks. "Around August 7" is a target, not a contract date.
- Benchmarks and pricing are blank: Grok 4.5 shipped with detailed scores; Grok 4.6 has zero third-party data. You cannot yet compare it head-to-head with Kimi K3, which already has verified numbers.
- Safety and industry pacing collide: The same day Musk posted the roadmap, 1,200+ employees at OpenAI, Anthropic, Google DeepMind, and Meta published the "Pacing the Frontier" letter asking the US government to help slow automated AI development — endorsed at the corporate level by OpenAI and Anthropic. xAI is absent from that list. In July, xAI also sued a user for allegedly using Grok to generate CSAM — relevant vendor-risk context for enterprise buyers.
In one sentence: the narrative is clear, but verifiable data is zero — any "best model" claim before an official model card is speculation.
02 What's confirmed vs. what's just Musk's word
Key timeline:
- July 8, 2026: xAI ships Grok 4.5 — built for coding and agentic work, co-trained with Cursor on real developer sessions. 500K-token context, $2/$6 per million input/output tokens, published model card.
- July 16–26, 2026: Moonshot AI's Kimi K3 moves from hosted preview to full open-weight release — 2.8T parameters, 1M-token context, topping Hugging Face's trending chart.
- July 28, 2026: Musk posts the Grok 4.6/4.7 roadmap in reply to Rauch — the primary source for this article.
- ~August 7, 2026 (target): Grok 4.6, 1.5T parameters, positioned as an SFT/RL upgrade rather than a raw scale-up.
- Late August–early September 2026 (estimated): Grok 4.7, 2.1T parameters — Musk says "better than 4.6 in every way, except slightly slower to serve, albeit with even better token efficiency."
| Model | Date | Parameters | Focus | Status |
|---|---|---|---|---|
| Grok 4.3 Beta | Apr 17, 2026 | Undisclosed | Baseline | Shipped |
| Grok 4.5 | Jul 8, 2026 | Undisclosed (single SKU, not MoE) | Coding/agentic, Cursor co-training | Shipped, benchmarked |
| Grok 4.6 | ~Aug 7, 2026 | 1.5T | SFT/RL upgrade | Announced via tweet, unshipped |
| Grok 4.7 | ~late Aug–early Sep 2026 | 2.1T | Broad upgrade over 4.6, better token efficiency | Announced via tweet, unshipped |
All Grok 4.6/4.7 figures are unverified vendor claims from a single social media post — treat them as directional, not confirmed specs.
03 Why xAI is emphasizing post-training, not just scale
SFT and RL, in plain terms: Supervised fine-tuning (SFT) trains a model on curated example outputs to shape behavior; reinforcement learning (RL) uses reward signals to teach which action sequences actually work — critical for multi-step agentic tasks. Musk's wording — "significantly improved SFT & RL" — signals xAI is doubling down on the same playbook that made Grok 4.5 competitive on agentic benchmarks: roughly 15,954 output tokens per SWE-Bench Pro task versus Opus 4.8's 67,020 — a 4.2x efficiency gap, largely credited to post-training on real Cursor developer sessions.
The scale-vs-speed trade-off: Grok 4.6's jump to 1.5T is a real scale increase, but Musk's framing of Grok 4.7 — bigger at 2.1T, "better in every way except slightly slower to serve" — suggests xAI is building two SKUs with different trade-offs, similar to Anthropic's Sonnet/Opus split or OpenAI's mini/full tiers.
Competitive pressure the timeline doesn't mention: Grok 4.6's target lands almost exactly 10 days after Kimi K3's full open-weight release. Moonshot's model topped the Frontend Code Arena leaderboard at 1,679 points — the first open-weight model to beat every closed model on that board — and ranked third on Artificial Analysis's Intelligence Index. Musk himself commented "impressive" on Kimi K3 benchmark posts. Chinese financial outlets reported the release wiped an estimated $314 billion off combined OpenAI/Anthropic valuation expectations — analyst estimates relayed through media, not independently confirmed, but directional sentiment is clear.
| Model | Vendor | Parameters | Context | Pricing (input/output per 1M tokens) |
|---|---|---|---|---|
| Grok 4.5 | xAI | Undisclosed | 500K | $2 / $6 |
| Grok 4.6 (announced) | xAI | 1.5T | Undisclosed | Undisclosed |
| Kimi K3 | Moonshot AI | 2.8T (MoE, ~16/896 experts active) | 1M | $0.30 (cache hit) / $3 (cache miss) in, $15 out |
| Claude Fable 5.1 (rumored) | Anthropic | Undisclosed | Undisclosed | Rumored unchanged from Fable 5 ($10 / $50) |
| GPT-5.6 Sol | OpenAI | Undisclosed | Undisclosed | Undisclosed |
04 Six steps to prepare before Grok 4.6 launches
- Build a verification checklist: Monitor Musk's X account, the xAI blog, and the model card page. Do not bake "1.5T parameters" into SLAs until an official announcement lands.
- Baseline on Grok 4.5 today: If you have not integrated Grok 4.5 yet, run your real task set on Cursor or the xAI API ($2/$6 per million tokens) and record token usage and pass rates as a comparison baseline.
- Benchmark Kimi K3 in parallel: Kimi K3 already has third-party scores (SWE-bench Verified 93.4%, Frontend Code Arena 1,679). Do not idle your selection window waiting for an unshipped model card.
- Reserve dual-SKU routing: Musk already hinted at 4.6 (faster) vs. 4.7 (stronger, slightly slower). Add latency/quality fallback logic in your agent gateway — see our OpenRouter dual-route integration guide.
- Assess vendor safety risk: Review xAI's recent CSAM lawsuit and Common Sense Media's child-safety ratings. If your product is consumer-facing or handles sensitive data, factor compliance into TCO, not just leaderboard rank.
- Secure stable compute for agent production: Whether you route to Grok 4.6 API or a mixed stack, local Cursor Agent and CI pipelines still need 24/7 stable environments. Evaluate bare-metal vs. virtualized Mac for Hypervisor overhead.
05 August 2026 release bottleneck, hard data, and FAQ
If Musk's timeline holds, Grok 4.6 and Grok 4.7 will land in the same month as rumored Claude Fable 5.1 (per leaks from 36kr and WinCentral, timed to beat OpenAI's anticipated GPT-6) and just weeks after Kimi K3's open-weight shock. For teams evaluating models, that compresses the useful shelf life of any single flagship to weeks — token efficiency and real-world task cost matter more than leaderboard rank alone.
- Grok 4.5 coding benchmarks (shipped): Terminal-Bench 2.1 83.3%, SWE-Bench Pro 64.7%; ~4.2x output-token efficiency vs. Claude Opus 4.8 (~15,954 vs. 67,020 tokens/task).
- Grok 4.6 announced specs: 1.5T parameters, SFT/RL post-training upgrade — unverified by third parties.
- Grok 4.7 announced specs: 2.1T parameters, "better than 4.6 except slightly slower, better token efficiency" — unverified.
- Kimi K3 reference benchmarks (verified): Frontend Code Arena 1,679 points (top open-weight), Artificial Analysis Intelligence Index rank #3 globally.
FAQ
When exactly is Grok 4.6 coming out?
Musk said "around August 7, 2026" in an X post, but xAI has not officially confirmed a date. Treat it as a target that could shift.
What's the difference between Grok 4.6 and Grok 4.7?
Grok 4.6 is a 1.5T-parameter model focused on SFT/RL post-training improvements. Grok 4.7, expected a few weeks later, is a larger 2.1T model that Musk says outperforms 4.6 across the board except for serving speed, where it trades some latency for better token efficiency.
Will Grok 4.6 beat Kimi K3 or Claude Fable 5.1?
Too early to tell. Grok 4.6 has no published benchmarks yet, Kimi K3 already has verified third-party scores, and Claude Fable 5.1 hasn't even been officially confirmed by Anthropic.
How much will Grok 4.6 cost?
Unknown. Grok 4.5 launched at $2 per million input tokens and $6 per million output tokens — a reasonable reference point, but xAI hasn't disclosed Grok 4.6 pricing.
Where will I be able to use Grok 4.6?
Based on Grok 4.5's rollout, expect Grok Build, the xAI API, and the xAI console first, with third-party integrations (like Grok 4.5's day-one Cursor availability) following — but this isn't confirmed for 4.6 yet.
Primary sources below. All Grok 4.6/4.7 details are vendor announcements or industry leaks — verify against official xAI channels before relying on any specific date, spec, or price.
xAI official blog: Introducing Grok 4.5
xAI Grok 4.5 model card (media.x.ai)
Moonshot AI official blog: Kimi K3
The Verge: "Pacing the Frontier" letter coverage
For teams running Cursor Agent, Grok API mixed routing, or iOS CI/CD around the clock, cloud APIs alone do not solve local Metal toolchain stability — virtualized Mac setups carry Hypervisor overhead and compatibility risk, and rapid model churn stresses the underlying environment. For production workloads that need zero-loss native compute, stable iOS CI/CD, and 24/7 AI Agent automation, ZUKCLOUD's bare-metal Mac mini cloud nodes are usually the better fit: dedicated Apple Silicon hardware, no Hypervisor tax, always-on, flexible daily/weekly/monthly billing. See our bare-metal architecture manifesto for the full engineering rationale.