Plugin4Shell: The RCE Bug Hitting 4 AI Coding Agents
xAI shipped Grok 4.7 on September 21, with no waitlist and no staged rollout, five months after the company first started talking about it and five walked-back timelines after that. The headline spec is a 2.1 trillion-parameter model, up 40% from Grok 4.6’s 1.5 trillion, trained in part on SpaceX and Starlink engineering data. The other headline, the one xAI put in its own launch copy, is that Grok 4.7 ships with the company’s “strongest safety guardrails yet.”
That claim landed nine days after Dario Amodei published his pacing essay, arguing labs should trade capability gains for safety headroom, and four days after King Charles convened Nvidia, Google DeepMind, and Anthropic in Scotland to make roughly the same case. Musk skipped that summit. xAI shipped a bigger, faster model into the same week anyway, wrapped in the language of caution. Worth checking whether the wrapping matches what’s inside.
Quick Summary: What Happened
Detail Info Shipped September 21, 2026, no waitlist Parameters 2.1 trillion, up 40% from Grok 4.6’s 1.5 trillion Context window 500,000 tokens (unchanged from Grok 4.6) Pricing $2 per million input tokens / $6 per million output tokens (flat vs. Grok 4.6) Training data Supplemental SpaceX/Starlink telemetry, manufacturing records, and engineering failure logs Available in Grok app, Cursor, Grok Build, the xAI API, GitHub Copilot Delays since late July At least five separate walk-backs of the release date Bottom line: Grok 4.7 is a real capability jump at unchanged pricing, shipped on the same week the rest of the industry was publicly debating whether to slow down. xAI’s safety framing deserves more scrutiny than a launch-day blog post is going to get on its own.
xAI’s own numbers are straightforward. Grok 4.7 runs on 2.1 trillion parameters, a 40% jump from the 1.5 trillion in Grok 4.6, which shipped August 12. The context window holds at 500,000 tokens — still the smallest ceiling among the major frontier models, a gap we flagged in our Grok 4.6 comparison against GPT-5.6 Sol’s 1.05 million and Claude Fable 5 Max’s 1 million. Pricing didn’t move at all: $2 per million input tokens, $6 per million output, the exact rate Grok 4.6 launched at.
Availability is the part that actually surprised people. No waitlist, no staged enterprise rollout — Grok 4.7 went live simultaneously in the consumer Grok app, the xAI API, Grok Build (xAI’s CLI), and GitHub Copilot, plus Cursor, which SpaceX agreed to acquire in June and closed in August for $60 billion. That last one isn’t a coincidence worth glossing over. Cursor and xAI now share a parent company, and a same-day launch across both is exactly the kind of distribution advantage that acquisition was supposed to buy.
The more interesting fact is buried lower in xAI’s announcement: supplemental training data pulled from SpaceX and Starlink operations. Satellite telemetry, manufacturing records, engineering failure logs — the kind of messy, real-world physical-systems data that doesn’t show up in a typical web-scale pretraining corpus. xAI’s stated goal is better reasoning about physical systems: engineering tolerances, failure modes, the sort of problem where a model needs to reason about how hardware actually breaks rather than how a textbook describes it breaking. No other frontier lab has a sister company that manufactures rockets and operates a satellite constellation. That data source is genuinely unique to xAI, whatever else you think about how it’s being spent.
It shipped ten days after that last “few more days,” on September 21. Every estimate in that chain expired right around the moment the next one replaced it, which is its own kind of pattern: not one slip, but a rolling one that reset every time the previous deadline came due.
“Strongest safety guardrails yet” is a claim xAI gets to make without anyone independently checking it on launch day. That’s true of basically every model release from every lab — third-party red-teaming takes weeks, not hours. But the claim doesn’t exist in a vacuum this particular week. It landed into a live, public argument about whether frontier labs are moving too fast to build guardrails properly at all.
Amodei’s essay wasn’t subtle: he argued labs should deliberately trade capability gains for more safety headroom. Nine days later, xAI’s own launch post put “strongest safety guardrails yet” right above a spec sheet showing 2.1 trillion parameters, a 40% jump over Grok 4.6, shipped into the same week the industry was publicly debating restraint. Agreeing that the industry should slow down and then shipping a bigger model faster than your last release cycle are not automatically contradictory — models can get both bigger and better-guardrailed at once. But xAI’s own copy makes the tension easy to see: a safety superlative sitting one paragraph above a capability jump, with nothing in between explaining how the two square up.
The SpaceX data angle adds a second, quieter thread. Training on satellite telemetry and manufacturing failure logs is a genuinely different data source than anything OpenAI, Anthropic, or Google DeepMind has access to, and it’s aimed at a genuinely useful capability: models that reason better about physical systems instead of just language about physical systems. That’s also exactly the kind of capability King Charles’s Dumfries House summit was implicitly worried about — models getting better at reasoning about the physical world, not just text about it. Nobody at that summit was talking about Grok specifically. The timing means Grok 4.7 became the first concrete example of the thing the summit was abstractly warning about, arriving four days after the warning.
If you’re already building on Grok 4.6, the upgrade path to 4.7 is essentially free — same pricing, same context window, same API surface, more capability. There’s no reason to hold off unless your workload specifically depends on 4.6’s exact benchmark behavior.
If you’re evaluating frontier models on cost, the math from our Grok 4.6 comparison still applies almost unchanged: $2/$6 per million tokens remains the cheapest frontier-adjacent pricing on the market, and it just got attached to a materially bigger model. Read your own context-length logs before assuming the discount holds at scale — xAI’s pricing has historically stepped up past a certain context threshold, and nothing in the Grok 4.7 announcement says that structure changed.
If you’re weighing xAI’s safety claims specifically, treat “strongest guardrails yet” as a marketing line until independent red-teaming or a published model card backs it up, not as a verified fact. That’s true of every lab’s launch-day safety language, not just xAI’s — but xAI shipped this one directly into a week where the entire industry was publicly arguing about exactly this question.
If your stack already runs through Cursor, this is the first real test of what SpaceX’s ownership of the editor actually means in practice. Same-day Grok 4.7 availability inside Cursor, on the same day it shipped everywhere else, is the kind of integration speed that’s only possible with common ownership. Watch whether it becomes the default model inside Cursor over the coming weeks — that’s the signal our acquisition coverage said to watch for.
Five delayed timelines from late July to September 21 is not, on its own, remarkable. Every frontier lab slips dates. What makes this stretch worth tracking is what shipped underneath the slippage: a model 40% larger than its predecessor, trained on data no competitor can access, released the same week two of the industry’s loudest voices were arguing publicly for less of exactly this kind of pace.
The SpaceX and Starlink training data is the more durable story here, longer than any single launch date. If it genuinely improves physical-systems reasoning, that’s a capability edge specific to xAI that compounds over future releases, not just a Grok 4.7 feature. Robotics, manufacturing, aerospace, anywhere reasoning about hardware failure matters, is a real market, and Musk owns the rocket company and the satellite constellation generating that training signal for free. No other lab has that pipeline. Whether xAI actually converts it into a durable product advantage, or whether it turns out to be a marketing footnote, is the question worth watching over the next few releases, not this one.
We think the safety framing is the weakest part of this launch, not because it’s necessarily false, but because it’s unverifiable on day one and xAI knows that. “Strongest guardrails yet” costs nothing to say and can’t be checked against anything until independent researchers get access. Landing that language nine days after Amodei’s essay and four days after a king convened the industry to discuss exactly this question reads less like a coincidence and more like xAI positioning itself on the right side of a debate it just shipped directly against.
The SpaceX data story is the part we’d actually pay attention to. It’s specific, it’s verifiable in the sense that the source of the data isn’t in dispute, and it’s a genuine structural advantage that has nothing to do with launch-day marketing copy. A model that reasons better about how hardware actually fails is useful independent of whatever xAI says about safety, and it’s the kind of edge that gets more interesting with each release rather than less.
On the delays: five walked-back estimates in under two months is a pattern, not bad luck, and it’s worth remembering the next time Musk puts a number on a release date. “10 days” on September 1 turned into 21. Treat future timelines from xAI the way you’d treat any vendor’s roadmap after watching it slip five times in a row — as a floor, not a forecast.
Grok 4.7 shipped September 21, 2026, with no waitlist, simultaneously across the Grok app, the xAI API, Grok Build, GitHub Copilot, and Cursor.
Grok 4.7 runs on 2.1 trillion parameters, a 40% increase over Grok 4.6’s 1.5 trillion. The context window holds steady at 500,000 tokens and pricing is unchanged at $2 per million input tokens and $6 per million output tokens. The main new ingredient is supplemental training data from SpaceX and Starlink, including satellite telemetry, manufacturing records, and engineering failure logs, aimed at improving physical-systems reasoning.
$2 per million input tokens and $6 per million output tokens, identical to Grok 4.6’s launch pricing. xAI did not raise prices despite the 40% increase in parameter count.
Musk pushed the release timeline back at least five times starting in late July 2026: from “four weeks out,” to “a few weeks,” to “3 to 4 weeks,” to “10 days” on September 1, to “needs a few more days to cook” on September 11. The model shipped ten days after that final estimate, on September 21.
xAI says it trained Grok 4.7 on supplemental data from SpaceX and Starlink operations, including satellite telemetry, manufacturing records, and engineering failure logs, aimed at improving the model’s reasoning about physical systems and hardware failure modes. No competing frontier lab has access to a comparable data source.
Yes, immediately at launch. SpaceX agreed to acquire Cursor’s parent company Anysphere for $60 billion in June 2026 and closed the deal in August, and Grok 4.7’s same-day availability inside Cursor is the first concrete sign of what that shared ownership means for product integration speed.
Not independently, as of launch. It’s xAI’s own characterization in its announcement, made the same week Dario Amodei’s essay and King Charles’s AI summit put the industry’s pace of development under public scrutiny. Treat it as a marketing claim until third-party red-teaming or a published model card backs it up.
Last updated: September 22, 2026. Sources: xAI — Grok 4.7 announcement · Dario Amodei — We Must Pace the Frontier · The Royal Family — The King convenes tech leaders for AI Summit in Scotland.
Related reading: Grok 4.6 Is 80% Cheaper. Here’s the Catch. · SpaceX Buys Cursor for $60B: What Devs Must Know · Dario Amodei Told AI Labs to Slow Down. OpenAI Blinked. · King Charles Summons AI CEOs After Trump’s Pushback · Grok 4.20 Review: xAI’s 4-Agent System Tested