Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RELEASE · MODELS · #751

SpaceXAI releases Grok 4.7 — faster, cheaper model for coding and knowledge work

SpaceXAI released Grok 4.7, a new base-model upgrade positioned as its most capable model for coding and knowledge work; it reportedly uses a larger base model, longer RL training focused on long-running tasks, and an entirely new safeguard stack. The company says Grok 4.7 matches the price and speed of Grok 4.6 while improving verification, longer-context handling, document/presentation generation, benchmark performance (CursorBench 4.0, GDPval, AA Briefcase), and safety metrics (LatchBio biosafety 62.4%, HackerBench v0.3 allowing 3.3% risky prompts); it is available in Cursor, Grok Build, the Grok API, third-party harnesses and cloud platforms, priced from $2 per million input tokens and $6 per million output tokens, with an optional faster (2x) variant at double the price.

KEY POINTS

  1. SpaceXAI released Grok 4.7, a new base-model upgrade positioned as its most capable model for coding and knowledge work; it reportedly uses a larger base model, longer RL training focused on long-running tasks, and an entirely new safeguard stack.
  2. The company says Grok 4.7 matches the price and speed of Grok 4.6 while improving verification, longer-context handling, document/presentation generation, benchmark performance (CursorBench 4.0, GDPval, AA Briefcase), and safety metrics (LatchBio biosafety 62.4%, HackerBench v0.3 allowing 3.3% risky prompts); it is available in Cursor, Grok Build, the Grok API, third-party harnesses and cloud platforms, priced from $2 per million input tokens and $6 per million output tokens, with an optional faster (2x) variant at double the price.
  3. This matters because Grok 4.7 claims frontier price‑performance for longer-running coding and knowledge tasks while introducing a new safeguard stack and partner red‑team access for cybersecurity defense.
MERIDIAN INTELLIGENCE

DECISION BRIEF

60/100
CONFIDENCE
01

WHAT CHANGED

SpaceXAI released Grok 4.7, a new base-model upgrade for coding and knowledge work that uses a larger base model, a longer reinforcement-learning training run weighted toward long-running tasks, and an entirely new safeguard stack; it is being offered in Cursor, Grok Build, via the Grok API, third-party harnesses and cloud platforms with pricing from $2 per million input tokens and $6 per million output tokens (optional 2x-speed variant at 2x price).

02

WHY NOW

The release matters now because Grok 4.7 is presented as a price‑competitive option for longer-running coding and knowledge tasks while introducing a redesigned safeguard stack and invite‑only red‑team access for select cybersecurity partners — claims that could influence developer tooling, benchmark rankings, and defensive research partnerships.

03

WHO IS AFFECTED

Directly affected actors: developers and knowledge‑workers using coding/document generation tools; customers and integrators of Cursor, Grok Build, Grok API, and third‑party harnesses/cloud platforms; and select cybersecurity partners granted invite‑only red‑team access.

04

CONFIRMED

Explicit, source-backed facts: - SpaceXAI announced Grok 4.7 as a new model for coding and knowledge work. (885, 893) - Grok 4.7 uses a larger base model than Grok 4.6 and was trained with a longer RL run focused toward longer tasks. (885, 893) - SpaceXAI says Grok 4.7 improves self‑verification, longer‑context handling, and document/presentation generation, and was trained to understand the Grok Bot harness. (885) - SpaceXAI reports improved benchmark outcomes on CursorBench 4.0, GDPval, and AA Briefcase relative to Grok 4.6. (885) - SpaceXAI reports safety benchmark results: LatchBio biosafety 62.4% and HackerBench v0.3 allowing 3.3% of risky dual‑use cyber prompts. (885) - SpaceXAI states Grok 4.7 is available in Cursor, Grok Build, via the Grok API, third‑party harnesses, model routers and cloud platforms, priced from $2/1M input tokens and $6/1M output tokens; a fast variant is offered at 2x speed and 2x price. (885) - The Decoder reports independent benchmark evaluations where Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index (v4.3.2) and performs mid‑pack versus competitors; Terminal‑Bench 4.0 agentic coding score reported at 26%. (893)

05

UNCERTAIN

Missing or conflicting evidence (explicit): - xAI's claims of 'frontier price‑performance' on CursorBench (885) conflict with The Decoder's reporting that Grok 4.7 scores mid‑pack on a composite independent index and lags leading models on several benchmarks (893). Both statements are explicit in sources but disagree on comparative standing. - Independent verification of xAI's reported safety/benchmarks (LatchBio 62.4%, HackerBench 3.3%) beyond SpaceXAI's announcement is not provided in the supplied excerpts. (885) - Details of the 'larger base model' architecture, the exact training data mix, and specifics of the 'entirely new safeguard stack' (implementation, scope, and failure modes) are not disclosed in the excerpts. (885) - Real‑world performance across diverse agentic coding workflows, long‑duration tasks outside cited benchmarks, and adoption/usage metrics are not reported in the supplied excerpts. (885, 893) - The practical impact and outputs from the invite‑only red‑team access for cybersecurity partners are not described. (885)

06

WHAT TO WATCH

Concrete observable signals to monitor next (and where they would appear): - Independent benchmark releases and comparative evaluations (CursorBench 4.0 follow‑ups, Artificial Analysis Intelligence Index updates, Terminal‑Bench, GDPval, AA Briefcase) for corroboration or contradiction of performance claims. (893) - Third‑party safety and red‑team reports or audits that validate or challenge the LatchBio/HackerBench results and the new safeguard stack. (885) - Public technical disclosures or papers from SpaceXAI with architecture/training details or from independent researchers reproducing/confirming model behavior. (885) - Announcements of wider platform adoption or pricing changes from cloud providers, third‑party harnesses, and enterprise customers. (885, 893) - Published outputs or research from cybersecurity partners given invite‑only red‑team access that show defensive utility or limitations. (885)

WHY IT MATTERS

This matters because Grok 4.7 claims frontier price‑performance for longer-running coding and knowledge tasks while introducing a new safeguard stack and partner red‑team access for cybersecurity defense.

EVIDENCE MAP

4

Editorial claims linked to specific sources, with support, contradiction and context shown separately.

xAI says Grok 4.7 improves self‑verification, handling of longer context, and generation of documents and presentations, with better results on CursorBench 4.0, GDPval, and AA Briefcase compared to Grok 4.6.

SUPPORTED

Checked 2026-09-24 · 1 supporting

SOURCES & TIMELINE

2