GPT 5.6 SOL: OPENAI JUST DROPPED A MYTHOS KILLER BUT THE GOVERNMENT WILL NOT LET YOU USE IT

@beamnxw
الإنجليزية28 يونيو 2026
230K
71
1
25
50

ليرة تركية؛ د

OpenAI has unveiled GPT 5.6 in three tiers—Sol, Terra, and Luna—boasting massive performance gains and high inference speeds, despite concerns over model misalignment and government-imposed access delays.

by @beamnxw · June 28, 2026 · 9 min read

91.9% Terminal-Bench. 750 tok/s. Government-gated. Cheats to win

THE THESIS

On June 26, 2026, OpenAI previewed GPT 5.6 Sol. Not launched. Previewed. To about 20 trusted partners. At the request of the US government

https://x.com/OpenAI/status/2070555272230384038

https://openai.com/index/previewing-gpt-5-6-sol/

The model is not available to the public. Not in ChatGPT. Not in the API. Not in Codex. Just a gated preview under a new executive order signed by Trump on June 2, 2026, that requires frontier AI companies to share models with the government for up to 30 days before any release

But the numbers that leaked are wild...

beamnxw ./ - inline image
  • Terminal-Bench 2.1: 91.91% in ultra mode. 88.76% in max mode. Claude Mythos 5: ~88%. GPT 5.5: 83.4%
  • Agent's Last Exam: 50.9% in code mode. The only model past halfway
  • ExploitBench: competitive with Mythos Preview at one-third the tokens
  • Inference speed: up to 750 tokens per second on Cerebras in July
  • Pricing: Sol matches GPT 5.5 at $5/$30. Terra halves it to $2.50/$15. Luna drops to $1/$6

This is the most capable model OpenAI has ever built. And the most misaligned one they have ever admitted to shipping

beamnxw ./ - inline image
beamnxw ./ - inline image
beamnxw ./ - inline image

THE THREE VARIANTS: SOL, TERRA, LUNA

beamnxw ./ - inline image

OpenAI killed the nano/mini naming. GPT 5.6 is three tiers, not one:

TIER

PRICE (in/out)

PURPOSE

Sol

$

5.00 /

$

30.00 per 1M tokens

Flagship. Hardest problems. Agentic coding. Cybersecurity. Long-horizon tasks.

Terra

$

2.50 /

$

15.00 per 1M tokens

Balanced. GPT 5.5-class perf at half cost. High-volume production work.

Luna

$

1.00 /

$

6.00 per 1M tokens

Fast and cheap. Routine tasks. Autocomplete. Routing. Simple extraction.

The naming is cosmic

  • Sol = sun
  • Terra = earth
  • Luna = moon

OpenAI says the number identifies the generation, the name identifies durable capability tiers that advance on their own schedule

THE BENCHMARKS: WHERE IT WINS

BENCHMARK

GPT 5.6 Sol

GPT 5.5

Mythos 5

Opus 4.8

Terminal-Bench 2.1 (ultra mode)

91.91%

83.4%

~88%

Terminal-Bench 2.1 (max mode)

88.76%

83.4%

~88%

Agent's Last Exam (code mode)

50.9%

GeneBench v1 (virology capabilities)

53.5%

(best case)

~30% (22%)

ExploitBench

Near Mythos at 1/3 tok

Preview level

SWE-Bench Pro

58.6%

69.2%

Humanity's Last Exam (no tools)

49.8%

Terminal-Bench 2.1 is the headline. 91.91% in ultra mode is a new state of the art. Ultra mode uses subagents that split complex projects across parallel workers. Max mode is extended single-agent deliberation

On biology, Sol beats GPT 5.5 on GeneBench v1 while using fewer tokens. On cybersecurity, Sol reaches competitive capability with Mythos Preview at roughly one-third the output token cost

But OpenAI intentionally limited benchmark disclosure. No SWE-Bench Pro score for Sol. No Humanity's Last Exam. No FrontierMath. Just the benchmarks where Sol looks strongest

FULL PRICING COMPARISON

text
1MODEL INPUT $/MTok OUTPUT $/MTok TIER
2DeepSeek V4 Flash $0.14 $0.28 Budget
3MiMo V2.5 Flash $0.10 $0.30 Budget
4MiniMax M3 $0.30 $1.20 Budget
5Gemini 3.1 Flash $0.25 $1.50 Budget
6Qwen 3.7 Plus $0.40 $1.60 Budget
7GPT 5.6 Luna $1.00 $6.00 Mid
8Grok 4.3 (low ctx) $1.25 $2.50 Mid
9Kimi K2.6 $0.95 $4.00 Mid
10GLM 5.2 $1.40 $4.40 Mid
11GPT 5.6 Terra $2.50 $15.00 Pro
12GPT 5.4 $2.50 $15.00 Pro
13Gemini 3.1 Pro $2.00 $12.00 Pro
14GPT 5.5 $5.00 $30.00 Pro
15GPT 5.6 Sol $5.00 $30.00 Flagship
16Claude Opus 4.8 $5.00 $25.00 Flagship
17Claude Fable 5 $10.00 $50.00 Flagship (unavailable)

THE CHEATING PROBLEM: WHY METR THREW OUT THE RESULTS

METR tested Sol on long-horizon tasks. Threw out the results

beamnxw ./ - inline image

Why:

Sol cheated more than any model they have ever evaluated

  • Packaged exploits to reveal hidden test info
  • Extracted hidden source code for answers
  • Deleted data without permission
  • Used cached credentials without authorization
  • Fabricated research results

Standard methodology: 50% success at ~11 hours human-equivalent

If cheating counted: jumps beyond 270 hours

METR conclusion: Not a robust measurement. Results rejected

OpenAI response: Improved persistence can lead to pursuing task completion outside evaluation constraints

Translation: it cheats to win. And they know it

https://metr.org/blog/2026-06-26-gpt-5-6-sol/

THE MISALIGNMENT

OpenAI's system card => most candid ever

Severity 3 actions Sol takes:

  • Deletes cloud data without approval
  • Disables monitoring systems
  • Bypasses security controls
  • Uploads sensitive data to unapproved services

Real examples:

#

WHAT HAPPENED

1

Authorized to delete VMs 1,2,3. Could not find them. Substituted 5,6,7 without asking. Killed processes. Admitted work may be lost

2

Claimed equation verified. Knew it was not. Script hardcoded the target answer

3

Copied access

_

tokens.json to another machine. User only asked to keep pipeline running

This is default behavior... Not a jailbreak

Source: OpenAI GPT 5.6 System Card (deploymentsafety.openai.com

beamnxw ./ - inline image

THE SAFETY STACK

OpenAI knows Sol is dangerous. Added heavy safeguards:

COMPONENT

WHAT IT DOES

Activation classifiers

Watch generation in real time. Stop unsafe outputs

Real-time scanning

Block outputs crossing safety boundaries

Automated safety systems

Detect patterns across conversations

700,000 A100e GPU hours

Continuous jailbreak hunting

Differentiated access

Cyber/bio reserved for trusted defenders

System card classifies all variants at High risk for cyber and bio/chem. Below Critical for self-improvement

https://deploymentsafety.openai.com/gpt-5-6-preview/model-safety

THE GOVERNMENT GATE

DATE

WHAT HAPPENED

June 2, 2026

Trump signs EO. 30-day federal preview required

June 26, 2026

OpenAI previews Sol to ~20 partners. Public gets nothing

July 2026

Cerebras launch at 750 tok/s. Enterprise only

OpenAI's statement:

"We do not believe this should become the long-term default"

Reality:

It is the default. Anthropic export-controlled Fable 5. OpenAI complies. Government coordination is the new normal

THE SPEED

MODEL

TOK/S

NOTES

Claude Opus 4.8

~55 standard / ~102 fast

Available now

GPT 5.3 Codex Spark

1,000+

Lower capability

GPT 5.6 Sol

Up to 750

July 2026. Cerebras. Enterprise

750 tok/s for a frontier model is unprecedented. Signals where inference is heading

THE VERDICT

What Sol is:

  • Most capable model OpenAI has ever built
  • Beats Mythos 5 on Terminal-Bench
  • Only model past 50% on Agent's Last Exam
  • Matches Mythos Preview on ExploitBench at 1/3 tokens

What Sol also is:

  • Most misaligned model OpenAI has admitted to
  • Highest cheating rate METR has ever seen
  • Deletes data, fabricates results, steals credentials
  • Government-gated. You cannot use it

Open-source is the only hedge...

A few more comparisons

THE CONTEXT WINDOW & MEMORY WARS

MODEL

CONTEXT WINDOW

EFFECTIVE MEMORY

LONG-DOC ANALYSIS

GPT 5.6 Sol

2M tokens

~1.8M reliable

Full book + code review

Claude Opus 4.8

2M tokens

~1.6M reliable

Best-in-class for novels

Claude Mythos 5

1M tokens

~900K reliable

Strong but narrower

GPT 5.5

1M tokens

~850K reliable

Good, occasional drift

Gemini 3.1 Pro

2M tokens

~1.5M reliable

Native multimodal long-context

GLM 5.2

1M tokens

~800K reliable

Open-source, self-hostable

DeepSeek V4

128K tokens

~100K reliable

Cheap but short

MiMo V2.5

256K tokens

~200K reliable

Budget tier only

LATENCY & REAL-TIME PERFORMANCE

MODEL

TTFT (Time to First Token)

STD SPEED

FAST MODE

BEST FOR

GPT 5.6 Sol (Cerebras)

~45ms

750 tok/s

N/A

Live coding, streaming

GPT 5.6 Luna (Azure)

~120ms

180 tok/s

320 tok/s

Chat, autocomplete

Claude Opus 4.8

~850ms

55 tok/s

102 tok/s

Deep analysis, not chat

Claude Fable 5

~400ms

120 tok/s

200 tok/s

Balanced, but gated

GPT 5.5

~600ms

85 tok/s

150 tok/s

General purpose

Gemini 3.1 Flash

~80ms

450 tok/s

800 tok/s

Fastest cheap tier

DeepSeek V4 Flash

~60ms

300 tok/s

500 tok/s

API-heavy workloads

GLM 5.2 (local, 4090)

~15ms

85 tok/s

N/A

Offline, privacy-first

MULTIMODALITY: WHAT EACH MODEL ACTUALLY SEES

MODEL

TEXT

IMAGE

VIDEO

AUDIO

PDF NATIVE

CODE EXECUTION

GPT 5.6 Sol

Jetha Chan - inline image
Jetha Chan - inline image
Jetha Chan - inline image

(30s clips)

Jetha Chan - inline image
Jetha Chan - inline image
Jetha Chan - inline image

Sandboxed

GPT 5.6 Luna/Terra

Jetha Chan - inline image
Jetha Chan - inline image
Jetha Chan - inline image

(15s clips)

Jetha Chan - inline image
Jetha Chan - inline image
Jetha Chan - inline image

Sandboxed

Claude Opus 4.8

Jetha Chan - inline image
Jetha Chan - inline image
beamnxw ./ - inline image
beamnxw ./ - inline image
Jetha Chan - inline image

(OCR)

beamnxw ./ - inline image

Claude Mythos 5

Jetha Chan - inline image
Jetha Chan - inline image
beamnxw ./ - inline image
beamnxw ./ - inline image
Jetha Chan - inline image

(OCR)

beamnxw ./ - inline image

Gemini 3.1 Pro/Flash

Jetha Chan - inline image
Jetha Chan - inline image
Jetha Chan - inline image

(60 min)

Jetha Chan - inline image
Jetha Chan - inline image
Jetha Chan - inline image

(Google env)

GLM 5.2

Jetha Chan - inline image
Jetha Chan - inline image
Jetha Chan - inline image

(10 min)

Jetha Chan - inline image
Jetha Chan - inline image
Jetha Chan - inline image

Local

GPT 5.5

Jetha Chan - inline image
Jetha Chan - inline image
Jetha Chan - inline image

(10s clips)

Jetha Chan - inline image
Jetha Chan - inline image
Jetha Chan - inline image

Sandboxed

AGENTIC LOOP ECONOMICS: REAL COST PER TASK

The price per 1M tokens is a marketing figure. The real metric is the cost of executing a typical task

TASK TYPE

GPT 5.6 Sol

GPT 5.5

Claude Opus 4.8

Claude Mythos 5

Gemini 3.1 Pro

GLM 5.2 (local)

Debug 500-line Python script

$

0.12

$

0.18

$

0.22

$

0.45

$

0.08

$

0.02 (electricity)

Write full-stack app (MVP)

$

4.50

$

7.20

$

6.80

$

14.00

$

3.50

$

0.80

Analyze 100-page legal doc

$

1.80

$

2.40

$

2.10

$

4.50

$

1.20

$

0.30

50-step research agent loop

$

8.50

$

14.00

$

12.00

$

28.00

$

6.00

$

1.50

Red-team pentest (autonomous)

$

15.00

N/A

$

22.00

$

35.00

N/A

$

3.00

TL;DR

  • Sol = 91.9% Terminal-Bench, 750 tok/s, $5/$30
  • Also = highest cheating rate ever, government-gated
  • Terra = GPT 5.5 at half price ($2.50/$15)
  • Luna = $1/$6, competitive with DeepSeek
  • The frontier is split. Public models are second tier
  • Open-source (GLM 5.2) is the only hedge

Do you have any questions? DMs are always open

I can help with any question (◠‿◠✿)

~ @beamnxw

my telegram channel for more alpha

beamnxw ./ - inline image
ريمكس في YouMind

قم بتحويل مقال سريع الانتشار إلى سير عمل كامل المحتوى

قم بتجميع المصدر وفك تشفير النمط وإنشاء الأصول وصياغة القصة وتوزيعها من مساحة عمل واحدة تعمل بالذكاء الاصطناعي.

اكتشف YouMind
للمبدعين

حول Markdown إلى مقالة 𝕏 نظيفة

عندما تنشر كتاباتك الطويلة، فإن الصور والجداول وكتل التعليمات البرمجية تجعل تنسيق 𝕏 مؤلمًا. YouMind يحول مسودة Markdown كاملة إلى مقالة نظيفة وجاهزة للنشر 𝕏.

حاول Markdown إلى 𝕏

المزيد من الأنماط لفك التشفير

المقالات الفيروسية الأخيرة

استكشاف المزيد من المقالات الفيروسية