← BACK TO BLOG
AI NEWS

Astra: The AI Model That Solved 27-Year-Old Math Problems for $2,000 (And Why It Changes Everything)

Ambitious SocietyAugust 202614 min read
Astra solved a 27-year-old math problem for $2,000 — AI breakthrough cover

Here is the thing nobody expected: AI's biggest breakthrough this year wasn't a bigger benchmark score or a flashier language model. It was proving that mathematicians were wrong about a problem they had been stuck on for 27 years.

On August 1, 2026, OpenAI announced that an internal version of its next major model, code-named Astra, solved ten open problems in mathematics and theoretical computer science. Not proposed solutions. Not approximations. Actual, verifiable proofs that can be checked by machine on GitHub.

✅ Here's what's verified: The proofs exist. They're formalized in Lean 4. They can be downloaded right now and verified on a laptop in seconds.

⚠️ Here's what's self-reported: OpenAI says the total token cost was approximately $2,000. The mathematical arguments came from Astra. The peer-review process is still pending.

🚩 Here's what's hype: "AI solved the hardest math problems." No — AI solved ten specific problems that mathematicians had made no progress on for decades. That's different, and more interesting.

But before we get into what happened, here's why this matters: this is the moment AI stopped being a tool for humans and became a research partner that generates knowledge humans can't.

On August 1, 2026, OpenAI's Astra model solved ten open mathematical problems — formalized in machine-checkable Lean 4 proofs anyone can verify.

What Astra Actually Is

Astra isn't shipping yet. You can't use it. OpenAI hasn't even decided whether to release it as GPT-5.7, GPT-6, or under its own name.

✅ Confirmed: It's described as OpenAI's "next major model" and is designed for "long-running workloads" where AI agents collaborate on complex problems over hours or days.

⚠️ Self-reported: It's "substantially" more capable than GPT-5.6 Sol, with particular strengths in long-horizon reasoning and formal verification.

🚩 Speculation: Some reports suggest Astra will be the first model subject to the Trump administration's pre-release AI safety review framework — meaning it might get held up for government evaluation before any public release.

The practical story: OpenAI built a model that can sustain deep reasoning over many steps, formalize arguments in machine-readable proof code, and catch its own mistakes without human intervention. Then they unleashed it on ten of the hardest unsolved problems in mathematics.

What came back wasn't luck. It was proofs.

The Ten Problems — And Why They Matter

Most of these are not household names. They don't need to be. They matter because they have resisted human effort for decades, and resolving them opens new research directions in fields like quantum complexity, group theory, and cryptography.

The headline result: Non-sofic groups exist (27 years open)

In 1999, mathematician Mikhail Gromov introduced a concept called "soficity" — the idea that certain infinite abstract groups can be approximated by finite groups. For 27 years, nobody could answer the central question: do non-sofic groups actually exist?

✅ Confirmed: Astra constructed an explicit non-sofic group. The proof spans property-(T) expanders and the binary Leavitt algebra.

Fields Medalist Tim Gowers's reaction: "If this proof is correct, I would recommend it for publication in a top journal without hesitation."

Other breakthrough results:

  • ✅ Disproof of Connes's rigidity conjecture — a 50-year-old conjecture about von Neumann algebras, dismantled.
  • ✅ First improvement to sphere-packing bounds since 1978 — in high dimensions, how efficiently can you pack non-overlapping spheres? Relevant to cryptography and error-correcting codes.
  • ✅ Three Erdős problems resolved — including problem #183 on Ramsey numbers. The Erdős catalogue is the "greatest hits of unsolved problems."
  • ✅ New lower bounds on circuit complexity — specifically, computing the permanent of a matrix. This affects computational complexity theory and cryptographic assumptions.
  • ✅ Parallel repetition theorem for quantum games — extends a principle from classical complexity theory into the quantum realm.

The pattern: these aren't toy problems. They sit at the foundations of their fields and have resisted serious attempts at resolution for decades.

💡 TIP
Want to start using AI the way the world's top researchers are — but for your everyday work?

Grab our 100 Free AI Prompts — 25 ready-to-use prompts each for ChatGPT, Claude, Gemini, and Grok. Copy. Paste. Get results.

Get The Free Guide
Astra's proofs were formalized in Lean 4 — a proof assistant that converts mathematical arguments into machine-readable code that can be instantly verified by anyone with a laptop.

The Verification Angle — Why Lean Matters

Here is where the story gets interesting.

In traditional peer review, when someone claims a major proof, experts spend months — sometimes years — checking the work. It's slow. It's expensive. It's error-prone.

Astra did something different.

Each proof was formalized in Lean 4, a proof assistant that converts mathematical arguments into machine-readable code. The Lean kernel then verifies every step automatically. If a single logical step is wrong — if one inference doesn't follow from the previous statement — Lean rejects the whole thing.

✅ Confirmed: The repository on GitHub ships 10 Lean 4 certificates. The repository's "sorry count" — lines where a proof is incomplete — stands at zero.

⚠️ Important caveat: None of these results have been through peer review in the traditional sense. Mathematicians have read and validated some. But none have been published in a journal and gone through the full review cycle.

Why this matters: this collapses the timeline between discovery and verification from months to seconds. If you want to claim a mathematical result, ship a Lean certificate. If your argument breaks on Lean, you have a problem. No interpretation required.

The Leiden Declaration and the Timing

In June 2026 — just two months before Astra — the mathematical community published the Leiden Declaration on Artificial Intelligence and Mathematics. ✅ Endorsed by the International Mathematical Union with over 1,000 signatories in the first 24 hours.

A coalition of 16 mathematicians from Oxford, Cambridge, ETH Zurich, Columbia, and Northwestern published a manifesto warning that AI companies are using published research to train models without researcher consent, bypassing peer review by announcing results via press release, threatening proper attribution when proprietary models generate proofs, and creating an unequal playing field where big tech controls computational infrastructure for mathematics.

The declaration called for responsibility: disclose when AI was used, submit work to peer review, get consent from researchers whose work trained the model.

Then, two months later, OpenAI announced Astra by publishing ten results through a press release, a GitHub repository with machine code, and no peer review process — exactly what the declaration warned against.

The timing is deliberate. OpenAI is publicly testing the boundaries the math community just drew.

Their implicit argument: the proofs are machine-verifiable. Anyone with a laptop can verify correctness in seconds. That is better peer review than papers can offer — automated, objective, trustless.

The mathematical community's counter-argument: that is not peer review. That is automated verification. Peer review also checks whether you understood the problem correctly, whether your approach is novel, whether the work opens new research directions. A machine-checkable certificate cannot tell you those things.

The Leiden Declaration tried to slow down AI's encroachment into mathematical research. Astra's announcement may have made that shift irreversible.

Why $2,000 Changes the Equation

Let's be concrete about the cost angle, because it reframes what's possible in mathematical research going forward.

If you wanted to hire a team of world-class mathematicians to work on these problems for a month, you'd spend at least six figures. Add postdoc salaries, faculty time, research assistants, and computational infrastructure. You're looking at $500K to $2M for a serious push on even one of these problems. And there's no guarantee of success.

✅ Astra cost $2,000 in tokens.

That's not a rounding error. That is a 250x to 1,000x cost reduction compared to the human equivalent research at scale.

Think about what that unlocks. Universities with $100K in AI budget can now fund serious mathematical research programs they couldn't afford before. Startup founders can tackle foundational mathematical problems as part of building products. Countries can invest in AI-assisted mathematical research as a strategic advantage.

⚠️ Reality check: this assumes the proofs are correct and will hold up to scrutiny. If peer review finds mathematical errors — or if Astra starts hallucinating proofs that seem right but aren't — the cost isn't $2,000. It's the credibility hit OpenAI takes and the wasted time from researchers pursuing dead ends.

But if the proofs hold up — and early signals from mathematicians like Tim Gowers and Thomas Bloom suggest they will — then AI has fundamentally altered the economics of mathematical research overnight.

The shift: the bottleneck moves from "can we afford elite researchers?" to "can we formalize the problem clearly enough for the model to understand it?" That's a different world than the one we've had for 300 years of mathematical tradition.

What This Means for Science — And Not Just Math

This is simultaneously a genuine breakthrough and a complicated inflection point for how science works.

The breakthrough part: AI can generate mathematical insights that advance human knowledge at scale. The non-sofic group result answers a 27-year-old foundational question. The sphere packing bounds improve on work from 1978 — 44 years of stagnation broken in one model run. These are genuine answers to questions humans have been asking for decades and couldn't resolve.

The complicated part: AI is generating these insights on a timeline and in a format that the traditional scientific apparatus wasn't designed for. The mathematical community suddenly has a new kind of authority facing them — a company with proprietary compute power, unilateral control over the model, and the ability to decide what gets announced and when.

⚠️ The epistemology problem: everything we know about Astra's reasoning comes from OpenAI. No independent lab has access to replicate these results. Science is built on reproducibility. If only OpenAI can run Astra, then only OpenAI can verify these results.

AI's ability to generate and verify mathematical proofs at this scale fundamentally changes who controls the future of scientific discovery.

The Hype Filter

Before we close, here's what's real and what's narrative.

Real: AI generated novel mathematical proofs that resolve decades-old problems. Machine-checkable Lean 4 certificates allow instant verification. This happened.

Real: The cost was dramatically lower than human research would be. The speed was shocking. The breadth — ten problems across eight different fields — is impressive.

Hype: "AI solved unsolvable problems." No. AI solved specific problems that humans hadn't solved yet. Some of these might have been solved eventually with more time and more humans.

Hype: "This proves AI is smarter than humans." No. It proves AI can sustain long chains of reasoning and formalize them correctly. Human mathematicians directed the problem selection. Humans will validate the significance. Humans will figure out what to do with these results next.

Real gap: We don't know what's going on inside Astra's reasoning. The chain-of-thought walkthrough OpenAI published is a summary written by the model itself, then cleaned up by humans.

What Comes Next

Three unknowns:

When does Astra release? OpenAI hasn't given a date. Reports suggest it might be the first model subject to the Trump administration's pre-release AI safety review — which could accelerate or delay the timeline significantly.

How is it released? As GPT-5.7? GPT-6? A separate product? The naming signals how OpenAI positions its next generation.

What do mathematicians do? The question is whether traditional peer review is enforceable now that Lean certificates exist and anyone can verify instantly.

The most likely outcome: a hybrid model emerges. Results get announced with Lean certificates for instant verification. Papers get submitted to journals for peer review of significance and novelty. Both happen in parallel, not sequentially.

Announcement plus verification on day one. Peer review over the next six months in parallel. That is a fundamentally different timeline than the 18-month journal review cycle that has governed mathematics for generations.

💡 TIP
AI is moving fast. The people who understand these tools will have the advantage.

Grab our 100 Free AI Prompts — 25 ready-to-use prompts each for ChatGPT, Claude, Gemini, and Grok. Start using the tools reshaping the world — today. Free. No catch.

Get The Free Guide

And when you are ready to go deeper — the full AI Mastery catalog covers everything from foundational prompting to advanced income strategies.

Explore The Full Catalog →

Ambitious Society exists to make AI education accessible to everyone. No jargon. No gatekeeping. Just real skills that translate into real results. Follow us on Threads @ambitious_society_1972.