$ cat briefs/daily/2026-08-02.md
2026-08-02
August 2, 2026 (Sun)
Generated by ingest + brief agents · Sources: Tier 1 (HuggingFace Daily, Anthropic, OpenAI, Google DeepMind, Meta AI, Mistral, xAI, alignment.anthropic.com, X feeds, Qwen + Z.ai rotation) + prefetch candidates
_3 top stories · 0 paper picks · 2 new pages · Sunday: lint + weekly synthesis_
Top Stories
1. OpenAI names its next model — and hands over ten proofs a program can check
OpenAI published "Ten advances in mathematics and theoretical computer science", crediting an internal version of Astra — the name it gives its next major model, used publicly for the first time. The problems are stated to have been open at least ten years, several far longer: the first explicit non-sofic group, a disproof of Connes' Rigidity Conjecture, a quantum parallel repetition theorem for general two-player entangled games, Ehrhart's volume conjecture, the first improvement to the general high-dimensional sphere-packing upper bound since 1978, new circuit-complexity lower bounds and three Erdős problems. Total token spend: roughly $2,000 at Sol API prices.
Each result ships with a Lean 4 certificate and a chain-of-thought walkthrough on GitHub, next to a 249-page manuscript. Coverage states humans organized the proofs into papers before the Lean conversion, so the pipeline is not autonomous end to end. OpenAI cites the Leiden declaration (June 2026, IMU-endorsed, signed by Tao, Scholze, Buzzard and Aaronson) and its five named risks.
Why it matters: this moves a capability claim off the leaderboard entirely — an open conjecture has no harness to configure, and a Lean certificate can be re-checked by anyone with a laptop and no access to the lab. The price is the other half of the same fact: the model that produced it is unreleased, so the generation cannot be reproduced by anyone outside OpenAI, which is one of the five risks in the declaration it chose to cite.
→ Astra (new page), AI for Mathematics (new page), OpenAI, Reasoning Models → Source · OpenAI · @SebastienBubeck · the-decoder
2. Yesterday's unverifiable number turned out to be right
On 2026-08-01 this wiki recorded, under Conflicting Reports, an r/LocalLLaMA claim that DeepSeek-V4-Flash-0731 scored 50 on the Artificial Analysis Intelligence Index — logged but explicitly not quoted as a benchmark, because it was a community report of a third-party index with no primary source behind it. Artificial Analysis has now published the figure itself, and it matches exactly: 50, one point behind GPT-5.6 Luna (max, 51) and GLM-5.2 (max, 51), a 10-point jump over the April Flash and 6 points ahead of DeepSeek V4-Pro — the larger sibling whose official release "will follow soon". AA puts it on its Pareto frontier for Intelligence vs Cost per Task, among the top 3 open-weight models, at roughly 60% below Luna's cost per task, helped by a ~98% cache-hit discount.
Why it matters: an open-weight MIT model is now within one index point of two closed frontier tiers at a fraction of the cost per task — but the durable lesson is procedural. Holding the figure for a day cost nothing and the wiki never published an unsourced number; the discipline that looks like pedantry on a quiet day is what makes the number trustworthy when it arrives.
→ DeepSeek V4-Flash, GLM-5.2, GPT-5.6 Sol (and Terra, Luna), DeepSeek → Source · Artificial Analysis · @ArtificialAnlys
3. Someone outside the room built on the new MCP spec within three days
The 2026-07-28 MCP specification removed protocol-level sessions — already recorded here as the largest revision since the protocol shipped. What arrived this week is the first evidence of it landing outside the vendors who wrote it: Simon Willison published "Stateless MCP has recaptured my interest" on 2026-07-31 and released mcp-explorer (a stateless Python CLI for probing an MCP server) and datasette-mcp. His reasons are operational rather than editorial — no server-side state to maintain, no routing a session back to the same backend machine.
Why it matters: a spec's rationale is a claim until someone with no stake in it builds against the thing and reports the same reasons back. One developer and two tools is a small sample, but it is three days rather than the twelve-month deprecation window, and it is the first datapoint of its kind recorded here.
→ MCP — Model Context Protocol, Agents (LLM Agents) → Source · Simon Willison
Paper Picks
None today. HuggingFace Daily Papers could not be resolved for a second consecutive day — huggingface.co returns 403 from this environment and no search result named an August 1–2 top entry. Recorded as a gap rather than guessed at. Two days without a paper pick is a source problem, not a quiet week; it is carried into today's lint as an action item.
Watch
- When Astra actually ships — no price, no date, no API identifier appears in any source read. Until one does, Astra carries
unknownin five of its seven Spec rows, which is the honest state and a countable one (source). - Whether a journal or the IMU says anything about Lean-certified results — the Leiden declaration named the risks in June; the Astra release is the first large test of them. Nothing read states how peer review treats a machine-checked proof (source).
- DeepSeek-V4-Pro's official release — still "will follow soon", and its smaller sibling now scores 6 points above it on AA's index. Whatever Pro ships as has to answer that (source).
New in Wiki
- Astra (new — OpenAI's unreleased next model; five of seven Spec rows are genuine
unknown) - AI for Mathematics (new — concept page; user review requested. It was created because four model pages already touched AI-for-math with no hub between them, and the generation/verification split looks like a long-lived category rather than an August story)
Updates
- OpenAI: the ten results, the Leiden context, the human-editing caveat; Astra added to Models & Products
- Reasoning Models: the May Erdős result is no longer a one-off — State of the Art 2026-07 → 2026-08-02
- DeepSeek V4-Flash: AA's own figures under Benchmarks; the Conflicting Reports entry closed as accurate rather than left standing
- MCP — Model Context Protocol: independent tooling response to the stateless spec — State of the Art 2026-07-29 → 2026-08-02