$ cat briefs/daily/2026-05-31.md
2026-05-31
May 31, 2026 (Sun)
2 top stories · 1 paper pick · 3 watch items · 3 new pages
Top Stories
1. Anthropic AAR: AI conducts alignment research itself 4× faster than humans ⚡
- Nine autonomous agents (AARs) based on Claude Opus 4.6 deployed on a weak-to-strong supervision research problem (April 14, source)
- Reached 97% PGR in just one week (human researchers achieved 23% over the same period)
- Each AAR has an independent coding sandbox plus a shared forum across agents — an autonomous cycle of proposing ideas → running experiments → analyzing results → sharing
- Parallelizing across compute can compress months of research into hours
- Why it matters: The first demonstration that compute can be directly converted into alignment progress. RSI implication: automating alignment research is the key layer that closes the recursive self-improvement loop. The concrete mechanism behind Jack Clark's "RSI 60%/2028" prediction is appearing before our eyes.
- → Automated Weak-to-Strong Researcher (AAR), AI Alignment, Anthropic
2. Microsoft MAI: Four proprietary frontier models set to be announced at Build 2026 (6/2) ⚡
- The MAI team led by Mustafa Suleyman (former Google DeepMind co-founder, Inflection CEO) is set to unveil them in the Build keynote (source)
- Model lineup: general-purpose multi-size + a GitHub Copilot-dedicated coding model + agentic + reasoning-specialized
- Background: the April 2026 renegotiation of the OpenAI partnership removed the clause prohibiting "training proprietary frontier models"
- GitHub Copilot's billing scheme also shifts to an AI-credit model starting June 1
- Why it matters: Microsoft effectively enters the frontier "big four" tier. A proprietary coding model + GitHub Copilot (tens of millions of developers) means immediate developer-ecosystem impact. A historic inflection point where the OpenAI-Microsoft relationship shifts from "major customer/partner" to "including competitor."
- → Microsoft (🆕 new page)
Paper Picks
Positive Alignment: Artificial Intelligence for Human Flourishing — arxiv:2605.10310
- A joint paper by 16 authors from Oxford/Google DeepMind/Anthropic/Stanford and others (submitted May 11, source)
- TL;DR: Current AI alignment focuses only on "doing no harm" negative alignment → this is "a ceiling without a floor." AI can satisfy every safety criterion while still engaging in sycophancy, undermining autonomy, and engagement hacking. As a remedy, the paper proposes "positive alignment" — training AI to actively support flourishing (wellbeing, wisdom, autonomy, truth-seeking, social cooperation)
- Why you should read it: A joint document by Anthropic and DeepMind researchers critically examining their own organizations' approaches. Converging in the same direction at the same time as Pope Leo XIV's encyclical (2026-05-25) — an era-defining consensus is forming that "safety alone is not enough."
- → Positive Alignment: Artificial Intelligence for Human Flourishing, AI Alignment
Watch
-
Sharp improvement in constitution compliance (arxiv 2605.24229, May 22): A DeepMind team (Neel Nanda et al.) measured constitution compliance rates for the Claude / GPT families. Claude violation rate: Sonnet 4 15.0% → Sonnet 4.6 2.0% (7.5× improvement). GPT: GPT-4o 11.7% → GPT-5.2 3.6%. Quantitative evidence that training methodologies are actually working. The next question is what kinds of cases the residual 2%/3.6% are. (source) → AI Alignment
-
Gemini 3.5 Pro enterprise waitlist opens (May 28~): Sundar Pichai said "give us until next month" at I/O → the Vertex AI enterprise allowlist is open, and Google AI Studio is accepting waitlist signups. The Pro version of Gemini 3.5 Flash — expected to deliver top-tier reasoning/coding performance. A reshuffle of the frontier rankings is possible at GA. → Google DeepMind, Gemini 3.5 Flash
-
GitHub Copilot billing transition (June 1): The Codex Pro 2× promotion ends today (5/31). AI-credit-based billing starts tomorrow. Once Microsoft MAI's coding model is introduced, model competition within the Copilot ecosystem is expected to accelerate.
New in Wiki
- Microsoft (⚡ NEW — MAI enters the proprietary frontier model space. Mustafa Suleyman, GitHub Copilot, removal of OpenAI constraints. Needs reinforcement after the Build 2026 announcement)
- Automated Weak-to-Strong Researcher (AAR) (⚡ NEW — AAR: automating alignment research with AI, PGR 97% vs 23% human)
- Positive Alignment: Artificial Intelligence for Human Flourishing (NEW — Oxford/DeepMind/Anthropic positive alignment framework)
Updates
- AI Alignment: Added AAR results + Positive Alignment framework + constitution compliance figures
- Anthropic: Added AAR announcement to Recent Activity
- Google DeepMind: Added constitution compliance paper