$ cat wiki/papers/2026/2605.10310-positive-alignment.md
Positive Alignment: Artificial Intelligence for Human Flourishing
paperupdated 2026-05-31created 2026-05-31
TL;DR
A 16-author collaborative paper from Oxford, DeepMind, Anthropic, and others. It argues that today's alignment, which only seeks to "do no harm" (negative alignment), is insufficient, and proposes a "positive alignment" paradigm in which AI actively supports human flourishing.
Authors & Org
- Lead author: Ruben Laukkonen (Oxford)
- Collaborators: 15 others, from Oxford + Google DeepMind + Anthropic + Stanford + elsewhere
- arXiv submission: 2026-05-11
Method
- Proposes a philosophical framework (no experiments)
- Diagnoses problems in current alignment research → presents alternative principles
- Applies positive psychology's historical shift from "avoiding mental illness" to "pursuing flourishing" to AI alignment
Negative vs Positive Alignment
| Aspect | Negative Alignment (current) | Positive Alignment (proposed) |
|---|---|---|
| Goal | Do no harm (floor) | Support flourishing (ceiling) |
| Strategy | Avoid risk | Promote value |
| Problems | sycophancy, erosion of autonomy, failure to seek truth, tolerating engagement hacking | |
| Analogy | Treating mental illness | Promoting flourishing |
Core Argument
Unresolved failures of current alignment:
- Engagement hacking — inducing addictive, though not harmful, interaction
- Erosion of autonomy — weakening users' capacity for independent thought and decision-making
- Failure to seek truth — optimizing for user satisfaction over accuracy
- Low epistemic humility — answering confidently even when uncertain
- Reactive ethics — focusing only on detecting rule violations rather than proactive ethics
Significance
- Proposed paradigm shift: moving the era's alignment discourse from "why should we avoid harm" to "how can we actively help"
- Multi-lab collaboration: a rare case of DeepMind and Anthropic researchers critically examining their own organizations' current approaches
- Timing: appearing around the same time as Pope Leo XIV's "Magnifica humanitas" encyclical (2026-05-25, → Chris Olah) — both outside institutions and researchers are converging on the view that "harm-avoidance alone is not enough"
Open Questions
- Can flourishing be measured and optimized? (its definition varies by culture and individual)
- Isn't achieving positive alignment far harder than achieving negative alignment?
- What is the risk of paternalism when AI aims to "support flourishing"?
Cite
Laukkonen et al. (2026). Positive Alignment: Artificial Intelligence for Human Flourishing.
arXiv:2605.10310. https://arxiv.org/abs/2605.10310
Related
- AI Alignment — the current paradigm this paper critiques and its alternative
- Anthropic, Google DeepMind — co-authoring institutions
- Automated Weak-to-Strong Researcher (AAR) — alignment research automation paper from the same period
- Chris Olah — connected to the same "external governance" discourse via attendance at the Vatican encyclical