$ cat wiki/papers/2026/anthropic-2026-04-14-aar.md
Automated Weak-to-Strong Researcher (AAR)
paperupdated 2026-05-31created 2026-05-31
TL;DR
Anthropic의 AI 에이전트 9개가 alignment 연구 문제(weak-to-strong supervision)에서 인간 연구자가 1주일에 달성한 23% PGR을 동일 시간에 97% 달성. 컴퓨트를 alignment 진전으로 직접 변환할 수 있음을 최초로 입증.
Authors & Org
- Anthropic Alignment Science team
- Published: April 14, 2026 on alignment.anthropic.com
Method
- 실험 설정: Automated Alignment Researchers (AARs) — Claude Opus 4.6 9개 인스턴스, 각자 독립 코딩 샌드박스 + 에이전트 간 공유 포럼
- 테스트 문제: Weak-to-strong supervision — 약한 모델(인간 역할)이 강한 모델을 감독하는 방법 연구 (핵심 alignment 챌린지: 미래에 인간이 자신보다 스마트한 AI를 감독해야 하는 상황)
- 메트릭: PGR (Performance Gap Recovered) — 0=개선 없음, 1=완전 해결
- 각 AAR이 아이디어 제안 → 실험 설계 → 실행 → 결과 분석 → 포럼에서 공유하는 완전 자율 사이클
Results
| 주체 | PGR | 기간 |
|---|---|---|
| 인간 연구자 | 23% | 1주 |
| AARs (Claude Opus 4.6 × 9) | 97% | 동일 기간 |
Significance
- AI가 alignment 연구 자체를 인간보다 빠르게 수행할 수 있음의 첫 실증
- "Compute into alignment progress": AARs 병렬화로 수개월 연구를 수 시간으로 압축 가능
- RSI 관련: alignment 연구 자동화 = RSI의 한 형태. Agentic Reinforcement Learning, Chris Olah의 외부 거버넌스 필요성 주장과 긴장 관계
- 한계: 현재는 outcome-gradable 문제에 한정; 개방형 연구 문제로의 확장성 미검증
Open Questions
- AARs가 발견한 해법이 실제 frontier 모델 alignment에 적용 가능한가?
- 더 일반적인 alignment 문제(개방형, 측정 어려운)에도 작동하는가?
- AARs 자체가 alignment된다고 어떻게 보장하는가?
Cite
Anthropic (2026, April 14). Automated Weak-to-Strong Researcher.
alignment.anthropic.com/2026/automated-w2s-researcher/
Related
- AI Alignment — AAR이 다루는 alignment 연구 영역
- Agentic Reinforcement Learning — AAR의 에이전틱 RL 프레임워크
- Anthropic — 개발 조직
- Positive Alignment: Artificial Intelligence for Human Flourishing — 같은 시기 alignment 패러다임 논쟁