Свежие переводы
Свежие посты с LessWrong.com
- How Would We Know? Reflections on trying to make AI go well amid deep uncertainty
- Self-Modeling Interventions Modulate Emergent Misalignment
- Current AIs out-persuade professionals in lab settings but (probably) not in the world
- Anthropic's updated Certificate of Incorporation
- Research Note: Filtering Subversion-Relevant Information From Pretraining Data Is Feasible


