The Road to Superalignment: From Motivation to Practical Algorithms

Speaker: HyunJin Kim

Location: 1 MetroTech Center, Room Global AI Frontier Lab (Floor 22)

Date: Monday, August 17, 2026

We are pleased to announce the next session of the Global AI Frontier Lab: Seminar Series on August 17th, 2026. HyunJin Kim will be presenting “The Road to Superalignment: From Motivation to Practical Algorithms". Dinner & networking will begin at 6:00 PM and the seminar will start at 7:00 PM EST. The seminar will be held at the Global AI Frontier Lab at 1 Metrotech Center, Brooklyn, NY 11201. This event will be in-person & online. In-person attendance is strongly encouraged for Lab researchers in NYC. All attendees must RSVP to participate. For online attendees, a Zoom link will be sent out prior to the event. Please reach out to global-ai-frontier-lab@nyu.edu with any questions. We hope to see you there!

 

Abstract: As AI systems become increasingly capable, human supervision may no longer be sufficient to evaluate and align their behavior reliably. This creates a fundamental supervision bottleneck. As a policy improves beyond the capabilities represented in its training data, its evaluator may become increasingly unreliable. In this talk, I will introduce superalignment as the problem of maintaining safe and reliable behavior as AI capabilities approach and potentially surpass human intelligence. I will review major scalable-oversight paradigms and discuss their limitations. I will then present the view that competence and conformity should be improved together through an ongoing, alternating training process, rather than treating alignment as a one-time post-training step. Finally, I will introduce UniPRO, a unified policy-and-reward framework for easy-to-hard generalization. UniPRO uses a shared backbone with separate policy and reward heads, alternating policy and reward updates so that the evaluator can adapt as the policy improves.  

 

Bio: HyunJin Kim is a Ph.D. student at Sungkyunkwan University, advised by Professor JinYeong Bak. His research focuses on language model alignment, scalable oversight, and unsupervised learning. He was previously a research intern at Microsoft Research Asia.