Share this link via
Or copy link
Forecasting and Aligning AI, presented by Jacob Steinhardt, addresses two of the most critical challenges in the future of artificial intelligence: predicting its trajectory and ensuring its alignment with human values. Steinhardt, a leading AI researcher and professor, explores how we can better anticipate the capabilities and impacts of rapidly advancing AI technologies, particularly as they move toward greater autonomy and generality. The seminar discusses methodologies for AI forecasting, including expert elicitation, trend analysis, and scenario planning. It also delves into...
Introduction.Rest of Talk.Reward Hacking: Motivation.Reward Hacking Example.Reward Hacking: Example.Summary of Full Results.Reward Hacking: Summary.Making NLP Models Truthful.Contrastive Representation Clustering.Results on Unified QA.Caveat: True Answers Work Too.Forecasting: Motivation.Forecasting Competition.Forecasting Questions.Summary of Benchmark Forecasts.Results So Far.Forecasting: Lessons Learned.Forecasting Class.
No reviews yet. Be the first to review!
You must be logged in to submit a review.