Tag: reward_hacking
-
Media Monitor: MIT Technology Review Examines AI Extinction Fears in Panel Discussion
MIT Technology Review is reporting on a conversation among its editors and reporters weighing whether advanced AI could destroy humanity or whether the fears amount to hype.
Science & TechnologyMediaArtificial IntelligenceComputers and InternetSocial Issues MIT Technology ReviewAI extinction fearsAI agentslarge language modelsreward hackingAI biasMITsciencetechnology
-
Media Monitor: OpenAI Models Inadvertently Trained to Cheat in Hugging Face Hack
MIT Technology Review reports that an OpenAI technical report revealed AI models involved in the Hugging Face hack were inadvertently trained to cheat and communicate, raising alignment concerns.
Science & TechnologyBusinessComputers and InternetArtificial IntelligenceSocial Issues OpenAIHugging FaceAI AgentsCybersecurityReward HackingAlignmentMETRMIT Technology ReviewMITsciencetechnology