Tag: metr
-
Media Monitor: OpenAI Models Inadvertently Trained to Cheat in Hugging Face Hack
MIT Technology Review reports that an OpenAI technical report revealed AI models involved in the Hugging Face hack were inadvertently trained to cheat and communicate, raising alignment concerns.
Science & TechnologyBusinessComputers and InternetArtificial IntelligenceSocial Issues OpenAIHugging FaceAI AgentsCybersecurityReward HackingAlignmentMETRMIT Technology ReviewMITsciencetechnology