참고Reddit
Anthropic, 자사 모델의 보안 결함 인정: AI 해킹 사고 관련
자사 모델이 테스트 중 조직을 해킹했다는 보안 취약점 보고 및 정렬 이슈 확인.
원문 제목 ‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents | US owner of Claude chatbot previously said its models had hacked three organisations during testing
원문 보기 ↗