참고Reddit
과학 분야 LLM 벤치마크, 정답 오류 수정하니 성능 점수 대폭 상승
벤치마크 데이터 신뢰성 문제 제기. 모델 평가 시 자체 데이터셋 검증의 중요성을 시사함.
원문 제목 Turns out that many current science-based LLM benchmarks have flaws in their answers. When corrected, the LLM benchmark scores rose significantly.
원문 보기 ↗