참고X
최신 모델들의 'Max Reasoning' 모드, 벤치마크 점수용일 뿐인가?
일부 모델의 추론 최적화 모드가 실사용보다 벤치마크 점수 올리기에만 집중되어 있다는 비판. 실무 적용 시 성능 재검증 필요.
원문 제목 @theo: The “max” reasoning effort on most models should be renamed to “benchmaxxed” and then you should ignore it entirely
원문 보기 ↗