참고Reddit
밀집형 모델을 MoE로 변환하는 'ToMoE v2' 연구
성능 손실을 최소화하며 모델 구조를 경량화하는 기법. 로컬 LLM 최적화 시 참고할만한 새로운 방법론.
원문 제목 What do you think about ToMoE v2 paper, converting dense model to MoE model at near lossless accuracy? I feel Qwen3.8-27B-A16B or something along those lines would be amazing, though there are architectural hurdles, as well as need folr training data.
원문 보기 ↗