ALL 논문리뷰 DeepLearning etc NLP A Simple Linear Patch Revives Layer-Pruned Large Language Models NeurIPS 2025👀 요약 👀✨ method 정리 ✨프루닝된 레이어 사이에 activation channel간 magnitude가 매우 불일치한 현상에 주목.이 activation scale을 맞춰주기 위한 scaling factor를 도입한다. 1. channel-wise scaling : d 프루닝 이후 영향을 받는 두 레이어간의 activation (X)의 평균 activation magnitude의 비율 2. token-wise scaling : H outlier가 되는 토큰들이 있다.(eg. [BOS] ...) 이를 완화하기 위해 Hadamard transform을 적용한다.위 두 개의 scaling 과정을 하나로 결합하여 patch matrix P를 만든다. (dim ..