Live page · Day archive

SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference

Read the original at HF Daily Papers

Summary

Researchers introduce SparseDecoding, a pruning framework that addresses distribution shifts and sparse matrix-vector operations to improve large language model inference efficiency.

Carried by: HF Daily Papers. First seen: .