Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time

Exploring foci of: arXiv (Cornell University) Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time October 2023 • Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Ré, Beidi Chen Large language models (LLMs) with hundreds of billions of parameters have sparked a new wave of exciting AI applications. However, they are computationally expensive at inference time. Sparsity is a natural approach to reduce this cost, but existing methods either require costly retraining, have to forgo LLM's in-context learning ability, or do not yield wall-clock time speedup on modern hardware. We hypothesize that contextual sparsity, which are small, input-dependent sets of attention heads and MLP parameters t… Open Article Page

Computer Science Machine Learning Artificial Intelligence Deep Learning Parallel Computing Computer Security Biology Paleontology Open Article