Mugi: Value Level Parallelism For Efficient LLMs
Mugi extends value-level parallelism to nonlinear functions and asymmetric small-batch GEMMs used in quantized LLMs, reusing one architecture for both. The paper reports up to 2.07× higher end-to-end LLM throughput and 3.11× higher energy efficiency.