To understand why grouped Mixture-of-Experts (MoE) expert execution improves prefill, we first need to clarify