Abstract: With the prevailing Mixture-of-Experts (MoE) architecture pushing the performance of Large Language Models (LLMs) to new limits, fine-tuning MoE models presents a significant challenge due ...
LLaMA-MoE is a series of open-sourced Mixture-of-Expert (MoE) models based on LLaMA and SlimPajama. We build LLaMA-MoE with the following two steps: ...