Llama-4-Scout-17B-16E-Instruct is a multimodal, instruction-tuned language model developed by Meta, designed to handle long-context tasks efficiently. It features 17 billion active parameters with a total of 109 billion parameters across 16 experts, utilizing a mixture-of-experts (MoE) architecture. The model supports a context window of up to 10 million tokens, making it suitable for applications requiring extensive context, such as multi-document summarization and large codebase exploration.
meta-llama/llama-4-scout-17b-16e-instructFeb 12, 20251.9s latency