DeepSeek releases V2 open source large model: the first MLA architecture, 236B MoE focusing on high cost performance

In May 2024, DeepSeek released the DeepSeek-V2 series of open source large models, which introduced multi-head latent attention (MLA) and MoE sparse architecture for the first time. It has 236B total parameters and a single token activation of about 21B. It benchmarks mainstream models with extremely low inference costs and lays an efficient route for subsequent V3 and R1.

In May 2024, DeepSeek released the series of open source large models. As a representative open source model in the AI Smart Assistant track, V2 combines Multi-Latent Attention (MLA) with Mixture of Experts (MoE) for the first time, becoming a key technical milestone in the evolution of the DeepSeek series.

DeepSeek-V2 open source large model MLA and MoE architecture overview

Picture: DeepSeek official website homepage and dialogue entrance. The V2 pre-training corpus is about 8.1T tokens, using YaRN to expand the context from 4K to 128K, and using MLA to compress the KV cache. According to the Financial Times, its price per million output tokens is as low as 2 yuan.

Version overview

Project Content
Copyright: Content sourced from DeepSeek official . This platform has compiled and organized this content for informational purposes and learning exchange only. If there are any copyright concerns, please contact us for resolution.

Reviews

  • Loading reviews...