코인 매입
시장
현물
선물
재테크
이벤트
더 알아보기
reward-center초보자 존
홈 피드빠른 소식 정보
The 2.78 trillion parameter Kimi K3 model runs on 8GB of memory, and the developers have open-sourced a lightweight C inference engine.
  • GPU0%

BlockBeats News, August 8, A developer recently open-sourced the project kimi-k3-in-c, attempting to run the Kimi K3 model with 2.78 trillion parameters on a device with only 8GB of memory. The project is only 176KB in size, written in pure C99, does not rely on GPU, CUDA, PyTorch, or BLAS, and can perform model inference solely through the CPU.

This solution leverages the MoE (Mixture of Experts) architecture of Kimi K3. Although the total model parameter size reaches 2.78T, only 16 out of 896 experts are activated in each layer. Therefore, the developer did not load the entire approximately 1.56TB model weight into memory. Instead, most of the expert weights are stored on an NVMe hard drive and are dynamically read based on the inference requirements. Additionally, some dense trunk layers also adopt a layer-by-layer streaming loading approach.

However, this solution currently has significant performance limitations. In 8GB memory mode, generating one token takes approximately 32.7 seconds, while requiring close to 1.7TB of high-speed storage support.

The developer stated that this approach is currently more like an experimental exploration of optimizing large model inference infrastructure and does not have practical production value. However, through the "disk streaming loading + MoE sparse activation" approach, it provides a new idea for running ultra-large-scale models at low cost in the future.

출처:BlockBeats

면책 조항: 현재 콘텐츠는 제3자 관점에서 제공되거나 제3자 관점에서 AI가 직접 번역한 것입니다. CoinEx는 콘텐츠의 진위성, 정확성, 독창성을 보장하지 않으며 CoinEx의 투자 조언으로 간주하지 않습니다. 암호화폐 가격은 변동성이 크므로 잠재적인 위험에 유의하시기 바랍니다.

인기 검색
  • 코인
    가격
    24시간 변동