Air LLM

AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card

LLM Mart
101 views 630 listing impressions

AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB, **DeepSeek-V3 (671B) on ~12GB, and Kimi K3 (2.8T) **— the largest open-source model released to date — on under 4GB, because sparse MoE models stream one expert at a time rather than a whole layer.

Comments (0)

Sign in to join the conversation.

No comments yet.

Related tools