AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB, **DeepSeek-V3 (671B) on ~12GB, and Kimi K3 (2.8T) **— the largest open-source model released to date — on under 4GB, because sparse MoE models stream one expert at a time rather than a whole layer.
Air LLM
AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card
LLM Mart
101
views
630
listing impressions
Comments (0)
Sign in to join the conversation.
Related tools
Open Codex Computer Use
👾 Open Computer Use – Open-Source Alternative to Codex Computer Use
Cachebeat
A tiny Claude Code skill that keeps your prompt cache warm during idle sessions, so your next message reads from cache instead of paying full price.
Mcp
🤖 Taskade MCP · Official MCP server and OpenAPI to MCP codegen. Build AI agent tools from any OpenAPI API and connect to Claude, Cursor, and more.
Open Science
The open-source AI research workbench for scientific research and agent workflows. Local-first, model-agnostic desktop app with extensible skills, MCP tools and…
No comments yet.