{"id":192,"date":"2026-10-08T18:42:58","date_gmt":"2026-10-08T14:42:58","guid":{"rendered":"https:\/\/demensdeum.com\/blog\/2026\/10\/08\/airllm-experiments\/"},"modified":"2026-10-08T19:11:07","modified_gmt":"2026-10-08T15:11:07","slug":"airllm-experiments","status":"publish","type":"post","link":"https:\/\/demensdeum.com\/blog\/2026\/10\/08\/airllm-experiments\/","title":{"rendered":"AirLLM: large language models on a weak graphics card"},"content":{"rendered":"<p>Typically, running a large language model with 70 billion parameters requires several tens of gigabytes of video memory. AirLLM approaches this problem more calmly and offers a different way: such a model can be run on a video card with only 4 GB of memory, and without quantization, distillation and trimming. In this note, I will try to tell you without haste how AirLLM works, how to use it, and I will share my small project AirLLM-experiments, which makes working with this library a little more convenient.<\/p>\n<h2>What is AirLLM<\/h2>\n<p>AirLLM is an open source library for large language model inference written by Gavin Li. Its idea is simple and yet elegant: instead of storing all the model weights in video memory at once, the library loads only one layer at a time onto the GPU. The weights are stored on disk in the form of shards cut into layers, and during generation, each layer is loaded, processed, and unloaded in turn.<\/p>\n<p>It follows that the amount of video memory required does not depend on the overall size of the model, but on the size of one of its layers. That&#8217;s why the 70B model fits in 4 GB, the DeepSeek-V3 on 671B fits in about 12 GB, and the Qwen3-235B fits in about 3 GB. But the size of the model itself is limited primarily by disk space, and not by the capabilities of the video card.<\/p>\n<p>These memory savings come at the cost of speed: each layer is read from disk when each token is generated, so output is slow, about seconds per token. So AirLLM is less about a quick interactive chat and more about the ability to even launch a model that otherwise wouldn\u2019t fit into the hardware.<\/p>\n<h2>Installation<\/h2>\n<p>The library is installed from PyPI using the usual command:<\/p>\n<div class=\"hcb_wrap\">\n<pre class=\"prism undefined-numbers lang-unknown\" data-lang=\"unknown\"><code>pip install airllm\n<\/code><\/pre>\n<\/div>\n<p>When you first launch, the model will be downloaded from Hugging Face and will be decomposed into layers. This process is slow and takes up a lot of disk space, so you should prepare free space in advance. If desired, you can enable 4bit or 8bit compression &#8211; then block quantization will be applied to the weights, which speeds up loading layers from disk by about three times, and the accuracy is lost quite a bit.<\/p>\n<h2>Basic example<\/h2>\n<p>You can work with AirLLM in almost the same way as with a regular model from transformers. It is enough to enter the model ID once on Hugging Face, and then everything will be familiar:<\/p>\n<div class=\"hcb_wrap\">\n<pre class=\"prism undefined-numbers lang-unknown\" data-lang=\"unknown\"><code>from airllm import AutoModel\n\nMAX_LENGTH = 128\nmodel = AutoModel.from_pretrained(\"Qwen\/Qwen3-32B\")\n\ninput_text = ['What is the capital of United States?']\n\ninput_tokens = model.tokenizer(\n    input_text,\n    return_tensors=\"pt\",\n    return_attention_mask=False,\n    truncation=True,\n    max_length=MAX_LENGTH,\n    padding=False)\n\ngeneration_output = model.generate(\n    input_tokens['input_ids'].cuda(),\n    max_new_tokens=20,\n    use_cache=True,\n    return_dict_in_generate=True)\n\nprint(model.tokenizer.decode(generation_output.sequences[0]))\n<\/code><\/pre>\n<\/div>\n<p>The AutoModel class itself determines the model type, so the same line can be used for Llama, Qwen, and DeepSeek. And if you take a model like DeepSeek-V3 with 671B parameters, it will also fit in about 12 GB of video memory &#8211; this is largely what makes the project attractive.<\/p>\n<h2>Launch on macOS<\/h2>\n<p>AirLLM runs not only on CUDA, but also on Apple Silicon. On macOS it uses the MLX backend, so will need mlx and torch installed, and only supports Apple chips. There is also a small limitation: the MLX implementation only supports Llama-like architectures, so Qwen and similar models cannot be launched on a MacBook in this way.<\/p>\n<h2>Use example<\/h2>\n<p>AirLLM-experiments are two small interactive console utilities for running large language models locally. They eliminate the need to write the same code every time to load the model and organize the chat.<\/p>\n<p>The first tool, airllm-lib-usage.py, runs the AirLLM library directly in the current Python process. It shows a numbered list of models from the general catalog, if necessary, it installs missing dependencies and opens a chat session. MLX runtime is used on macOS, torch is used on other platforms.<\/p>\n<p>The second tool, airllm-openai-server-client.py, helps you work with the airllm-openai-server server, which runs in Docker and provides an OpenAI-compatible API. When you launch it for the first time, the utility asks for settings, collects the image, picks up the container and opens a streaming chat to the running server. When launched repeatedly, it finds an existing container and simply reuses it, and stores the model cache in a separate Docker volume so as not to download the model again.<\/p>\n<p>The list of available models is included in the common module model_catalog.py, so both tools work with the same set. Chat commands are also common: \/help, \/clear, \/exit. The source code and instructions are in the repository:<\/p>\n<p><a href=\"https:\/\/github.com\/zefir1990\/AirLLM-experiments\" rel=\"noopener\" target=\"_blank\">https:\/\/github.com\/zefir1990\/AirLLM-experiments<\/a><\/p>\n<h2>What to pay attention to<\/h2>\n<p>The first response can take several minutes to prepare: first, AirLLM downloads the model and cuts it into layers, and only then starts generating. Further, the speed, however, also remains low &#8211; this is a deliberate compromise for the sake of saving memory. For gated models like meta-llama, you need to accept the terms on Hugging Face and pass an access token. It\u2019s best not to delete the cache of downloaded models and shards, otherwise everything will have to be downloaded again.<\/p>\n<h2>Links<\/h2>\n<p><a href=\"https:\/\/github.com\/lyogavin\/airllm\" rel=\"noopener\" target=\"_blank\">https:\/\/github.com\/lyogavin\/airllm<\/a><br \/>\n<a href=\"https:\/\/github.com\/mkamranr\/airllm-openai-server\" rel=\"noopener\" target=\"_blank\">https:\/\/github.com\/mkamranr\/airllm-openai-server<\/a><br \/>\n<a href=\"https:\/\/github.com\/zefir1990\/AirLLM-experiments\" rel=\"noopener\" target=\"_blank\">https:\/\/github.com\/zefir1990\/AirLLM-experiments<\/a><br \/>\n<a href=\"https:\/\/pypi.org\/project\/airllm\/\" rel=\"noopener\" target=\"_blank\">https:\/\/pypi.org\/project\/airllm\/<\/a><br \/>\n<a href=\"https:\/\/huggingface.co\/\" rel=\"noopener\" target=\"_blank\">https:\/\/huggingface.co\/<\/a><\/p>\n<h2>Sources<\/h2>\n<p><a href=\"https:\/\/github.com\/lyogavin\/airllm\" rel=\"noopener\" target=\"_blank\">https:\/\/github.com\/lyogavin\/airllm<\/a><br \/>\n<a href=\"https:\/\/github.com\/zefir1990\/AirLLM-experiments\" rel=\"noopener\" target=\"_blank\">https:\/\/github.com\/zefir1990\/AirLLM-experiments<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Typically, running a large language model with 70 billion parameters requires several tens of gigabytes of video memory. AirLLM approaches this problem more calmly and offers a different way: such a model can be run on a video card with only 4 GB of memory, and without quantization, distillation and trimming. In this note, I [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[],"class_list":["post-192","post","type-post","status-publish","format-standard","hentry","category-notes"],"translation":{"provider":"WPGlobus","version":"3.0.6","language":"en","enabled_languages":["en","ru","zh","de","ja","fr","es","pt","hi"],"languages":{"en":{"title":true,"content":true,"excerpt":false},"ru":{"title":true,"content":true,"excerpt":false},"zh":{"title":true,"content":true,"excerpt":false},"de":{"title":true,"content":true,"excerpt":false},"ja":{"title":true,"content":true,"excerpt":false},"fr":{"title":true,"content":true,"excerpt":false},"es":{"title":true,"content":true,"excerpt":false},"pt":{"title":true,"content":true,"excerpt":false},"hi":{"title":true,"content":true,"excerpt":false}}},"_links":{"self":[{"href":"https:\/\/demensdeum.com\/blog\/wp-json\/wp\/v2\/posts\/192","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/demensdeum.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/demensdeum.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/demensdeum.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/demensdeum.com\/blog\/wp-json\/wp\/v2\/comments?post=192"}],"version-history":[{"count":2,"href":"https:\/\/demensdeum.com\/blog\/wp-json\/wp\/v2\/posts\/192\/revisions"}],"predecessor-version":[{"id":194,"href":"https:\/\/demensdeum.com\/blog\/wp-json\/wp\/v2\/posts\/192\/revisions\/194"}],"wp:attachment":[{"href":"https:\/\/demensdeum.com\/blog\/wp-json\/wp\/v2\/media?parent=192"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/demensdeum.com\/blog\/wp-json\/wp\/v2\/categories?post=192"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/demensdeum.com\/blog\/wp-json\/wp\/v2\/tags?post=192"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}