The Colibrì proof-of-concept allows the GLM-5.2 AI model to run on consumer machines equipped with only 25 GB of RAM. By leveraging streaming technology, the implementation supports the 744-billion-parameter Mixture-of-Experts (MoE) architecture efficiently, storing expert pathways on disk rather than in RAM.
The model operates in pure C without dependencies like Python or GPUs, using a disk-based approach to manage memory needs. Only about 9.9 GB of RAM is used to host the dense parts, while expert pathways are streamed from a 370 GB disk space, utilizing techniques like Least Recently Used (LRU) caching.
Despite the innovative memory management, practical usability is affected by low processing speeds, ranging between 0.05 to 0.1 tokens per second. This rate results in very slow interactive response times, significantly below the 20-30 tokens per second required for real-time applications.
This method addresses accessibility challenges posed by the high costs of AI model subscriptions and the need for privacy in home lab setups. While current speed limitations hinder broader use, the ability to run advanced AI models on consumer-grade hardware is a promising step toward democratizing AI capabilities.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
The Colibrì proof-of-concept allows operation of the 1.5-TB GLM-5.2 AI model on systems with just 25 GB of RAM. This approach addresses high subscription costs for AI models and enhances accessibility for home lab setups, despite significantly lower processing speeds.
A new implementation allows the GLM-5.2 model to run on consumer machines with approximately 25 GB of RAM by utilizing streaming technology for its mixture-of-experts (MoE) architecture. This approach significantly reduces memory requirements and enables advanced model capabilities without relying on high-end hardware or dependencies like Python or GPUs.