colibri 1.11.0 reports a 744B model in 9.9 GB of resident RAM. Here's how expert streaming works, and what that memory number leaves out.
As of September 14, 2026, JustVugg's GitHub repo describes a pure C engine with zero engine dependencies and nine supported model families. colibri 1.11.0 streams experts from NVMe, placing weights across storage, RAM and VRAM. My rough arithmetic puts 744 billion int4 weights at about 370 GB before overhead. The weights still need somewhere to live. Mixture-of-experts routing means only a fraction needs to work on each token.
The two GLM-5.2 demos tell different stories:
- STREAMING CPU: int4, ready in 32 seconds, 9.9 GB resident. No tokens-per-second figure or exact host and drive published for that run.
- SIX RTX 5090s: full expert residency, 4 tokens per second, 1.6 seconds to first token, zero disk traffic.
That second speed number does not belong to the first memory number. Kimi K3 at 2.8 trillion parameters appears on the supported roster, but the narration identifies no hardware-and-speed demo for it.
The interesting part is colibri's DEFAULT POLICY: less fast memory can mean more waiting, but it must not silently change precision or router behaviour. Ask for GLM-5.2 at int4 and the engine's stated rule is to preserve that choice, including the experts the router selects. Its caching policies get the same treatment: history can overfit, lookahead can lose on some machines, and dual-SSD striping still needs broader end-to-end comparisons. Those qualifications are useful engineering information.
Then there's the part I'd spend too long watching. Brain displays all 19,456 experts, coloured by storage tier and lit by routing activity. Atlas maps 13,260 characterised experts, including 1,041 replicated specialists. Their positions come from measured routing affinity: experts sit near each other because they receive similar kinds of tokens. You can see the memory hierarchy and the model's routing behaviour, then open the C file responsible.
Watch to the end for the missing measurements and the hardware tradeoffs I'd check before planning a local setup.
⏱️ Chapters:
00:00 colibri 1.11.0: 744B, 9.9 GB
00:21 Whose Hardware, Which Speed?
00:49 Frontier Models Meet Memory Multitiering
01:14 Nine Families, Up To 2.8T
01:55 One C File Per Family
02:14 GLM-5.2 Says Ciao
02:42 The Mixture-Of-Experts Connection
02:48 What The Router Actually Does
03:21 Where Should Idle Experts Live?
03:35 VRAM, RAM And NVMe
04:01 The 370 GB Weight Estimate
04:26 A JIT For Weights
04:38 LRU, Hot-Store And Prefetch
05:12 Route, Fetch, Compute, Repeat
05:29 When Prefetch Can Lose
05:50 Four Storage I/O Techniques
06:12 O_DIRECT And Dual-SSD Striping
06:48 The SSD Evidence Still Needed
07:03 CPU, CUDA And Metal Together
07:29 Six RTX 5090s: The Dashboard
07:57 Two Demos, Two Different Claims
08:28 Preserve The Model By Default
08:43 Precision And Routing Stay Intact
09:14 Experiments Must Earn Their Place
09:43 Measure The Whole Chat
10:06 The Repo's Visual Side
10:10 Brain: Watch 19,456 Experts
10:39 Atlas: Mapping Expert Affinity
11:09 SQL Has Neighbours
11:18 Published Figures And Missing Measurements
12:10 What Developers Can Take Away
12:13 Private Access Through Local Storage
12:33 Choose Your Memory Tier Mix
12:53 Readable C, Measurable Contributions
13:24 Four Languages, Community Contributions
13:38 Run, Watch And Change colibri
13:49 Holding The Model Locally
🔗 Sources:
- Repository: github.com/JustVugg/colibri
- README, Demos and Reported Figures: github.com/JustVugg/colibri
#Colibri #MixtureOfExperts #LocalAI #LLMInference #SignalCoders
👇 Subscribe — a new plain-English breakdown of AI engineering, open source tools, and local LLMs every single day.
YouTube:
www.youtube.com/@Signalcoders
Twitter/X: @SignalCoders
Our Playlists:
https://www.youtube.com/playlist?list=PLL70bDS-hH7Y
https://www.youtube.com/playlist?list=PLTnqXOgLPogE
https://www.youtube.com/playlist?list=PLZVB6navai50
https://www.youtube.com/playlist?list=PLEj9VzjWmmho
⚠️ Disclaimer: Educational purposes only. Signal Coders is not affiliated with or sponsored by JustVugg, colibri or any model or hardware provider mentioned. This is a reading of the public repository as presented on September 14, 2026. Nothing was installed, run, benchmarked or independently reproduced here. Demo figures and capability claims are the project's own; the roughly 370 GB weight estimate is my arithmetic before overhead. The streaming CPU and six-GPU demos describe different configurations, and the published GPU speed does not establish streaming CPU performance. Opinions are commentary, not measured findings. Verify current documentation and hardware requirements before use.
Comments (0)