aiminute. ← All AI news
Tools AI Minute Newsroom 2026-08-04

Swiftlet squeezes an 80-billion-parameter model into 4.3 GB of RAM — and a 35B onto an iPhone

Swiftlet squeezes an 80-billion-parameter model into 4.3 GB of RAM — and a 35B onto an iPhone

A new open-source Swift-and-Metal runtime called Swiftlet streams mixture-of-experts weights from disk on demand: attention layers, routers and a cache of frequently used 'experts' stay in memory while the rest is fetched from storage. The result is Qwen3-Next-80B running on an M5 Mac in about 4.3 GB of RAM at 4.5–5 tokens per second, and a 35B model running on an iPhone 17 at roughly a token per second. Trade-offs remain: you need 42 GB of free disk for the big model, and the author warns that such sparse models 'chat like large models but recall facts like small ones.' The project topped Hacker News and passed 280 GitHub stars within a day.

Why it mattersThe memory wall — not the chip — has been the reason big models don't run on consumer devices. If expert streaming holds up, the hardware you already own gets a class of models that until now required a GPU workstation, which matters for privacy, offline use and anyone who can't pay for API calls.

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

TikTok put a shopping chatbot inside the video you are watching.
2026-10-06
Meta's assistant keeps an hourly file on everyone in your life.
2026-10-05
One command restores the Apple Intelligence off switch Apple deleted.
2026-10-05
OpenAI will ship a Codex upgrade daily for 28 days or reset limits.
2026-10-05
Meta open-sourced the firmware for building your own Muse gadget.
2026-10-04