aiminute. ← All AI news
Tools 2026-08-04

Swiftlet squeezes an 80-billion-parameter model into 4.3 GB of RAM — and a 35B onto an iPhone

Swiftlet squeezes an 80-billion-parameter model into 4.3 GB of RAM — and a 35B onto an iPhone

A new open-source Swift-and-Metal runtime called Swiftlet streams mixture-of-experts weights from disk on demand: attention layers, routers and a cache of frequently used 'experts' stay in memory while the rest is fetched from storage. The result is Qwen3-Next-80B running on an M5 Mac in about 4.3 GB of RAM at 4.5–5 tokens per second, and a 35B model running on an iPhone 17 at roughly a token per second. Trade-offs remain: you need 42 GB of free disk for the big model, and the author warns that such sparse models 'chat like large models but recall facts like small ones.' The project topped Hacker News and passed 280 GitHub stars within a day.

Why it mattersThe memory wall — not the chip — has been the reason big models don't run on consumer devices. If expert streaming holds up, the hardware you already own gets a class of models that until now required a GPU workstation, which matters for privacy, offline use and anyone who can't pay for API calls.

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Apple Music will tell you when a song was made by a machine — but the uploader decides whether to say so
2026-08-21
Stripe has just paid $7.5 billion for a model router. Days later Ramp built one and is giving it away until January.
2026-08-21
Meta's assistant is now a Mac app that reads your screen and types into any window — and what it sees can train the model
2026-08-21
One click on a news site now tells Google to show you more of it — and you stay on the page you were reading
2026-08-21
Adobe will now generate the music, the voiceover and the door slam — and it says the licence covers you
2026-08-21