A Phone That Doesn’t Need the Cloud to Think
Apple in talks with startup that shrinks AI models to run on an iPhone
https://www.cnbc.com/2026/07/14/apple-prismml-ai-compression-iphone.html
PrismML says its compressed version of Alibaba’s Qwen model uses up to 15 times less memory, potentially advancing Apple’s AI push.
Filed July 15, 2026 at 12:49 pm

Apple's in early talks with PrismML, a Caltech spinout that just showed it can take a 27-billion-parameter model, shrink it from 54GB down to under 4GB, and run the whole thing on an iPhone — no cloud round-trip required. The technique strips each weight down to almost nothing, one to three possible values instead of the usual 16-bit precision, and the payoff is real: 10-15x less memory, 6-8x faster responses, a third to a sixth of the power draw, with only a small hit to accuracy that shows up mostly in factual recall rather than reasoning or math. It's early, unproven at scale, and Apple hasn't confirmed anything. But if it holds up, it's a preview of where this is all heading — the intelligence moving off the server and onto the device in your pocket, closer to you, faster, and harder for anyone else to see.