Apple in talks with startup that shrinks AI models to run on an iPhone

CNBC · 2026-07-14

Apple is reportedly evaluating technology from PrismML, a Caltech spinout that claims its method can shrink powerful AI models to run directly on an iPhone, using up to 15 times less memory, generating responses 6 to 8 times faster, and consuming 3 to 6 times less energy. This breakthrough could significantly enhance Apple's AI strategy by enabling faster, more private Siri interactions through on-device processing, reducing reliance on cloud computing, lowering costs, and supporting offline functionality. While PrismML CEO Babak Hassibi describes the discussions with Apple as early but "progressing nicely," analysts like Carolina Milanesi of Creative Strategies emphasize the advantage of on-device AI for privacy-sensitive data like health information. PrismML achieves this by simplifying internal information storage within AI models, reducing values from 16 bits to just 1 or 3 possible values, a greater compression than the chip industry's move from eight-bit to four-bit computing. Despite a potential slight performance trade-off, particularly in factual recall, the technology promises to reshape memory and datacenter compute demand, though it's acknowledged that overall chip demand may not necessarily decrease but rather shift from datacenters to individual devices.

The full article also explores expert skepticism regarding PrismML's claims, especially concerning real-world testing, battery consumption during multitasking, and performance reliability across millions of queries and diverse device combinations.

Read the original report at CNBC