News

Apple joins hands with startups to promote AI model compression technology so that iPhones can run large-scale AI

3 min read
Apple teamed up with startups to promote AI model compression technology so that iPhones can also run large-scale AI! Apple is reportedly in talks with Silicon Valley startup PrismML to develop a revolutionary AI model compression technology. The goal of this technology is to compress otherwise large AI models to the point where they can be run natively on an iPhone. If this technology is successfully verified, it will be expected to enhance Apple's advantage in privacy protection and provide strong support for the currently slow-progressing Siri upgrade. The CEO of PrismML revealed that Apple has begun to evaluate the feasibility of its technology, and although the negotiations are still in the early stages, the progress looks quite smooth. It is worth noting that this news coincides with the day after Apple launched the iOS27 public beta version. This version marks the first time Siri is open to the public for testing after a substantial revision, with the purpose of competing with AI assistants such as OpenAI and Anthropic. By keeping more of its AI computing tasks local to the device, Apple can reduce response latency, reduce cloud computing costs, and use some features without an Internet connection. This is highly consistent with Apple’s long-term emphasis on privacy protection. Counterpoint analyst Pathak pointed out that the combination of cloud and device-side AI will provide users with a more comprehensive, more efficient and more privacy-focused AI experience: complex tasks are handled by the cloud, while sensitive information is completed on the device. PrismML was originally incubated at Caltech and is backed by Khosla Ventures. The company has publicly released a compressed version of Alibaba’s open source model Qwen, compressing the size of the original model from approximately 54GB to less than 4GB, allowing the model’s 27 billion parameters to run on iPhone 15 and newer models. In the future, PrismML also plans to handle other large open source models to further advance the technology. At a technical level, PrismML's compression solution can reduce cloud models that originally required 8 GPUs to only 1, and allow models that rely on data centers to be migrated to mobile phones and laptops. Although this approach can effectively reduce the memory and computing power required for a single AI task, it does not mean that overall chip demand will decrease accordingly. Highlights: 🔍 Apple is negotiating with PrismML in an effort to compress large AI models to run on the iPhone. 📱 This technology is expected to improve privacy protection and facilitate the upgrade process of Siri. 💡Pr