Qualcomm's latest smartphone silicon can run a 30-billion-parameter mixture-of-experts model locally, pushing AI inference off the cloud.
Qualcomm launched two new smartphone chips built with AI as the core selling point, with its top-tier chip capable of running a 30B parameter mixture-of-experts model directly on the device. That's a significant jump in on-device capability that reduces dependency on cloud inference for everyday AI tasks.
The move fits a broader industry push to shift inference workloads to edge hardware, cutting latency and cloud compute costs for app developers building AI-native mobile experiences.
On-device inference at this scale changes the cost calculus for mobile AI apps — companies paying per-token API fees today should evaluate what workloads can move to the edge as this hardware ships into phones.
The daily signal, curated. Get it in your inbox.
Subscribe on LinkedIn →