The Maple-Preview model, a ternary 20B Mixture-of-Experts (MoE) architecture, has been shown to operate at a speed of 120 tokens per second when running on an iPhone. This performance metric highlights its capability for local inference on consumer mobile hardware.
Maple-Preview utilizes a ternary quantization scheme, which reduces the precision of the model's weights to three values. This approach, combined with the Mixture-of-Experts design, contributes to its efficiency and ability to run on devices with limited computational resources like smartphones.
The ability to run a 20-billion parameter model at high speeds on an iPhone suggests advancements in optimizing large language models for edge devices. This could lead to more powerful AI applications that operate without constant cloud connectivity, offering benefits in privacy, latency, and offline functionality.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
A new AI model, Maple-Preview, a ternary 20B Mixture-of-Experts (MoE) model, has demonstrated performance of 120 tokens per second on an iPhone. This development indicates progress in running large language models efficiently on mobile devices.