TL;DR: A private Qwen 35B MoE model was successfully tested on a Samsung S26 Ultra, demonstrating significant on-device LLM capabilities.
Summary: An independent developer reported successfully running a private Qwen 35B MoE (Mixture of Experts) LLM on a Samsung S26 Ultra. Early tests indicate the model fits within the device's memory, achieving approximately 90 input processing tokens per second and 8 output tokens per second after optimization.
Why it matters: This showcases the increasing feasibility of deploying large, powerful MoE models directly on high-end mobile devices, opening new avenues for offline AI applications. AI builders should explore optimizing models for on-device execution and consider the potential for advanced mobile-first AI experiences.
Source: reddit