A new era of 'pocket-scale benchmarking' has begun, measuring the performance of AI models running directly on our smartphones rather than in massive data centers, with the iPhone 17 Pro demonstrating the highest performance to date.
Imagine this: even in the remotest area without any internet connection, your smartphone’s AI assistant seamlessly retouches your photos, summarizes long documents, and provides instant interpretations of complex foreign languages. Until now, AI was considered to operate only “in the clouds”—that is, within the supercomputers of massive data centers. However, AI is now entering our small smartphones. This is professionally referred to as ‘Pocket-Scale Inference’ (the process of running AI models directly on the device to derive results).
Just how smart is the AI packed into your smartphone compared to the AI in the clouds? A new standard has been established to verify this.
Why Is This Important?
Until recently, most AIs we use, like ChatGPT, relied on powerful servers. The questions you input would travel over the internet to a distant server, generate an answer there, and send it back to your phone. In contrast, pocket-scale AI completes all calculations within your phone.
There are two main reasons why this is significant. The first is ‘Privacy.’ Your personal conversations or sensitive photo data do not leave your device to go to an external server, making it much safer. The second is ‘Speed.’ You can react instantly regardless of network conditions. However, the catch is that smartphones are much smaller and have more limited performance than server-grade supercomputers. The AI performance we experience depends on how efficiently the smartphone runs this ‘small AI.’
In Simple Terms
If server-grade AI is like a ‘large hotel kitchen with top-tier chefs,’ pocket-scale AI is like a ‘mini kitchen in a studio apartment.’ A large kitchen can prepare hundreds of meals at once, but a mini kitchen has clear limits on what it can produce at one time.
Recently, the AI performance analysis firm Artificial Analysis released a benchmark (a performance measurement standard) to gauge how fast and accurately AI produces results in the cramped kitchen of a smartphone. Source: Artificial Analysis
However, this measurement is trickier than it seems. Unlike server environments in data centers, the runtime (the software environment for running AI) on smartphones is still technically immature. Source: Artificial Analysis It is as if chefs were using completely different sets of tools; the speed and quality of the AI’s responses change drastically depending on the settings. This makes it much harder to measure true capability.
Current Status
Which device is currently leading this ‘pocket-scale AI’ race? According to recent analysis results, the iPhone 17 Pro has topped the chart, recording the best performance in both intelligence (the model’s reasoning power) and speed (response time). Source: Zeli
Artificial Analysis is collaborating with Liquid AI to collect real-world data on how well AI functions on actual devices. Source: Artificial Analysis They are basing their criteria on actual ‘response speed’ and ‘context awareness’ that we feel when using apps in daily life, rather than just theoretical figures. Source: GIGAZINE
Of course, challenges remain. Due to the small memory capacity of smartphones, there are significant gaps compared to data-center-grade AI in areas like ‘context limits’ (the amount of information the AI can remember at once) and the time it takes to produce an answer. Source: Zeli
Future Outlook
In the future, the key metric for smartphone performance will quickly shift from ‘how high-resolution a video it can shoot’ to ‘how smart an AI it can run internally.’ In the open-source community, technologies that analyze a user’s smartphone chipset environment and automatically apply optimal AI settings are already emerging. Source: PocketTune GitHub
Soon, we will enter an era where we no longer ‘rent’ smart AI assistants from servers, but instead carry them inside our own smartphones, allowing us to ask questions anytime, anywhere. Checking a smartphone’s ‘AI benchmark score’ before purchasing one might soon become a basic necessity.
References
-
[Intelligence at pocket scale: Benchmarking small models and mobile phones Artificial Analysis](https://artificialanalysis.ai/articles/mobile-phone-intelligence-inference) -
[Benchmarking Pocket-Scale Inference Hacker News](https://news.ycombinator.com/item?id=49469786) - Benchmarking Pocket-Scale Databases
-
[Vue HN 2.0 Intelligence at pocket scale: Benchmarking small models and mobile phones](https://vue-hackernews-ssr-5cavbdjcta-ew.a.run.app/item/49420960) -
[Artificial Analysis (@ArtificialAnlys) Vanlett](https://vanlett.net/ArtificialAnlys) - Artificial Analysis has published the results of its… - GIGAZINE
-
[Consumer Inference Systems Artificial Analysis](https://artificialanalysis.ai/hardware-inference-stack/mobile-phones) - iPhone 17 Pro tops pocket-scale AI benchmark
-
[Open-Source Agentic Inference Benchmark InferenceX](https://inferencex.semianalysis.com/) - GitHub - ayanbag/PocketTune: On-device tuning of local-LLM
- Google Scholar
- DBpia - Academic AI Platform providing domestic papers, journals, and magazines
- NVIDIA Blackwell Sets New Standard for Gen AI in MLPerf Inference…
-
[Benchmark MLPerf Inference: Datacenter MLCommons V3.1](https://mlcommons.org/benchmarks/inference-datacenter/)
- To measure the smartphone's battery life
- To verify actual AI performance in real-world usage environments rather than in data centers
- To increase the frame rate of mobile games
- Galaxy S26
- iPhone 17 Pro
- Google Pixel 11
- Because communication speeds are too fast compared to data centers
- Because the mobile runtime environment is less mature than data centers, causing results to vary significantly based on settings
- Because smartphones do not have AI chips installed