Appearance
Mobile Deployment
1. Definition: Running Deep Learning Under a "5 W Power Budget"
Mobile inference means executing neural network inference on smartphones, tablets, and embedded devices (Raspberry Pi, Jetson, IoT modules). Its essential difference from server inference is that the constraint dimensions are different:
| Constraint | Server inference | Mobile inference |
|---|---|---|
| Compute | Tens of TFLOPS and up | Single-digit TFLOPS |
| Memory | 80–640 GB HBM | 4–16 GB unified memory |
| Power | Hundreds of watts | Under 5 W (phone) |
| Cooling | Active air / liquid cooling | Passive; throttles within seconds |
| Latency SLA | Milliseconds | ~100 ms (user-perceived) |
| Storage | Terabytes | A few GB (no room for large models) |
The mobile "power wall" is the core constraint: the iPhone 15 Pro sustains ~5 W — beyond that it throttles and dims the screen. Any design that "runs the GPU at 100% for 10 seconds straight" will trigger throttling. This is why mobile optimization cares about "how many tokens per joule," not just "peak tokens per second."
The key to understanding the mobile inference stack is the hardware abstraction layer (HAL): upper-layer frameworks (PyTorch Mobile / TFLite) → intermediate computation graphs (ONNX / TFLite FlatBuffer) → lower-layer hardware backends (CoreML / NNAPI / Hexagon / ANE). See Hardware Primer.
2. The Landscape of Mobile Inference Backends
| Platform | Official framework | Acceleration backend | Typical use case |
|---|---|---|---|
| iOS / macOS | CoreML | Apple Neural Engine (ANE), Metal | Inference embedded in iOS apps |
| Android | NNAPI | Hexagon DSP, NPU, Mali GPU | Inference embedded in Android apps |
| Cross-platform | PyTorch Mobile | CPU / Metal / OpenCL / Vulkan | Migrating server-side stacks |
| Cross-platform | TensorFlow Lite | NNAPI / CoreML / GPU Delegate | Continuing the TF ecosystem |
| Cross-platform | ONNX Runtime Mobile | EP mechanism (see ONNX Runtime) | Fallback across many platforms |
| LLM-specific | MediaPipe LLM Inference | GPU + built-in quantization | Google's LLM tooling |
| LLM-specific | MLC-LLM | Metal / Vulkan / OpenCL | Cross-platform LLM compiler path |
| LLM-specific | llama.cpp | Metal / Vulkan / NEON | Running LLMs on iOS / Android |
Terminology for ANE, Hexagon, and NPUs
- ANE (Apple Neural Engine): the NPU inside iPhone A-series chips, a 16-core architecture peaking at 11 TOPS (A17 Pro);
- Hexagon DSP: the DSP/HTA (Hexagon Tensor Accelerator) inside Qualcomm Snapdragon;
- Huawei NPU (Da Vinci architecture): an embedded NPU in the Mate series, with the CANN inference framework;
- MediaTek APU: the NPU in the Dimensity series. All of these are called "NPUs" (neural processing units), but their APIs and operator support differ widely — which is exactly why the NNAPI / CoreML abstraction layers matter.
3. Model Conversion and Quantization
iOS: CoreML
python
import coremltools as ct
import torch
import torchvision
# PyTorch → CoreML
model = torchvision.models.resnet50(weights=torchvision.models.ResNet50_Weights.IMAGENET1K_V1).eval()
sample = torch.rand(1, 3, 224, 224)
traced = torch.jit.trace(model, sample)
mlmodel = ct.convert(
traced,
inputs=[ct.TensorType(shape=sample.shape, name="input")],
minimum_deployment_target=ct.target.iOS15,
compute_units=ct.ComputeUnit.ALL, # auto-select CPU+GPU+ANE
)
# Quantize
mlmodel_int8 = ct.optimize.coreml.optimize_linear_quant_weights(mlmodel)
mlmodel.save("Resnet50.mlpackage")
mlmodel_int8.save("Resnet50_int8.mlpackage")CoreML quantization tiers:
- float16: 2× size reduction, near-lossless
- int8 dynamic: 4× reduction, activations quantized at runtime, CPU-friendly
- int8 static: 4× reduction + speedup, requires calibration data
- 4-bit palettization: 8× reduction, accuracy loss must be evaluated
Android: TFLite
python
import tensorflow as tf
converter = tf.lite.TFLiteConverter.from_keras_model(model)
# Quantization options
def representative_dataset():
for _ in range(100):
yield [np.random.rand(1, 224, 224, 3).astype(np.float32)]
converter.optimizations = [tf.lite.Optimize.DEFAULT] # dynamic int8
converter.representative_dataset = representative_dataset # static int8
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type = tf.int8
converter.inference_output_type = tf.int8
tflite_int8 = converter.convert()
with open("model_int8.tflite", "wb") as f:
f.write(tflite_int8)TFLite quantization tiers:
| Type | API | Size | Speed | Accuracy |
|---|---|---|---|---|
| FP32 (baseline) | - | 1× | 1× | 100% |
| FP16 | target_spec.supported_types = [tf.float16] | 0.5× | 1.3× | ~99.9% |
| Dynamic-range quantization (INT8 weights) | optimizations=[DEFAULT] | 0.25× | 1.5–2× | ~99% |
| Full INT8 | + representative_dataset | 0.25× | 2–3× | ~98% |
| INT4 (experimental) | custom | 0.125× | - | ~95% |
See Model Quantization Fundamentals and Pruning and Sparsification.
4. NPU Acceleration: A Love-Hate Relationship
In theory the NPU is the most powerful mobile accelerator; in practice, the experience is another story:
Apple Neural Engine
- Pros: CoreML schedules automatically; the ANE suits convolutions and matrix multiplications
- Cons: the ANE supports a limited set of operators; hitting an unsupported operator "falls back to GPU/CPU," and switching devices costs overhead
- Diagnostics:
puts(mlmodel.get_compute_unit()to see where it actually runs
Qualcomm Hexagon DSP
- Pros: extremely strong INT8 performance — ResNet-50 INT8 in ~3 ms on Snapdragon 8 Gen 3
- Cons: requires the Qualcomm QNN SDK / SNPE and a long toolchain; only works on Snapdragon devices
- Diagnostics: use Snapdragon Profiler to inspect dispatching
Huawei NPU (Da Vinci Architecture)
- Pros: strong INT8 performance on Mate 60+
- Cons: requires CANN / MindSpore Lite, disconnected from the cross-platform ecosystem
- Diagnostics: MindSpore Lite provides benchmark tools
The NPU is not a silver bullet
- Low operator coverage: many custom attention / new activation functions are unsupported by NPUs, and automatic fallback to CPU is often slower;
- Poor cross-device consistency: the iPhone 13 and iPhone 15 NPUs differ by 2× in performance, and older models may have no NPU at all;
- Hard to debug: an NPU is a black box, and performance bottlenecks are hard to locate;
- Power throttling: sustained full NPU load throttles within seconds, so long tasks lack stability. Production experience: validate correctness on CPU FP16 first, then add GPU + NPU acceleration; benchmark against the 90th percentile rather than the peak (to account for throttling).
5. LLMs on Mobile: The New Direction After 2024
Running LLMs on mobile is challenging — an 8B model in INT4 still needs 4 GB of memory, while the iPhone 15 Pro has 8 GB in total. Feasible paths:
1. llama.cpp on iOS / Android
The most direct option (see llama.cpp):
bash
# iOS: build llama.cpp as an iOS framework
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
mkdir build && cd build
cmake -DLLAMA_METAL=ON -DLLAMA_BUILD_IOS=ON ..
make
# Integrate into an Xcode project and call from SwiftMeasured numbers:
| Device | Model | Quantization | Speed | Memory footprint |
|---|---|---|---|---|
| iPhone 15 Pro 8GB | Llama-3-8B | Q4_K_S | 10 tokens/s | 4.5 GB |
| iPhone 15 Pro 8GB | Phi-3-mini (3.8B) | Q4_K_M | 22 tokens/s | 2.3 GB |
| iPad Pro M4 16GB | Llama-3-8B | Q4_K_M | 35 tokens/s | 4.5 GB |
| Samsung S24 Ultra | Llama-3-8B | Q4_K_M | 12 tokens/s | 4.5 GB |
2. MLC-LLM: The Compiler Path
MLC-LLM compiles LLMs into Vulkan / Metal / OpenCL kernels, running across iOS / Android / WebGPU. Pros: consistent cross-platform behavior; cons: complex debugging and lagging support for new models.
python
# MLC-LLM compiles Llama-3-8B for iOS
from mlc_llm import compile_model
compile_model(
model="meta-llama/Meta-Llama-3-8B-Instruct",
quantization="q4f16_1",
target="iphone",
output_dir="dist/llama3-8b-iphone",
)3. MediaPipe LLM Inference
The on-mobile LLM inference framework Google launched in 2024, supporting Gemma / Llama / Phi and more. Pros: deeply integrated with the Android ecosystem; cons: weak cross-platform support (limited iOS support).
4. ONNX Runtime Mobile + LLM
ONNX Runtime Mobile is the lightweight version of ORT. It can run LLMs but is less optimized than the options above. See ONNX Runtime.
6. The Constraint Checklist for On-Mobile LLMs
The checklist for squeezing an 8B model into a phone
- Total memory: iPhone 8 GB / Android 8–12 GB; the OS takes 2–3 GB, the app takes 1–2 GB, leaving 4–6 GB for the model + KV cache;
- Model size: 8B INT4 ≈ 4 GB; it must be decompressed on the fly at download time (decompressing to disk is too slow);
- First-token latency: model loading takes 1–3 s, and prefilling a 200-token prompt takes 0.5–1 s — show a loading state in the UX;
- Sustained generation thermals: an iPhone throttles to 60% performance within seconds under full NPU load — reflect this in benchmarks;
- Battery: generating 1000 tokens drains about 1–2% of the battery — users notice;
- Privacy compliance: local inference is an advantage for privacy-sensitive scenarios (healthcare, finance) and can be a selling point;
- First-token optimization: use prefix caching for the system prompt (supported since llama.cpp 0.2+);
- Streaming UX: streaming output is a must, or users won't wait.
7. Performance Data: Baseline Reference
Baselines on the iPhone 15 Pro (A17 Pro) (see Inference Benchmarking in Practice):
| Model | Type | Backend | Speed | Power |
|---|---|---|---|---|
| ResNet-50 INT8 | Vision | ANE | 4 ms / 250 img/s | 1.5 W |
| MobileNetV3 INT8 | Vision | ANE | 1.5 ms / 650 img/s | 0.8 W |
| YOLOv8-n INT8 | Detection | GPU + ANE | 8 ms | 2.5 W |
| BERT-base INT8 | NLP | ANE | 12 ms | 2.0 W |
| Whisper-tiny INT8 | Speech | ANE | 60 ms (real-time factor 0.2) | 1.8 W |
| Llama-3-8B Q4 | LLM | Metal GPU | 10 tokens/s | 4.0 W |
| Phi-3-mini (3.8B) Q4 | LLM | Metal GPU | 22 tokens/s | 2.5 W |
8. Trade-offs
- Quantization accuracy vs. speed: INT4 / INT8 is the only path, but every tier down must be evaluated for task-specific loss;
- NPU vs. GPU: the NPU is fast but has incomplete operator coverage, with the GPU as fallback; many models run "part NPU + part GPU";
- Local vs. cloud: local offers low latency and privacy; cloud offers bigger, more capable models; hybrid architectures (small requests local, large ones cloud) are common;
- Pre-download vs. download-on-demand: users won't install apps over 200 MB — models should download after first launch;
- Offline vs. online: offline scenarios (airplane mode, the subway) must be local, but model capacity is limited;
- App store compliance: Apple has review rules about "dynamically downloaded models" — assess them in advance.
9. Comparison with Similar Solutions
| Solution | Relationship with mobile deployment |
|---|---|
| llama.cpp | The first choice for on-mobile LLMs |
| ONNX Runtime | ORT Mobile is the cross-platform fallback |
| OpenVINO | Usable on Intel edge devices, but mobile is not its main battlefield |
| MLC-LLM | The cross-platform LLM compiler path |
| MediaPipe | Google's full-stack mobile solution |
10. Where to Go Next
- Concept pages: Model Quantization Fundamentals, Pruning and Sparsification, Knowledge Distillation, Weight-Only Quantization and Mixed Precision, Latency, Throughput, and Concurrency, The GPU Memory Hierarchy and the Bandwidth Wall
- Case-study pages: llama.cpp, ONNX Runtime, OpenVINO, vLLM
- Practice pages: Inference Engine Comparison, Tuning and Performance Optimization, Inference Benchmarking in Practice, Portfolio Projects, Common Pitfalls and Anti-Patterns
- Resource pages: Hardware Primer, Glossary, Curated Resources
References
- Apple. CoreML Tools documentation — iOS conversion and quantization
- Apple. Metal Performance Shaders — Metal compute
- Google. TensorFlow Lite — Android deployment
- Google. MediaPipe LLM Inference — on-mobile LLMs
- Qualcomm. QNN SDK — Hexagon DSP
- Huawei. CANN / MindSpore Lite — Huawei NPUs
- MLC-LLM — cross-platform LLM compiler
- PyTorch Mobile — PyTorch on mobile
- ONNX Runtime Mobile — the trimmed mobile version of ORT