Skip to content

Mobile Deployment

At a glance iOS CoreML/Metal, Android NNAPI/Hexagon, cross-platform PyTorch Mobile/TFLite — squeezing models into the triple constraints of memory, power, and thermals on iPhone and Android. This article dissects model conversion, quantization, NPU acceleration, and the state of on-mobile LLMs.

Mobile Deployment ​

1. Definition: Running Deep Learning Under a "5 W Power Budget" ​

Mobile inference means executing neural network inference on smartphones, tablets, and embedded devices (Raspberry Pi, Jetson, IoT modules). Its essential difference from server inference is that the constraint dimensions are different:

ConstraintServer inferenceMobile inference
ComputeTens of TFLOPS and upSingle-digit TFLOPS
Memory80–640 GB HBM4–16 GB unified memory
PowerHundreds of wattsUnder 5 W (phone)
CoolingActive air / liquid coolingPassive; throttles within seconds
Latency SLAMilliseconds~100 ms (user-perceived)
StorageTerabytesA few GB (no room for large models)

The mobile "power wall" is the core constraint: the iPhone 15 Pro sustains ~5 W — beyond that it throttles and dims the screen. Any design that "runs the GPU at 100% for 10 seconds straight" will trigger throttling. This is why mobile optimization cares about "how many tokens per joule," not just "peak tokens per second."

The key to understanding the mobile inference stack is the hardware abstraction layer (HAL): upper-layer frameworks (PyTorch Mobile / TFLite) → intermediate computation graphs (ONNX / TFLite FlatBuffer) → lower-layer hardware backends (CoreML / NNAPI / Hexagon / ANE). See Hardware Primer.

2. The Landscape of Mobile Inference Backends ​

PlatformOfficial frameworkAcceleration backendTypical use case
iOS / macOSCoreMLApple Neural Engine (ANE), MetalInference embedded in iOS apps
AndroidNNAPIHexagon DSP, NPU, Mali GPUInference embedded in Android apps
Cross-platformPyTorch MobileCPU / Metal / OpenCL / VulkanMigrating server-side stacks
Cross-platformTensorFlow LiteNNAPI / CoreML / GPU DelegateContinuing the TF ecosystem
Cross-platformONNX Runtime MobileEP mechanism (see ONNX Runtime)Fallback across many platforms
LLM-specificMediaPipe LLM InferenceGPU + built-in quantizationGoogle's LLM tooling
LLM-specificMLC-LLMMetal / Vulkan / OpenCLCross-platform LLM compiler path
LLM-specificllama.cppMetal / Vulkan / NEONRunning LLMs on iOS / Android

Terminology for ANE, Hexagon, and NPUs

  • ANE (Apple Neural Engine): the NPU inside iPhone A-series chips, a 16-core architecture peaking at 11 TOPS (A17 Pro);
  • Hexagon DSP: the DSP/HTA (Hexagon Tensor Accelerator) inside Qualcomm Snapdragon;
  • Huawei NPU (Da Vinci architecture): an embedded NPU in the Mate series, with the CANN inference framework;
  • MediaTek APU: the NPU in the Dimensity series. All of these are called "NPUs" (neural processing units), but their APIs and operator support differ widely — which is exactly why the NNAPI / CoreML abstraction layers matter.

3. Model Conversion and Quantization ​

iOS: CoreML ​

python
import coremltools as ct
import torch
import torchvision

# PyTorch → CoreML
model = torchvision.models.resnet50(weights=torchvision.models.ResNet50_Weights.IMAGENET1K_V1).eval()
sample = torch.rand(1, 3, 224, 224)
traced = torch.jit.trace(model, sample)

mlmodel = ct.convert(
    traced,
    inputs=[ct.TensorType(shape=sample.shape, name="input")],
    minimum_deployment_target=ct.target.iOS15,
    compute_units=ct.ComputeUnit.ALL,           # auto-select CPU+GPU+ANE
)

# Quantize
mlmodel_int8 = ct.optimize.coreml.optimize_linear_quant_weights(mlmodel)

mlmodel.save("Resnet50.mlpackage")
mlmodel_int8.save("Resnet50_int8.mlpackage")

CoreML quantization tiers:

  • float16: 2× size reduction, near-lossless
  • int8 dynamic: 4× reduction, activations quantized at runtime, CPU-friendly
  • int8 static: 4× reduction + speedup, requires calibration data
  • 4-bit palettization: 8× reduction, accuracy loss must be evaluated

Android: TFLite ​

python
import tensorflow as tf

converter = tf.lite.TFLiteConverter.from_keras_model(model)

# Quantization options
def representative_dataset():
    for _ in range(100):
        yield [np.random.rand(1, 224, 224, 3).astype(np.float32)]

converter.optimizations = [tf.lite.Optimize.DEFAULT]                  # dynamic int8
converter.representative_dataset = representative_dataset              # static int8
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type = tf.int8
converter.inference_output_type = tf.int8

tflite_int8 = converter.convert()
with open("model_int8.tflite", "wb") as f:
    f.write(tflite_int8)

TFLite quantization tiers:

TypeAPISizeSpeedAccuracy
FP32 (baseline)-1×1×100%
FP16target_spec.supported_types = [tf.float16]0.5×1.3×~99.9%
Dynamic-range quantization (INT8 weights)optimizations=[DEFAULT]0.25×1.5–2×~99%
Full INT8+ representative_dataset0.25×2–3×~98%
INT4 (experimental)custom0.125×-~95%

See Model Quantization Fundamentals and Pruning and Sparsification.

4. NPU Acceleration: A Love-Hate Relationship ​

In theory the NPU is the most powerful mobile accelerator; in practice, the experience is another story:

Apple Neural Engine ​

  • Pros: CoreML schedules automatically; the ANE suits convolutions and matrix multiplications
  • Cons: the ANE supports a limited set of operators; hitting an unsupported operator "falls back to GPU/CPU," and switching devices costs overhead
  • Diagnostics: puts(mlmodel.get_compute_unit() to see where it actually runs

Qualcomm Hexagon DSP ​

  • Pros: extremely strong INT8 performance — ResNet-50 INT8 in ~3 ms on Snapdragon 8 Gen 3
  • Cons: requires the Qualcomm QNN SDK / SNPE and a long toolchain; only works on Snapdragon devices
  • Diagnostics: use Snapdragon Profiler to inspect dispatching

Huawei NPU (Da Vinci Architecture) ​

  • Pros: strong INT8 performance on Mate 60+
  • Cons: requires CANN / MindSpore Lite, disconnected from the cross-platform ecosystem
  • Diagnostics: MindSpore Lite provides benchmark tools

The NPU is not a silver bullet

  1. Low operator coverage: many custom attention / new activation functions are unsupported by NPUs, and automatic fallback to CPU is often slower;
  2. Poor cross-device consistency: the iPhone 13 and iPhone 15 NPUs differ by 2× in performance, and older models may have no NPU at all;
  3. Hard to debug: an NPU is a black box, and performance bottlenecks are hard to locate;
  4. Power throttling: sustained full NPU load throttles within seconds, so long tasks lack stability. Production experience: validate correctness on CPU FP16 first, then add GPU + NPU acceleration; benchmark against the 90th percentile rather than the peak (to account for throttling).

5. LLMs on Mobile: The New Direction After 2024 ​

Running LLMs on mobile is challenging — an 8B model in INT4 still needs 4 GB of memory, while the iPhone 15 Pro has 8 GB in total. Feasible paths:

1. llama.cpp on iOS / Android ​

The most direct option (see llama.cpp):

bash
# iOS: build llama.cpp as an iOS framework
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
mkdir build && cd build
cmake -DLLAMA_METAL=ON -DLLAMA_BUILD_IOS=ON ..
make

# Integrate into an Xcode project and call from Swift

Measured numbers:

DeviceModelQuantizationSpeedMemory footprint
iPhone 15 Pro 8GBLlama-3-8BQ4_K_S10 tokens/s4.5 GB
iPhone 15 Pro 8GBPhi-3-mini (3.8B)Q4_K_M22 tokens/s2.3 GB
iPad Pro M4 16GBLlama-3-8BQ4_K_M35 tokens/s4.5 GB
Samsung S24 UltraLlama-3-8BQ4_K_M12 tokens/s4.5 GB

2. MLC-LLM: The Compiler Path ​

MLC-LLM compiles LLMs into Vulkan / Metal / OpenCL kernels, running across iOS / Android / WebGPU. Pros: consistent cross-platform behavior; cons: complex debugging and lagging support for new models.

python
# MLC-LLM compiles Llama-3-8B for iOS
from mlc_llm import compile_model

compile_model(
    model="meta-llama/Meta-Llama-3-8B-Instruct",
    quantization="q4f16_1",
    target="iphone",
    output_dir="dist/llama3-8b-iphone",
)

3. MediaPipe LLM Inference ​

The on-mobile LLM inference framework Google launched in 2024, supporting Gemma / Llama / Phi and more. Pros: deeply integrated with the Android ecosystem; cons: weak cross-platform support (limited iOS support).

4. ONNX Runtime Mobile + LLM ​

ONNX Runtime Mobile is the lightweight version of ORT. It can run LLMs but is less optimized than the options above. See ONNX Runtime.

6. The Constraint Checklist for On-Mobile LLMs ​

The checklist for squeezing an 8B model into a phone

  1. Total memory: iPhone 8 GB / Android 8–12 GB; the OS takes 2–3 GB, the app takes 1–2 GB, leaving 4–6 GB for the model + KV cache;
  2. Model size: 8B INT4 ≈ 4 GB; it must be decompressed on the fly at download time (decompressing to disk is too slow);
  3. First-token latency: model loading takes 1–3 s, and prefilling a 200-token prompt takes 0.5–1 s — show a loading state in the UX;
  4. Sustained generation thermals: an iPhone throttles to 60% performance within seconds under full NPU load — reflect this in benchmarks;
  5. Battery: generating 1000 tokens drains about 1–2% of the battery — users notice;
  6. Privacy compliance: local inference is an advantage for privacy-sensitive scenarios (healthcare, finance) and can be a selling point;
  7. First-token optimization: use prefix caching for the system prompt (supported since llama.cpp 0.2+);
  8. Streaming UX: streaming output is a must, or users won't wait.

7. Performance Data: Baseline Reference ​

Baselines on the iPhone 15 Pro (A17 Pro) (see Inference Benchmarking in Practice):

ModelTypeBackendSpeedPower
ResNet-50 INT8VisionANE4 ms / 250 img/s1.5 W
MobileNetV3 INT8VisionANE1.5 ms / 650 img/s0.8 W
YOLOv8-n INT8DetectionGPU + ANE8 ms2.5 W
BERT-base INT8NLPANE12 ms2.0 W
Whisper-tiny INT8SpeechANE60 ms (real-time factor 0.2)1.8 W
Llama-3-8B Q4LLMMetal GPU10 tokens/s4.0 W
Phi-3-mini (3.8B) Q4LLMMetal GPU22 tokens/s2.5 W

8. Trade-offs ​

  • Quantization accuracy vs. speed: INT4 / INT8 is the only path, but every tier down must be evaluated for task-specific loss;
  • NPU vs. GPU: the NPU is fast but has incomplete operator coverage, with the GPU as fallback; many models run "part NPU + part GPU";
  • Local vs. cloud: local offers low latency and privacy; cloud offers bigger, more capable models; hybrid architectures (small requests local, large ones cloud) are common;
  • Pre-download vs. download-on-demand: users won't install apps over 200 MB — models should download after first launch;
  • Offline vs. online: offline scenarios (airplane mode, the subway) must be local, but model capacity is limited;
  • App store compliance: Apple has review rules about "dynamically downloaded models" — assess them in advance.

9. Comparison with Similar Solutions ​

SolutionRelationship with mobile deployment
llama.cppThe first choice for on-mobile LLMs
ONNX RuntimeORT Mobile is the cross-platform fallback
OpenVINOUsable on Intel edge devices, but mobile is not its main battlefield
MLC-LLMThe cross-platform LLM compiler path
MediaPipeGoogle's full-stack mobile solution

10. Where to Go Next ​

References ​