Bundling a Quantized Model Inside a Mobile App
10 min read · updated August 11, 2026
The decision to put a model inside the app rather than download it is made on two numbers: how many bytes the weights take, and which published store ceiling those bytes hit first. Both are knowable before you write any code, and one of them is arithmetic you can do in your head.
The arithmetic, before the policy
Weight storage is parameters multiplied by bits per weight, divided by eight. Every quantization format adds a little on top for its scales and zero points or its lookup tables, so the honest form is “plus overhead” rather than a clean figure. Worked through, with the assumptions stated:
bytes = parameters x bits_per_weight / 8 0.5B at 4 bits = 0.5e9 x 4 / 8 = 250 MB (+ scales) 1B at 4 bits = 1.0e9 x 4 / 8 = 500 MB (+ scales) 1B at 8 bits = 1.0e9 x 8 / 8 = 1000 MB (+ scales) 3B at 4 bits = 3.0e9 x 4 / 8 = 1500 MB (+ scales) A 4-bit scheme that stores one fp16 scale per group of 32 weights carries 16 bits / 32 weights = 0.5 extra bits per weight, so "4-bit" is really about 4.5 bits in practice: 3B at 4.5 bits = 3.0e9 x 4.5 / 8 = 1688 MB
That is the number to carry into the rest of this page. It is not a measurement of any particular file — check the actual artefact — but it is the right order of magnitude and it is derived from two things you know. Note also what it excludes: the tokenizer, the runtime library itself, and the KV cache, which is allocated at run time and does not appear in the bundle at all.
What Apple limits
Apple publishes maximum build file sizes in App Store Connect reference documentation (Apple Developer). For an app whose minimum deployment target is iOS 9.0 or later, the maximum uncompressed app size is 4 GB, and the maximum executable file size is 80 MB for the total of all __TEXT sections in the binary.
Both numbers matter and they constrain different things. The 4 GB ceiling is generous enough that a 4-bit 3B model fits with room to spare, which is why the App Store limit is rarely the thing that stops you. The 80 MB executable limit is much tighter, and it is the one people trip over when a runtime is statically linked with large embedded kernels — but note that model weights shipped as a resource file are not part of __TEXT, so a weights file does not count against it. Apple points at Background Assets as the supported route for hosting large assets outside the build (Background Assets).
What Google Play limits
Google Play’s limits are stated in terms of compressed download size as calculated by Play Console on upload, which is a different measurement from Apple’s and catches people out (Play Console Help):
- Base module: 500 MB. This is the one a bundled model hits first, and it is far tighter than Apple’s ceiling.
- Each feature module: 500 MB.
- Each individual asset pack: 1.5 GB.
- All modules and install-time asset packs together: 4 GB.
- On-demand and fast-follow asset packs together: 30 GB, for a stated total maximum of 34 GB.
The practical reading is that a 1B model at 4 bits is roughly at the base-module ceiling on its own, before your application code, and that Play Asset Delivery — specifically an asset pack rather than the base module — is the mechanism Google expects you to use. Play’s own guidance in the same document is that most developers should stay well below these limits, because install size correlates with install completion.
Why the download is not smaller than the file
There is a comfortable assumption that a 500 MB weights file becomes a much smaller download once the store compresses it. It does not, and the reason is mechanical. Quantization is, among other things, a compression step: it has already removed the redundancy that a general-purpose compressor would find. Quantized weights are close to high-entropy data, and high-entropy data does not deflate.
So for planning purposes, treat the compressed download of a quantized weights file as approximately its size on disk. Where compression does help is everything else in the bundle — code, layouts, strings — which is exactly the part that is not your problem. If you want a smaller download, the lever is a smaller model or a lower bit width, not the packaging.
The corollary is more useful than it looks: because the ratio is close to one, the arithmetic at the top of this page is directly comparable to Play’s limits, which are stated in compressed bytes. A 4-bit 1B model at roughly 500 MB is not “about half the base module budget” — it is the base module budget. There is no compression headroom to plan against and no point measuring it twice.
One exception is worth knowing. An unquantized fp32 or fp16 checkpoint does still compress somewhat, because the exponent bytes across a tensor are highly correlated even when the mantissas are not. That is a reason a float model’s download is a smaller multiple of its disk size than you would guess — and not a reason to ship one, since the quantized version is smaller in both measurements.
The answer is usually not to bundle it
Weigh the two failure modes honestly. Bundling means the model is there on first launch with no network, no server, and no way for the feature to be unavailable — and it means every user pays the bytes, including the ones who never touch the feature, and every update to the model is an app update through review. On-demand delivery means a first-run download you have to design an interface for, and it means the feature can fail, and it means you can ship a better model on Tuesday without a release.
There is also a licensing dimension that a bundle makes concrete. Open weights come with licences, and some are gated behind an access request on the model host. Redistributing weights inside your app is distribution, and it is governed by that licence — read it, and where access is gated, treat the gate as a condition to satisfy rather than an obstacle to route around.