On-Device AI Runs Inside a Budget: Latency, Memory and Heat on an iPhone
A model that scores well on a desk, on the newest phone, plugged in, tells you very little about how the feature behaves in a real pocket. The apps whose on-device features last are the ones whose owners decided what the feature was allowed to cost before they decided what it could do.
When a team adds an on-device model to an iOS app, the first question is almost always about accuracy: does it recognise the thing, segment the subject, transcribe the sentence. That is the right question for a week. It is the wrong question for the life of the product, because accuracy is measured on a desk, on a fast phone, plugged in, with nothing else running. Users meet the model on a four-year-old device at 14% battery, in a warm pocket, with a music app and a map in the background. What decides whether the feature survives that is not the model's score. It is whether the model fits inside a budget.
I have written before about what changes when inference moves onto the phone: the economics, the offline behaviour, the things that get harder. This post is about the discipline that makes the harder parts manageable. It is the same discipline a finance person applies to a company: decide what you are allowed to spend before you start spending it, and write the limits down where someone will check them.
Three budgets, not one
An on-device model draws on three resources at once, and they fail in different ways.
The first is latency: how long the user waits between an action and a result. Latency is the budget people already think about, and it has a human threshold rather than a technical one. A result that arrives while the user is still looking at the screen feels like part of the interface. A result that arrives after they have looked away feels like a job running in the background, and they treat it differently — they wonder whether it worked, they tap again, they leave.
The second is memory. This is the budget that kills apps quietly. iOS does not give an app a fixed allowance it can read in advance; it gives an allowance that depends on the device, on what else is resident, and on how the system feels about the app at that moment. When the app crosses the line, it is not slowed down. It is terminated, without a dialog, and the user sees the app return to its launch screen and loses whatever they were doing. A model that loads comfortably on the newest phone can be the direct cause of a crash report that only ever appears on older ones, and the crash will not mention the model, because the system ended the process rather than the code failing.
The third is heat, and it is the one that is least intuitive because it has no single-call signature. One inference is cheap. A thousand inferences in a row — a video processed frame by frame, a camera feed analysed live — sustained over a few minutes, raise the temperature of the device, and the system responds by reducing clock speeds for everything, including your app. The first ten seconds of a session are not representative of the tenth minute. A feature that demos beautifully and feels sluggish after five minutes of use is almost always a thermal problem, not a code problem, and it cannot be found by testing quickly.
On a server, a model that is too expensive shows up on an invoice. On a phone, it shows up as a one-star review that says the app got slow — and nobody connects the review to the model.
Write the budget down before choosing the model
The mistake I see most often, and have made myself, is choosing the model first and discovering its footprint afterwards. The order should be reversed. Before evaluating any model, write down three numbers for the feature: the slowest acceptable response on the oldest device you support, the most memory the feature may add to the app's footprint at its peak, and the longest continuous session the feature must survive without the device visibly warming or the frame rate dropping.
These numbers will be somewhat arbitrary the first time. That is fine; their job is to turn a vague argument — “is this model too heavy?” — into a measurable one, and to make a decision about trade-offs explicit rather than accidental. A smaller model that scores a few points lower but fits inside the budget on every supported device is a better product than a larger one that works on half of them. Whether that trade is worth it is a product judgment, and it should be made by a person looking at the budget, not discovered by a user.
The same habit applies to the device floor. Every model has a point below which it should not be offered at all, and the honest options are to raise the minimum device for that feature, to ship a smaller model for older devices, or to fall back to a server for them. What is not honest is to ship the same model to everyone and accept that a tail of users will have a bad time. That tail is usually the one leaving reviews.
Where the budget actually goes
Once the limits exist, the work is finding where they are being spent. In my experience the biggest savings rarely come from the model itself.
Input size is the largest lever. Models are almost always run on images or frames larger than they need, because the camera produces large images and nobody downsamples before inference. The cost of most vision models scales with the number of pixels, so halving each dimension cuts the work roughly by a factor of four, often with no visible change in the result. The question to ask is not “what resolution does the camera give me” but “what is the smallest input at which the output is still good enough”. Finding that threshold takes an afternoon of comparing outputs side by side, and it is usually the cheapest performance work in the project.
Frequency is the second lever. Live features tend to run the model on every frame because that is the easiest thing to write. Most of the time the answer does not change between adjacent frames. Running inference a few times a second and smoothing or interpolating between results is often indistinguishable to the user and reduces sustained load substantially. It is also the difference between a feature that can run for ten minutes and one that can run for ten seconds.
Residency is the third. Loading a model is expensive in both time and memory, and keeping it loaded costs memory for as long as it is held. The design question is when it needs to be resident at all. A model used once per session should be loaded when the user is about to need it and released when they are done — not at launch, where it inflates the footprint of every screen that has nothing to do with it, and not held forever because unloading felt like a risk. If two models are used in sequence, they do not need to be in memory at the same time, and arranging for that is often the difference between fitting and not fitting.
Only after those three is it worth reaching for quantisation and model compression. They are valuable, and on Apple hardware the tooling is good, but each one trades a measurable amount of quality for footprint, and it is easier to accept that trade knowingly once the free savings have already been taken.
For each model-backed feature, write down: the slowest acceptable response on the oldest supported device; the peak memory the feature may add; the longest continuous session it must survive; the minimum device it is offered on and what older devices get instead; and the input size, call frequency and load/unload moments you have chosen. If any of those lines is blank, nobody has decided it — the system will decide it for you, in production.
Test the way users hold the phone
A budget is only as good as the measurement behind it, and most on-device testing is done in the conditions least like real use. A few changes make it far more honest.
Test on the oldest device you claim to support, not the newest device you own. If you do not own one, the oldest supported phone is an inexpensive purchase compared with one bad release. Test with the device warm: run the feature for a realistic session length, then measure again, because the second measurement is the one users experience. Test on battery, including with Low Power Mode enabled, which changes performance in ways the system does not announce to your app. And test with other work happening — a large photo library loading, audio playing, a navigation app in the background — because memory pressure is relative to everything else that is resident.
It is worth adding measurement to the app itself. Apple's own tools will show you the thermal state of the device and let you observe changes in it, and memory footprint can be tracked through a release. Logging these, in aggregate and without personal data, turns a vague sense that the app feels slow for some people into a number you can compare between versions. For a privacy-first product it is worth being deliberate here: what you record should describe the device's condition, never the content the user was processing.
Degrade on purpose
Even with a good budget, devices will sometimes be hotter, older or busier than you planned for. The difference between a robust feature and a fragile one is what it does at that point. A fragile one keeps running at full cost until the system intervenes. A robust one notices and steps down deliberately: it lowers the input size, reduces the call frequency, switches to the smaller model, or pauses a non-essential background analysis and tells the user plainly that it has done so.
This is worth building early, because it has the same character as a financial contingency plan. You do not know which of your assumptions will fail, but you can decide in advance what you will cut first. Cutting deliberately, on your terms and in a sequence you chose, is a very different experience for the user than having the operating system cut for you.
The budget changes with the product
The last point is about time. The budget you write down at launch is not permanent. New devices raise the ceiling; new OS releases shift the rules; your own app grows and each new feature takes some of the memory that was previously free. A feature that fit comfortably in its first release can quietly stop fitting two years later because of things added around it, with no change to the model at all.
This is another reason to keep the budget as a written artefact. A simple review at each major OS release and each major feature addition, checking the numbers against the current state of the app, catches drift while it is still cheap. It connects directly to the carrying cost of an app in the store: a model in production is not a thing you finish building, it is a standing commitment, and the budget is how you keep track of what it costs.
None of this is exotic, and none of it requires a large team. It requires deciding what the feature is allowed to cost before falling in love with what it can do, measuring under the conditions people actually use it, and building in a graceful way to give something back when those conditions turn out worse than expected. The apps whose on-device features hold up for years are almost never the ones with the cleverest model. They are the ones whose owners treated the phone as a shared resource with a limit, rather than a free one.