Mobile system design
iOS system design interview: a worked answer from the hiring side
The round is not scored on the diagram. It is scored on the decisions. A worked answer for an offline-capable messaging app on iOS, minute by minute, with the follow-ups interviewers use and how the Android answer differs.
7 min read
The candidate drew a clean diagram in eight minutes. View, view model, repository, API client, Core Data. Then I asked what happens when the user sends a message in a tunnel and closes the app. He added a box called "Retry". That box was the whole interview.
I have run 500+ technical interviews from the hiring side over 15 years in mobile, and the iOS system design round is the one where strong engineers lose the most points for the least reason. The diagram is the entry fee. What gets scored is whether each box carries a decision you can defend.
This article is a worked answer. One prompt, the way I would want a senior candidate to walk it in 45 minutes. If you want the argument for why mobile design is its own discipline, that is in mobile system design is not backend system design. Here we just do the round.
- The prompt: design a messaging app on iOS that works offline.
- The shape: requirements, client architecture, local store, sync, pagination, images, push, background limits, observability, rollout.
- The test underneath: can you say what breaks, and what you chose to accept.
The 45-minute round, minute by minute
Most candidates lose time, not knowledge. They spend twenty minutes on a server they were not asked to design and run out of clock before sync. A plan you say out loud in the first minute also tells the interviewer you can run a meeting.
| Minutes | What you do | What the interviewer is checking |
|---|---|---|
| 0 to 5 | Clarify scope, users, scale, offline expectations | Do you ask before you build? |
| 5 to 10 | State requirements and what is out of scope | Can you cut? |
| 10 to 20 | Client architecture and local store | Is there one source of truth? |
| 20 to 32 | Sync engine, outbox, pagination | Do you know where data is lost? |
| 32 to 38 | Images, push, background execution | Do you know the platform's limits? |
| 38 to 43 | Observability and rollout | Have you shipped to a real fleet? |
| 43 to 45 | Trade-offs you accepted, what you would do next | Can you summarise a decision? |
The interviewer will interrupt this plan. That is fine. The plan is what lets you return to it.
Requirements: say what you will not build
I ask the candidate to design "a messaging app". A good first five minutes sounds like questions, not features. One-to-one or groups? How large can a group get? Text only or media? Must the user be able to read and send with no network? How long can a device stay offline and still be expected to sync cleanly?
Then the candidate writes the answer down as requirements:
- Functional: one-to-one and group chats, text and photos, read receipts, push for new messages.
- Offline: the last conversations are readable offline; sends made offline are delivered later, in order, exactly once.
- Non-functional: the chat list opens from local data with no spinner; scrolling history stays smooth on an older device; battery and data use stay modest.
- Out of scope: end-to-end encryption design, voice and video, the server's fan-out. I will describe the API contract I need, not the backend.
That last line earns points. Naming what you will not design is how you protect the next 40 minutes.
Client architecture: one source of truth
The rule I want to hear is short: the UI reads only from the local store. The network writes into the store. Nothing on screen comes straight from a response.
- Views in SwiftUI observe a view model per screen.
- View models read from repositories, which expose observations of the local database.
- Repositories write user intent to the store and to an outbox in one transaction.
- A sync engine, owned outside any screen, drains the outbox and applies server changes to the store.
- An API client knows HTTP, auth refresh and a WebSocket for live events. Nothing else.
Why this shape: the chat list opens instantly because it never waits on the network, offline is the same code path as online, and a server event and a local send land in the same place, so the screen cannot show two versions of the truth.
Local store: what goes in it and why
SQLite underneath, through Core Data, SwiftData or a thin SQLite wrapper. The choice matters less than the schema and the threading rule. I want to hear the tables: conversations, messages, attachments, outbox and a sync state row per conversation holding the server cursor.
Two details separate a senior answer. First, every message has a client-generated id from the moment it is typed, and a nullable server id. The UI keys cells on the client id, so a message does not jump or duplicate when the server confirms it. Second, writes go through one serial writer, and reads come from observations. That removes most of the threading bugs before they exist.
Eviction is part of the design, not an afterthought. Keep the last N conversations with recent history, drop old media first, and let older history page back in from the server.
Sync engine: the outbox and idempotency
This is where the round is decided. Sending a message is not a network call. It is a durable intent that must survive a killed app, a reboot and a dead network, and must reach the server exactly once.
struct OutboxOperation: Codable, Identifiable {
enum Kind: String, Codable { case sendMessage, markRead, deleteMessage }
enum State: String, Codable { case pending, inFlight, failed }
let id: UUID // also the idempotency key sent to the server
let kind: Kind
let conversationID: String
let payload: Data // encoded request body
let createdAt: Date
var attempts: Int = 0
var state: State = .pending
var nextAttemptAt: Date = .distantPast
}The walk-through I want to hear: the message row and the outbox row are written in one transaction, so there is never a message without an intent or an intent without a message. The engine drains operations in order per conversation, so replies do not overtake the message they reply to. Each request carries the operation id as an idempotency key.
Then the case that separates levels: the request times out. Did the server receive it? You cannot know. Without an idempotency key, retrying may send the message twice, and not retrying may lose it. With the key, the server stores the result against it, and a retry becomes a lookup that returns the original message. That is the mechanism that makes "exactly once" a claim you can defend.
func drain() async {
while let op = await store.nextReadyOperation() { // oldest pending, per conversation order
await store.mark(op.id, .inFlight)
do {
let result = try await api.send(op.payload, idempotencyKey: op.id)
await store.complete(op, with: result) // sets server id, deletes op, one transaction
} catch let error as APIError where error.isPermanent {
await store.mark(op.id, .failed) // surfaced to the user, not retried
} catch {
await store.scheduleRetry(op, after: backoff(op.attempts + 1))
}
}
}Three more points finish the section. Retries use exponential backoff with jitter so a fleet coming back online does not hit the server at the same second. Permanent errors, such as a blocked recipient, stop retrying and show a state the user can act on. On launch, anything left in inFlight is reset to pending, because a crash mid-request is a normal event.
Pagination: cursors, not pages
History loads backwards from the newest message. Offset pagination breaks here: new messages arrive at the top of the list while you page, and offsets shift, so you get gaps or duplicates. The answer is a cursor, an opaque token or the id and timestamp of the oldest message you hold, and the server returns the next page before it.
Pages are written into the store, not appended to an array in memory. The list observes the store. Prefetch the next page when the user is a screen or two from the end, not when they hit it. Catch-up after reconnect is a separate query: "everything since my last sync cursor", which is how the device learns about edits and deletions it missed.
Image pipeline
Photos are where memory and scroll performance go wrong. The answer I look for has four parts.
- Downsample to the display size with ImageIO before decoding. A full-resolution photo decoded into a small cell holds far more memory than the cell needs.
- Two caches: an in-memory cache for decoded images, bounded and cleared on memory warnings, and a disk cache for the downloaded bytes.
- Cancel the request when the cell is reused or scrolls away, and deduplicate requests for the same URL.
- Uploads go through a background URLSession from a file on disk, so they continue when the app is suspended. The outbox operation waits for the attachment id before sending the message.
Push and background execution limits
Candidates lose the most credibility here, because they describe iOS as if it were a server. The accurate version:
- A visible push carries the notification. A Notification Service Extension can modify it, for example to download a thumbnail, within a short time budget.
- A silent push with
content-availablecan wake the app to sync, but the system throttles it and may not deliver it at all. It is a hint, never the transport. - BGAppRefreshTask gives short, system-scheduled windows. BGProcessingTask suits longer work, typically when the device is idle and often charging. Neither runs on your schedule.
- Background URLSession is the one reliable way to keep a transfer going after suspension.
The design consequence: correctness never depends on background time. The source of truth is the catch-up query on launch and on foreground. Push and background refresh only make the data fresher when the system allows it.
Observability and rollout
A senior answer names the numbers it would watch, because the sync engine is invisible when it fails. Outbox depth per user, age of the oldest pending operation, send success on first attempt, duplicate sends caught by the idempotency key, time from app open to first rendered chat list, crash-free sessions. MetricKit gives hang and launch diagnostics from the field; the sync metrics you log yourself, without message content.
Rollout: the new sync engine ships behind a remote feature flag, off by default. Turn it on for internal users, then a small percentage, and watch the numbers above. App Store phased release limits how fast a binary reaches automatic updates; the flag limits how fast the behaviour does, and it can be turned off without a new build. Plan the store migration so the old code can still open the database if you roll the flag back.
How the Android answer differs
The architecture is the same: local store as truth, outbox, idempotency keys, cursors. Two platform answers change.
Deferred work. On Android, WorkManager runs the outbox drain as unique work with a network constraint, and it persists across reboots. That is a stronger guarantee than BGTaskScheduler, and interviewers expect you to use it. Process death. Android can kill the process while the user is on a screen and restore that screen later. UI state that matters goes into SavedStateHandle or the database, never only into a view model's memory. An Android answer that ignores process death reads the same way an iOS answer that trusts silent push does.
Mid, senior and staff: what the answers sound like
| Topic | Mid-level | Senior | Staff |
|---|---|---|---|
| Offline send | "Retry when the network returns." | Durable outbox, idempotency key, per-conversation order. | Defines the server contract for the key and who owns deduplication across teams. |
| Local store | "Cache responses in Core Data." | Store is the only source for the UI; client ids; one writer. | Schema migration and rollback plan tied to the flag. |
| Background | "Sync in the background." | Knows what iOS guarantees; correctness on foreground. | Sets a freshness budget and measures it. |
| Failure | Handles errors. | Separates retryable from permanent; resets in-flight on launch. | Names the accepted risk and the metric that would show it. |
The staff column is not more components. It is the same design, with owners, budgets and a way to know when it is wrong.
The follow-ups interviewers use to push
- "The request timed out. Did the message send?" Tests idempotency.
- "The user deletes a message that is still in the outbox." Tests whether operations can cancel each other.
- "The same account is on a phone and an iPad." Tests catch-up by cursor and read state across devices.
- "The database migration fails on launch for 1 percent of users." Tests rollback and whether the app can still start.
- "Silent pushes stop arriving." Tests whether correctness ever depended on them.
- "A group has thousands of members." Tests pagination of members and receipts, not just messages.
Prepare one sentence for each. If a follow-up makes you add a box, your design was missing a decision. If it makes you point at a box that already handles it, you are having the conversation the round is for.
Questions engineers ask about iOS system design interviews
What is asked in an iOS system design interview?
Usually one product: a feed, a chat, a photo app, an offline notes app. You are expected to drive requirements, draw the client architecture, choose a local store, design sync and pagination, and explain what happens when the network, the OS or the user does something you did not plan for.
How is an Android system design interview different?
The structure is the same. The platform answers change: WorkManager for deferred work instead of BGTaskScheduler, and process death as a normal event you design for, not an edge case. Interviewers listen for whether you know which guarantees your platform gives and which it does not.
How should I practise mobile system design?
Pick one product and answer it out loud against a 45-minute timer. Then change one constraint, such as offline for a day or two devices per user, and answer again. The second pass shows you where your design was a list of components instead of a set of decisions.