Mobile architecture
Android system design interview: a worked answer from the hiring side
On Android the round is decided by one question: what does the green check promise when the process dies a second later? A worked answer for an offline field inspection app, minute by minute, with the follow-ups interviewers use.
12 min read
The diagram usually arrives fast. Compose screens, a ViewModel per screen, a repository, Retrofit, Room. Then I ask one question: the technician taps Submit with no signal, sees the green check, and puts the tablet in a bag. Ten minutes later Android kills the process. What survived?
I have run 500+ technical interviews from the hiring side over 15 years in mobile. On Android, that question decides more system design rounds than any box on the diagram. A green check is a contract. The round is scored on whether you can say what it promises and prove the design keeps the promise.
This is a worked answer: one prompt, walked the way I would want a senior candidate to walk it in 45 minutes. The iOS version of the round, with an offline messaging app, is in the iOS system design worked answer. The argument for why mobile design is its own discipline is in mobile system design is not backend system design.
- The prompt: design an Android app for field technicians who fill inspection forms, take photos and collect signatures, often with no network for hours.
- The shape: requirements, the draft contract, process death, background work, upload lanes, conflicts, schema changes, low storage, observability, rollout.
- The test underneath: can you name what the user was promised, and what breaks it.
The 45-minute round, minute by minute
Strong candidates lose this round on the clock, not on knowledge. Twenty minutes on a server nobody asked for, and there is no time left for background work, which is where the Android points are.
| Minutes | What you do | What the interviewer is checking |
|---|---|---|
| 0 to 5 | Clarify users, offline duration, attachments, shared records | Do you ask before you build? |
| 5 to 10 | Requirements, and what you will not design | Can you cut? |
| 10 to 18 | The draft contract and the local store | What does "saved" mean here? |
| 18 to 28 | Process death and background work | Do you know what Android guarantees? |
| 28 to 35 | Upload lanes and conflicts | Do you know where data is lost? |
| 35 to 41 | Schema changes, low storage, observability | Have you run a real fleet? |
| 41 to 45 | Rollout and the trade-offs you accepted | Can you summarise a decision? |
Say the plan out loud in the first minute. The interviewer will break it. The plan is how you come back.
Requirements: say what you will not build
A good first five minutes is questions. How long is a technician offline: minutes in a lift, or a full day on a rural site? How large are the photos? Can two technicians touch the same asset on the same day? Are the devices company tablets on old Android versions? Is any field safety-relevant, where a wrong value matters to a regulator?
Then the candidate writes the answers down:
- Functional: download the day's assignments, fill versioned forms, attach photos and a signature, submit, see what still needs to sync.
- Offline: a technician can complete and submit a whole inspection with no network; everything reaches the server later, once, in order.
- Non-functional: no submit ever waits behind photo uploads; the app opens and works on an old tablet with a nearly full disk; battery lasts a shift.
- Out of scope: the back office, the server's storage design, authentication beyond token refresh. I will state the API contract I need, not build the backend.
Naming what you will not design is not modesty. It is how you protect the next 40 minutes.
The draft contract: durable before acknowledged
The sentence I want to hear before any library name: the UI may show success only after the edit is committed to disk. Not after it reaches a ViewModel. Not after an autosave timer fires. After the database transaction commits.
That turns the draft into a small state machine, and saying the states out loud is the strongest five minutes of the round:
| State | What it means | Who moves it on |
|---|---|---|
| capturing | Edits in memory, on screen | A committed local transaction |
| durable | On disk, survives process death. The earliest state the green check may show | The sync engine, when it enrols the draft |
| queued | In an upload lane, attachments addressed by content hash | A worker |
| submitted | Form data accepted; photos may still be uploading | The server's acknowledgement |
| acknowledged | The server returned the record's new version | Nobody; this is the end state |
The write that matters is one transaction: the draft row and an operation in the outbox, together. Either both survive a kill or neither does.
@Dao
interface DraftDao {
@Transaction
suspend fun submit(draft: DraftEntity, op: OutboxOperation) {
upsertDraft(draft.copy(state = DraftState.DURABLE))
insertOperation(op) // op.id is a client UUID, sent later as the idempotency key
}
@Upsert suspend fun upsertDraft(draft: DraftEntity)
@Insert suspend fun insertOperation(op: OutboxOperation)
}Then state the invariant so a test can assert it: after any interruption, the newest state the UI ever showed is at most the newest state committed to disk. A design you can test that way is a design. "We autosave every few seconds" is a hope.
Process death: the ViewModel is not storage
Android can kill the process while the app is in the background and later restore the screen the user left. The ViewModel does not survive that. Anything that exists only in its memory is gone.
The answer I listen for splits state by what it costs to lose:
- Data the user produced, form fields, photos, the signature: Room, written as the user goes, not on Submit.
- Small UI state, which tab, which section is open, the scroll position: SavedStateHandle.
- Everything else is rebuilt from those two on restore.
A candidate who says "the ViewModel survives rotation, so we are fine" has answered the first question and walked into the follow-up. Rotation is not the hard case. Process death is.
Background work: what WorkManager guarantees, and what it does not
WorkManager is the right tool, and saying its name earns nothing. What earns points is knowing the guarantee precisely. Its queue lives in its own database, so an enqueued request survives process death and a reboot, and a worker that returns retry runs again with backoff. The minimum backoff is ten seconds.
What it does not guarantee: that your work remembers where it stopped. A running worker can be killed with no callback at all. WorkManager will run it again. Resuming at the right place is your design, built from the durable state in the outbox.
fun scheduleSync(context: Context) {
val request = OneTimeWorkRequestBuilder<OutboxWorker>()
.setConstraints(
Constraints.Builder()
.setRequiredNetworkType(NetworkType.CONNECTED)
.build()
)
.setBackoffCriteria(BackoffPolicy.EXPONENTIAL, 30, TimeUnit.SECONDS)
.build()
WorkManager.getInstance(context).enqueueUniqueWork(
"outbox-sync",
ExistingWorkPolicy.KEEP, // one drain at a time, however many triggers fire
request
)
}Two more points finish the section. A request is a hint, not a schedule: the system decides when deferred work runs, and Doze batches it into maintenance windows that can be hours apart. And a token minted when the work was enqueued may be expired when it runs, so the worker refreshes credentials at run time. A 401 in the queue means "re-authenticate and replay", never "discard".
Upload lanes: a submit never waits behind a photo
Durable capture creates the next problem. A technician who took forty photos in a dead zone resurfaces with hundreds of megabytes queued, and the one submit the office needs today sits behind them.
The answer is lanes with independent scheduling:
- Critical metadata: submits and status changes. Small, first, always.
- Form bodies.
- Attachments: chunked, resumable, and preemptible, so they yield the network to anything waiting in lane 1.
Each chunk is addressed by the hash of its content. A retried chunk is then idempotent by construction, and the server can deduplicate a photo it already has. The user sees the inspection as submitted while the photos finish in the background, and the screen shows which photos are still pending.
Conflicts on shared records
Two technicians inspect the same asset on the same day, from two dead zones, and submit different values hours apart. With last write wins, the later submit silently replaces the earlier one.
The senior answer separates detection from resolution. Detection is mechanical: the server keeps a version per record, and a submit whose base version is not the current one is a conflict. Resolution is a policy per field. Free-text notes can merge or take the newer version. Safety fields, a pass or fail, a defect severity, never resolve automatically: the record goes to review, both values are kept with author and time, and a supervisor decides.
One trap catches good candidates here. Ordering by the device clock looks reasonable and is wrong: a tablet whose clock runs ahead lets a stale edit beat a correct later one. Order by the server's version, never by device time.
Schema changes and the oldest tablet
Forms change when regulations change. Drafts started under last month's form must stay editable, and some of them are on tablets that have been offline for weeks.
The answer: every draft records the form version it was created under. The app renders any version it has seen. A draft upgrades lazily, when the technician next opens it, through a migration function that is small and testable. Each old version gets a published end date tied to how many drafts still use it.
The instinct to reject out loud is a server-side backfill that rewrites every stored draft at once. It cannot reach drafts that are still offline, and it runs migration code against data nobody has looked at. Lazy migration keeps the damage of a bug to one draft.
Low storage: the failure that stops the app opening
Field fleets keep old tablets with small disks. A duplicated attachment cache or a large update can leave so little space that the app cannot write its own database, and then it cannot start.
The design has two parts. On launch, check free space before writing to the durable log, and degrade to a mode that can still finish and submit what is pending. When reclaiming space, delete only what can be proven redundant, for example a duplicate attachment whose content hash matches an original, and drain pending operations before deleting anything.
The answer that fails this section is the one that sounds decisive: "wipe local data and resync". On a device with unsynced inspections, that turns a disk problem into permanent data loss.
Observability and rollout
The sync engine is invisible when it fails, so a senior answer names its numbers. Completion after interruption: the share of drafts that reach acknowledged after at least one process kill. Queue age: how long the oldest pending operation has waited, at p90. Time from submit to acknowledged, which should not grow with photo backlog. Conflicts per thousand submits, by field type. Crash-free sessions and storage-guard activations, segmented by app version and device model.
Rollout: the new engine ships behind a remote flag, off by default. Turn it on for an internal crew, then a small share of the fleet, and watch the numbers above. A staged rollout on Google Play limits how fast the binary spreads; the flag limits how fast the behaviour does, and it can be switched off without a new build. Assume several app versions stay active for months, because some devices update only in a quarterly maintenance window.
How the iOS answer differs
The architecture carries over: durable outbox, idempotency keys, server versions, cursors. Two platform answers change. iOS gives you no persistent job queue like WorkManager: BGTaskScheduler requests are hints with short or opportunistic windows, and a background URLSession is the reliable way to keep a large transfer going after suspension. The "what survived?" question stays, with different vocabulary: the local store plus scene state restoration instead of Room plus SavedStateHandle. The iOS worked answer walks that version.
Mid, senior and staff: what the answers sound like
| Topic | Mid-level | Senior | Staff |
|---|---|---|---|
| Submit offline | "Queue it and retry." | Durable before acknowledged; outbox in the same transaction; idempotency key. | Defines what the green check promises and the test that proves it after a kill. |
| Background work | "Use WorkManager." | Unique work, constraints, resume from durable state, refresh tokens at run time. | Classifies work by urgency and owner; measures completion after interruption. |
| Conflicts | Last write wins. | Server versions; policy per field; review for safety fields. | Names who ratifies, what the losing technician sees, and the conflict budget. |
| Migrations | "Migrate the database." | Versioned forms, lazy migration on open. | Funds the old-version tail, owns end dates, sets the number at which to reverse. |
The staff column is not more components. It is the same design with owners, budgets and a way to know when it is wrong.
The follow-ups interviewers use to push
- "The process dies one second after the green check. What survived?" Tests the draft contract.
- "The worker is killed halfway through a 20 MB upload." Tests resumable chunks and durable progress.
- "The token expired while the work waited overnight." Tests whether a 401 discards data.
- "Two technicians edited the same asset offline." Tests detection, policy and who decides.
- "A new form version ships while drafts are offline." Tests lazy migration and old versions.
- "The disk is full and the app will not start." Tests recovery without data loss.
Prepare one sentence for each. If a follow-up makes you draw a new box, your design was missing a decision. If it makes you point at a box that already handles it, you are having the conversation the round is for.
Questions engineers ask about Android system design interviews
What is asked in an Android system design interview?
One product, designed as a client: a feed, a chat, a delivery or field app, a media downloader. You are expected to drive requirements, choose a local store, design sync and background work, and explain what happens when the network, the OS or the user does something you did not plan for. On Android, process death and deferred work are where most of the points are.
How is it different from an iOS system design interview?
The structure is the same: local store as the source of truth, an outbox, idempotency keys, cursors. The platform answers change. WorkManager persists queued work across reboots, which iOS does not offer in the same form, and process death is a normal event you design for. The iOS version of the round is worked through in a separate answer.
Do I need to know WorkManager APIs by heart?
No. You need to know what it guarantees and what it does not. It guarantees that enqueued work runs again after a kill or a reboot. It does not guarantee that your work remembers where it stopped. That second half is your design, and it is what the interviewer is listening for.
How should I practise Android system design?
Pick one product and answer it out loud against a 45-minute timer. Then change one fact, such as the device being offline for two days or the disk being nearly full, and answer again. The second pass shows which parts of your design were decisions and which were boxes.
