Skip to content
All essays

Mobile architecture

How to choose a sync solution for a mobile app

Most sync decisions start with a comparison table of products. The table is the last step. Before it come three decisions no product makes for you: what the app promises offline, who owns the truth for each kind of data, and which conflicts a machine may settle. A vendor-neutral guide for iOS and Android teams, with a build-or-adopt rule.

13 min read

The question usually arrives as a spreadsheet: four sync products in the columns, features in the rows, a tick in most cells. Someone has to pick one by Friday. The spreadsheet is rarely wrong about the features. It is silent about the only question that decides whether the choice survives production: when two copies of the same data disagree, who is allowed to decide?

I have run 500+ technical interviews from the hiring side over 15 years in mobile, and sync is where design reviews and design rounds go wrong in the same way. A product name gets offered as an answer. A sync solution is not a library choice. It is a decision about who owns the truth for each kind of data, and the library only carries that decision out.

This is the decision guide. If you are preparing for the interview round instead, the worked answer is the offline sync system design interview, and the reference page is mobile data synchronization design. The cost model below comes from my book Controlled Change.

  • The decision: adopt a sync solution, build a narrow one, or use both, split by data kind.
  • The tools: the offline promise per feature, authority per fact, a conflict policy per kind of data, and eight questions for any candidate.
  • The exit: what leaving costs, and how to test a candidate the way production will.

Start with the promise, not the product

Controlled Change opens its offline chapter with one line I would put above every sync evaluation: "Offline-first architecture is a contract about continuity and truth, not a cache strategy." The contract differs per feature, so the book splits "supports offline" into five levels:

LevelWhat the user getsWhat it requires
Graceful failureA clear message that a connection is needed; navigation keptNothing to sync
Read continuityEarlier data stays readable, with visible freshnessA durable local copy and freshness rules
Draft continuityWork the user typed survives, not yet submittedUser-owned local state
Deferred intentThe app accepts an action now and sends it laterA durable operation ledger and idempotency
Local authorityThe device settles a bounded decision on its ownExplicit policy and later reconciliation

The levels are not a ladder to climb. Each feature should stop at the lowest level that keeps its promise. Browsing saved items needs read continuity. Booking scarce inventory offline should at most save a request, and never claim a reservation.

Do this table before you open a vendor page, because it changes the shopping list. A feature set that stops at levels one to three needs a cache with freshness and a pull cursor, not a sync product. Only deferred intent and local authority need the machinery the spreadsheet is comparing.

Name who owns the truth, per fact

Authority is not global. The book's chapter on data authority lists it per fact for its fictional reference system, Atlas: the backend owns confirmed bookings and payments, the device's ledger owns whether it accepted an unsent action, the user owns a draft until it is submitted, and the server owns inventory, with the local copy as a timestamped projection.

Kind of dataWho owns the truthWhat the device may do
Drafts, notes, personal listsThe userEdit freely; sync to the user's other devices
Shared lists, tags, membership-like setsShared, merge-shapedPropose changes that merge by a known rule
Bookings, inventory, ordersThe serverSave a request; show it as pending until a receipt arrives
Payments, permissions, entitlementsThe serverRead a projection; never decide

This table is the filter for every candidate. A solution built around replicating a database to the device is a natural fit for the first two rows. For the last two, the device is not a replica of the truth. It is a client that sends commands the server may reject, and a solution that cannot express "the server said no, here is why" will push you to fake it.

Keep one rule from the book in mind when a product page says the local database is your source of truth. That is a sound rule for the UI, which should read from one local model. It does not make the device the authority for money or access. The local database is a projection, and it should carry its freshness and its provenance.

Three problems that all get called conflict

Comparison tables have one row called "conflict resolution". Controlled Change separates three problems that can all surface as the same HTTP 409, and each needs a different fix:

  • Duplicate delivery: the same intended action arrives twice, because a response was lost and the client retried. The fix is a stable idempotency key, created once and reused on every replay, and a server ledger keyed by account, operation kind and key.
  • Concurrent valid intent: two different actions touch the same state. The fix is a domain conflict policy. Deduplication does nothing here.
  • Version incompatibility: client and server disagree about the shape or meaning of a command. The fix is contract evolution, capability checks or an explicit rejection.

Ask any candidate how it handles each of the three, separately. A solution that answers all three with one merge strategy has answered one of them.

Choose the conflict model before the solution

The book's table of policies is the most useful page to bring to an evaluation, because it names both where each policy fits and where it does damage:

PolicyFitsDangerous for
Last writer winsLow-value preferences where a lost edit is acceptableMoney, inventory, safety evidence
Field-level mergeIndependent profile or draft fieldsFields whose meaning is coupled
Set union, observed removeTags, membership-like collectionsOrdered lists, exclusive choices
Optimistic concurrencyInventory, booking, controlled editsFlows that cannot explain a rejection
Server arbitrationPricing, authorization, fraud, policyUser-owned drafts the server should not discard
Human resolutionHigh-value semantic disagreementHigh-frequency trivial edits

Then the line that should end most arguments about algorithms: "A merge algorithm is a product rule." A CRDT, or any automatic merge, decides how two edits combine. It does not decide whether removing a traveler should beat another device adding a document for that traveler. Your product decides that.

The book's casebook shows the cost of skipping this step. Atlas built one engine with pluggable merge strategies and ran every synced entity through it, bookings included. A traveler modified a reservation offline while the host changed availability; the engine merged locally and showed a confirmed change, and the server later rejected it. Atlas moved bookings out of the engine, to commands with server authority and a visible reconciliation state. The principle it wrote down: "Conflict policy belongs to the domain invariant, not to the transport."

For a sync evaluation, that means one thing. A single solution for all your data is only right if all your data has the same invariant. Most apps do not, and that is fine: adopt or build for the merge-shaped data, and keep server-owned data on commands.

Eight questions to ask any sync solution, including your own

Put the same questions to every candidate, and to the version you would build yourself. A build is not exempt because its weaknesses are your own.

QuestionWhat a good answer showsWhere a weak answer hurts
1. Who resolves conflicts, and where?Resolution per entity or field; the server can reject; your policy is code you ownServer-owned data merged on the device
2. Can each device get only its part?Sync scoped by account and permission; revoked data leaves the deviceFull copies on every device; access kept after revocation
3. What happens to queued work when the schema changes?Queued operations keep a payload version and are migrated, replayed or quarantinedAn update silently drops or reinterprets what the user asked
4. Who checks write permission per record?The server, on every write, whatever the client versionRules enforced only in the app
5. What does leaving cost?Useful export, your own identifiers, the ability to run in parallelVendor IDs as primary keys across mobile and backend
6. Can you see it in production?Oldest pending age, conflict rate, cursor lag, without payloads"Syncing" as the only state, and support reading user data
7. How do old app versions behave?Unknown fields and values tolerated; old operations accepted or clearly rejectedA server change that strands last year's installs
8. Where does the data live?Known regions for storage and processing; deletion you can evidenceCopies in places your compliance team never approved

Questions 3 and 7 are where mobile differs from every other client. The book puts it in one line: "Compatibility is not a courtesy. It is a production topology." An operation created by an old app version may be sent by a newer one, after an update, to a server that has changed twice since. A sync solution that only works when every clock moves together has not met a real install base.

Price the exit before the entrance

Controlled Change lists the synchronization platform among the high-risk candidates for replacement, next to payment, identity and the persistence engine. Its vendor checklist is short and worth running before adoption. It includes: can data be exported in a useful form, are identifiers portable, what happens if pricing changes tenfold, can the capability be disabled per region or cohort, can it run in parallel with an alternative, and who owns migration and contract termination.

The book does not argue against lock-in: "Lock-in can be rational; a provider may deliver unique value. Make the choice explicit and price the exit." Pricing it means two habits. Store your own identifiers and map the provider's to them. And let features talk to a narrow port in your product's language, never to the provider's types.

Swift
// What features see. No vendor type crosses this line.
struct OperationID: Hashable, Sendable { let raw: String }

enum OperationState: Sendable {
    case savedOnDevice
    case waitingToSend
    case sending
    case confirmationDelayed
    case completed(serverVersion: Int)
    case needsAttention(reason: String)
}

protocol NotesSync: Sendable {
    func submit(_ edit: NoteEdit) async throws -> OperationID
    func state(of operation: OperationID) async -> OperationState
    func changes() -> AsyncStream<NoteChange>
}

Keep the port small, and keep it about what notes need. The book's warning applies to the opposite move: "A generic vendor-shaped interface merely moves lock-in one file inward." The states come from the book's chapter on offline-first, which separates saved on this device, waiting to send, sending, confirmation delayed, completed and needs attention, instead of one word, syncing, for all of them. The same port works on Android as a Kotlin interface with a Flow.

What building your own means

Teams that decide to build usually underestimate it, because they picture a retry loop. The book's chapter on the subject opens by correcting that: "A synchronization engine is not a loop that retries failed requests." It is a small distributed system on a device that can be killed, wake with expired credentials and meet server versions it has never seen.

The book splits the engine into seven jobs: an operation ledger for durable intent, a claim coordinator that leases work to one worker at a time, a transport that sends versioned and idempotent requests, a resolution policy that classifies each outcome, a pull synchronizer that advances a cursor, a projection writer that updates the local model in a transaction, and an observer surface that reports status per operation.

You do not need all seven on day one. The book's migration path starts with one idempotent mutation and one operation table, then adds explicit states, lookup of unknown outcomes and server receipts, leases only when concurrent workers exist, and shared mechanics only after two domains prove the common model. The thing to avoid is the reverse: each feature team inventing its own outbox, each failing in its own way.

One fact applies to both paths. In the book's words, exactly-once delivery across a network boundary is usually an illusion, whichever engine you pick. Design for at-least-once delivery and idempotent handling on the server, and the server work is the same in both cases.

A build-or-adopt rule

This is my rule of thumb, built on the decisions above. It is a starting position for the discussion, not a verdict.

Your data and promiseDefaultWhy
Read continuity and local drafts onlyNo sync product: a cache with freshness and a pull cursorNothing to reconcile
A few offline writes on server-owned dataBuild narrow: commands, idempotency keys, server receipts on your APIThe server must arbitrate anyway
Many user-owned, merge-shaped entities across devicesAdopt, if it passes the eight questions and the exit is pricedMerge, partial sync and multi-device are the hard part a solution can carry
Both kindsBoth, split by data kindNever route server-owned data through a client-side merge

Two signals should reopen the decision later. Conflict frequency is one: the book treats a domain that keeps producing human conflicts as evidence that the state model or the ownership boundary is wrong. The other is a change in who owns the truth, which a new feature can cause without anyone noticing.

Test the finalists the way production will

A proof of concept against a healthy local server tells you the API is pleasant. The book is blunt about what it does not tell you: "A sync engine tested only against a healthy local server is untested." Run each finalist through the book's failure list, and write down the expected user-visible state for each, because "no crash" is not enough:

  • Kill the app right after the local commit.
  • Lose the network after the server commits but before the response arrives.
  • Switch accounts while operations are pending.
  • Expire the credentials while work is queued.
  • Send the same request from two workers.
  • Run an old client against a new server rule.
  • Revoke an item while a stale local copy remains.
  • Reconnect after a week with hundreds of remote changes.

Mobile platforms make one more test mandatory. iOS and Android both decide when background work runs, and both can stop it. The book's rule: persist intent before you ask for background time, make workers restartable, and let a foreground launch resume pending work. A candidate whose correctness depends on a background task finishing will lose work the first time the system defers or stops it.

A one-page record for the decision

Before you sign anything or create the first table, write six lines. They fit in an ADR or an architecture proposal review:

  • The offline level each feature needs, and the ones deliberately left out.
  • Who owns the truth for each kind of data.
  • The conflict policy per kind of data, and which conflicts reach a person.
  • The answers to the eight questions, for the chosen option and the runner-up.
  • The exit: export, identifiers, and who owns a migration.
  • The signal that reopens the decision: conflict rate, a new server-owned domain, or a pricing change.

If you cannot fill in the second line, no product will fill it in for you. If you cannot fill in the fifth, you have not chosen a solution yet, only a dependency. For the interview version of the same ground, with cursors, tombstones and resync, see the worked offline sync answer; for the product side, offline-first mobile architecture.

Questions engineers ask about choosing a sync solution

How do you choose a data sync solution for a mobile app?

Decide what the app promises offline, feature by feature, and who owns the truth for each kind of data, before you compare products. Then choose a conflict policy per kind of data, and put the same eight questions to every candidate, your own build included: conflicts, partial sync, queued work across schema changes, authorization, exit cost, observability, old app versions and data residency. Test the finalists against failures, not a healthy local server.

Should you build your own sync engine or adopt one?

Build a narrow one when the server owns the truth for the data that matters, such as bookings, payments or permissions, and you only need a few offline writes: an operation table, idempotency keys and server receipts on your existing API. Adopting a solution makes more sense when most synced data is user-owned and merge-shaped, such as notes, lists and drafts, across several devices. Many apps end up with both, split by data kind.

What does a mobile data synchronization framework actually do?

It stores local intent durably, sends it with retries that are safe to repeat, pulls remote changes from a cursor, applies a conflict policy, writes the local read model in a transaction, and reports the state of each operation. Each of those is a separate job, and a framework is only as good as the weakest one for your data.

Is a local database with sync enough to make an app offline-first?

No. Offline-first is a promise to the user about what still works, what has been accepted and what is still uncertain. A sync layer is one part of keeping that promise. The interface states, the authority rules and the recovery paths are the rest, and no library ships them for your domain.

Share this essay