Mobile architecture
Mobile architecture migration interview: change the app without stopping the release train
"How would you migrate this app's architecture without stopping the release train?" is a question about the months between two diagrams, not about the destination. A worked answer from the hiring side, for iOS and Android alike.
12 min read
The question usually comes late in a staff loop, once a target design is on the board. This app has years of history, a weekly release train and several teams shipping into it every week. How would you get it from what exists to what you just drew, without stopping the train?
I have run 500+ technical interviews from the hiring side over 15 years in mobile. The weak answers to this one share a shape: they describe the destination again, in more detail, or they ask for a feature freeze, or they propose a rewrite in a new framework. All three skip the part the question is about. The destination is the easy half. The question is about the months in between.
Controlled Change, my book on mobile architecture, says it without softening: "A migration that requires stopping the release train to complete has already failed; it just has not admitted it yet." This answer follows its method. It is platform-neutral; for how a reviewer reads the same kind of plan, see how to review a mobile architecture proposal.
- The prompt: move a mature iOS and Android app to a new architecture while regular releases continue.
- The shape: pressure, characterization, seam, journey slices, shadowing, data, rollout, migration states, roadmap, stop conditions.
- The test underneath: can old and new run in the same release, and can you prove when the old one is gone.
The 45-minute round, minute by minute
| Minutes | What you do | What the interviewer is checking |
|---|---|---|
| 0 to 5 | Ask why migrate, what hurts today, who ships into the app | Do you ask for the pressure before proposing a target? |
| 5 to 12 | State the pressure as a measurable outcome | Is this a problem or a taste? |
| 12 to 20 | Characterize current behavior, find the seam | Do you know what the old code protects? |
| 20 to 30 | Branch by abstraction, slice by journey, shadow | Can old and new coexist in one binary? |
| 30 to 38 | Data, queued work, rollout and rollback | What can you not take back? |
| 38 to 45 | Migration states, owners, stop conditions, deletion | Does this migration ever end? |
Candidates who spend twenty minutes on the target architecture often run out of time before the coexistence plan, and the coexistence plan is the part being scored.
Name the pressure, not the destination
"Move to Clean Architecture" and "modernize the app" are not problems. A problem is something a product owner would recognize: a booking change takes six weeks because four teams touch it, payment incidents cannot be traced across client and backend, a shared state holder breaks two features every release. Controlled Change puts the difference in one line: "Reduce booking-change lead time from six weeks to two while maintaining the duplicate-reservation invariant" is actionable, "modernize" is not.
The outcome does two jobs in the interview. It tells you which journey to migrate first, the one where the cost is paid today. And it tells you when to stop, because a migration with no outcome has no finish line.
If the interviewer pushes toward a rewrite, answer with the conditions under which you would accept one: the platform is unsupported, the runtime blocks a capability the product needs, or incremental change is demonstrably more expensive. A rewrite also hides costs: years of undocumented behavior, accessibility and localization edge cases, analytics semantics, offline data, support procedures, security controls added after incidents, and parity criteria that keep growing while the product keeps shipping.
Characterize before you change
A mature codebase is a record of product bets, deadlines, platform transitions and incidents. Some of its strangest code is the only thing protecting a client version that is still installed, or an SDK that breaks if initialized twice. Before moving anything, capture what the current journey actually does:
- state transition tests around the known paths, and request and response fixtures;
- database snapshots and the migrations they went through;
- screenshots of key states at large accessibility text sizes, and the analytics event contracts;
- performance baselines, field error and completion rates, and the feature-flag combinations still live;
- the edge cases support already knows about.
The book is precise about what these tests mean: "Characterization tests are not an endorsement." They record behavior so you can decide, case by case, what is intentional, what is accidental and what is obsolete. Without them, a cleaner implementation can quietly remove a fix that took an incident to learn.
Find a seam, then branch by abstraction
A seam is a place where you can observe or substitute behavior without rewriting the system around it: a repository facade, a navigation entry point, a network interceptor, a feature factory, a feature flag. Pick the smallest seam that enables the next step. Starting with a universal dependency injection container and fifty protocols is a migration of its own, with no user benefit.
Then branch by abstraction. A stable facade first delegates to the old code. The new implementation fulfils the same product contract behind it, and a flag or cohort policy decides which one runs. The facade speaks product language, submit this booking, not the old class names: a facade that mirrors legacy classes keeps the old architecture alive in a new file.
// The stable product contract. Callers never learn which implementation ran.
interface CheckoutFlow {
suspend fun submit(intent: BookingIntent): SubmitOutcome
}
enum class CheckoutPath { LEGACY, NEW }
class CheckoutFacade(
private val legacy: CheckoutFlow,
private val rebuilt: CheckoutFlow,
private val record: (CheckoutPath, SubmitOutcome) -> Unit,
) {
// The path is read once, when the journey starts, and kept for the whole
// journey: a config refresh cannot switch implementations mid-checkout.
fun startJourney(path: CheckoutPath): CheckoutFlow {
val chosen = if (path == CheckoutPath.NEW) rebuilt else legacy
return object : CheckoutFlow {
override suspend fun submit(intent: BookingIntent): SubmitOutcome =
chosen.submit(intent).also { record(path, it) }
}
}
}Two details score here. The path is chosen once per journey, so a background configuration refresh cannot move a user from the old checkout to the new one halfway through a payment. And every outcome is recorded with the path that produced it, so the two implementations can be compared in the field, not argued about in a meeting.
Slice by journey, never by layer
The most common plan that fails is horizontal: migrate the whole data layer, then the whole domain layer, then the UI. Months pass, two systems run side by side everywhere, and no user journey is better yet. Progress is impossible to show and easy to cancel.
Migrate vertical slices instead: one deep link entry, one booking type, one message composition path, one local projection, one API endpoint behind an adapter. Each slice delivers something on its own and leaves the old path working for everything else. Start with the costly critical journey from your outcome statement, release it to a controlled cohort with rollback, and expand slice by slice.
Shadow the logic, never the side effects
For deterministic logic, run the old and new implementations side by side and compare results without showing the new one to the user. Price formatting, route parsing, validation, conflict classification and local projection building are good candidates.
For anything with side effects, shadowing must not duplicate writes. Compare the commands each path would send, serialize requests without sending them, or replay recorded fixtures in a controlled environment. Two implementations that both charge the card are not a comparison, they are an incident.
Then read the comparison carefully. Mismatches need a stable comparison key and privacy-safe categories, and the rate alone is not enough: a 99.9 percent match can still hide a catastrophic mismatch in the most valuable case. Say which case you would check by hand.
The data you cannot take back
Code can be routed back to the old path. Data written by the new path often cannot be read by the old one. This is where an interviewer probes hardest, because it decides whether rollback is real.
- One writer if you can. Prefer one authoritative writer and an asynchronous projection over dual writes.
- If you must dual-write, define which store wins, order the writes, record partial failure, make retries idempotent, reconcile continuously, cap how long both exist, and test rollback in both directions.
- Queued work is a historical contract. An offline operation created by last year's version may be sent by this year's version to a newer backend. Keep a decoder for its payload version, migrate it only when its meaning is preserved, or quarantine it with a path to recovery.
- Expand, then contract. For a breaking contract change: the server accepts both forms, clients understand the new form, new behavior activates for capable clients, you measure old traffic, and only then does anything get removed.
Above all, preserve what the user wrote and what is still pending. A local database migration that falls back to a reset is acceptable for a cache and unacceptable for unsent work.
The rollout path is part of the design
On mobile, the binary is installed and stays installed. That is why the book treats the release sequence as architecture: "The code path and the rollout path are one design." A sequence a staff candidate can say in one breath:
- deploy backward-compatible backend support;
- ship the new client capability dormant, behind the facade;
- validate the schema migration and the telemetry;
- enable for employees, then widen by cohort and platform version;
- watch business correctness, not only crashes;
- keep a server-side way to turn it off;
- remove the old path only after the compatibility window closes.
Then say what rollback means, because it means several things after a binary ships: turn off the flag, route back to the old implementation, change a server response or capability, block an unsafe operation on the server, ship a patched binary, and only as a last resort raise the minimum version. Test them before you need them. A kill switch that depends on the broken code path is not containment. Write the halt conditions before exposure, not during the incident.
Give every journey a state and an owner
A migration across many teams fails quietly when nobody can say where each journey stands. The book tracks every consumer or journey through the same states, each with entry criteria, a rollback, an owner and evidence:
enum MigrationState: Sendable {
case notStarted, compatible, shadowing, partialRollout
case defaultNew, oldDisabled, oldDeleted
}
struct JourneyMigration: Sendable {
let journey: String // "checkout", not "the data layer"
let owner: String // one accountable team
var state: MigrationState
let rollbackTarget: MigrationState
let haltConditions: [String] // written before exposure, not after
// "Default new" is not done. Done is when the old path is gone.
var isComplete: Bool { state == .oldDeleted }
}The line that separates levels is the last one. Making the new path the default is not completion: the old code, schemas, flags, dashboards, documentation and support paths still have to be removed. A migration declared done at default-new leaves two systems to maintain, test and debug during every incident.
Keep the roadmap paying for itself
A migration that eats the whole roadmap gets cancelled or bypassed. Structure the work so each wave delivers a product capability, retires a measurable risk or makes delivery faster. A useful mix: enabling work (the facade, fixtures, observability), vertical slices where new product work is built on the new path, risk retirement for crash-prone or insecure old behavior, deletion, and platform improvements that make the next slice cheaper.
Negotiate capacity with product leadership in their terms, release speed, reliability, support cost, regulatory need, not "cleaner architecture". And reserve capacity for deletion after each wave, because deletion is the work nobody schedules otherwise.
The organizational half, when every team agrees the migration is right and nobody wants to pay for it first, is a chapter in Before the Room, my field guide to leading outcomes without authority: a wave plan with funded first movers, a support owner and a decommission date.
Say when you would stop
Staff candidates name the stop conditions before anyone asks. Pause the migration when field quality regresses outside budget, when the new architecture needs more coordination than expected, when product assumptions change, when the bridge becomes the dominant complexity, when tooling support is not there, or when the business case no longer pays the cost.
Stopping is not failure if what remains is owned and contained. A bounded legacy area with an owner and no new callers is a legitimate end state. A migration that can never stop is the one to worry about.
Mid, senior and staff: what the answers sound like
| Topic | Mid-level | Senior | Staff |
|---|---|---|---|
| Why | "The code is messy." | Names the journeys that hurt. | States a measurable outcome and the invariant it must keep. |
| How | "Rewrite it, then switch." | Facade, flag, one journey at a time. | Branch by abstraction, path fixed per journey, shadowing without side effects. |
| Data | "Migrate the database." | Migrations with fixtures. | One writer, queued operations kept decodable, expand then contract. |
| Release | "Ship it behind a flag." | Staged rollout with a kill switch. | Halt conditions written before exposure, several rollback mechanisms tested. |
| Done | "When it's on for everyone." | When the new path is default. | When the old path, flags and bridges are deleted, with evidence. |
The staff column is not a bigger diagram. It is the same migration with outcomes, owners, evidence and an end.
The follow-ups interviewers use to push
- "Product wants a feature in this journey next month." Tests whether new work lands on the new path behind the facade.
- "The new path has a bug at 10 percent rollout." Tests which rollback mechanism you use and how fast it reaches devices.
- "Offline operations were queued by last year's version." Tests payload versioning and quarantine.
- "Leadership asks why this takes three quarters." Tests whether each wave delivers value on its own.
- "The bridge code is now bigger than the feature." Tests stop conditions.
- "How do you know the old path is really gone?" Tests deletion evidence: no imports, no runtime traffic, no config keys.
Prepare one sentence for each. If a follow-up forces you to invent a new mechanism, the plan was missing a decision. If you can point at a part of the plan that already handles it, you are having the conversation the round is for. For the incident side of the same skill, see the production incident interview question.
Questions engineers ask about migration interviews
How do you answer "migrate the architecture without stopping releases" in a staff interview?
Name the pressure as a measurable outcome, characterize the current behavior, put one product-shaped seam in front of one critical journey, run old and new side by side behind a flag chosen once per journey, shadow the logic but never the side effects, keep one writer for the data, and give every journey a migration state, an owner, halt conditions and a deletion date.
When is a rewrite the right answer?
When the platform is unsupported, when the runtime blocks a capability the product needs, or when you can show that incremental change costs more. Discomfort with old code is not on that list. Even a justified rewrite still has to coexist with the installed app versions, so the coexistence plan does not go away.
Do I have to migrate every screen to the new architecture?
No. A migration can stop with a bounded legacy area, as long as that area has an owner and does not gain new callers. Saying when you would stop is part of a staff answer, not an admission of failure.
How should I practise this question?
Pick an app you know, choose one costly journey, and answer out loud against a 30-minute timer: the outcome, the characterization tests, the seam, the first slice, the rollout stages, the halt conditions and what gets deleted. Then change one fact, for example a queued offline operation created by last year's version, and answer again.