Mobile architecture
Mobile release train: cut-off, hotfix and go/no-go
A mobile release train is a schedule plus a set of decision rights: who cuts, who may board late, who calls go or no-go, who can stop it at night. How to set the cadence, run the cut-off, ship a hotfix while a train is in flight, and measure the train itself.
13 min read
Most pages about the mobile release train describe the timetable: cut on Tuesday, submit on Thursday, repeat. The timetable is the easy part. What makes a train work, or quietly fail, is who is allowed to decide what: who sets the cut, who may add a commit after it, who calls go or no-go, and who can stop exposure at two in the morning without asking anyone.
I have run 500+ technical interviews from the hiring side over 15 years in mobile. When a senior candidate says "we had a weekly train", the follow-up that separates answers is "and who could break the cut-off?" A team that cannot answer has a calendar, not a release process.
My book Controlled Change frames the stakes in its chapter on migrations: "A migration that requires stopping the release train to complete has already failed; it just has not admitted it yet." The train is the thing everything else has to keep running around. So it deserves design.
- The cadence, and the two signals that say it is wrong.
- The cut-off, what may board after it, and who decides.
- How flags let the train leave without the feature.
- A hotfix while the next train is already in flight.
- Who can stop it, a go/no-go checklist, and the metrics of the train itself.
A release train is decision rights with a timetable
The train has one rule that everything else serves: the departure does not wait for any single piece of work. Whatever is ready at the cut ships; whatever is not rides the next train. That rule only holds if a short list of decisions has named owners before anyone needs them:
| Decision | Typical owner | What goes wrong without one |
|---|---|---|
| Cadence and cut-off time | Release captain or release engineering | The cut slides for the loudest feature |
| Late boarding after the cut | Release captain, on evidence from the change author | Cherry-picks by negotiation in a group chat |
| Go or no-go for submission | Release captain | A meeting that ends without a decision |
| Exposure of each feature | The feature team, through its flag | Shipping the binary and turning the feature on become one act |
| Halting exposure in an incident | On-call and the feature owner, each alone | Every problem waits for the next departure |
| Removing the flag afterwards | The feature team, with a dated ticket | The release flag becomes permanent architecture |
The Mobile System Design Blueprint's model answer for a release-train exercise puts the split in one line: a release captain owns the train and the go or no-go, product teams own their flags. Notice what that separates. The captain decides whether a binary leaves. The feature team decides whether users see a feature. Those are different decisions with different speeds, and a train only stays on time when they stay apart.
Pick the cadence from your bottleneck
Weekly or every two weeks is the question people search for, and the useful answer is that the period is an output, not a preference. Controlled Change says the branching model "depends on release cadence, team size, and regulatory needs", and its trade-offs are blunt: frequent releases reduce batch risk and increase operational cadence; long stabilization improves test time and increases divergence.
Work it out from three numbers you already have:
- Cut to store. How long from the cut to a build in review: stabilization, regression suite, signing, notes. If that takes most of a week, a weekly train leaves no slack for a bad candidate.
- Wait for an ordinary fix. The period is the longest a non-urgent fix waits after it lands. A long period pushes more fixes into the hotfix lane.
- Batch size. How many teams and changes ride each train. The more that ride, the harder a regression is to attribute, and the more a shorter period pays.
The fictional Atlas system in Controlled Change starts with manual weekly releases when it has two engineers, and later runs a weekly release train with progressive exposure at 150 engineers. Weekly is a common answer for teams that can keep the cut-to-store path short and automated. It is not a law. A regulated product may need approvals and evidence, but the book's line there is worth repeating to any compliance reviewer: that "does not justify infrequent giant releases."
Two signals tell you the cadence is wrong after the fact. A hotfix rate that keeps climbing says the train is too slow to carry fixes. A cut-to-store lead time that eats most of the period says the pipeline cannot sustain the cadence, and the fix is the pipeline, not more weekend work.
The cut-off is a revision, a branch and a rule
"Code freeze" is the phrase, and it misleads: trunk does not freeze. What happens at the cut is narrower. The model in Controlled Change:
- Integrate continuously on trunk.
- At the cut, take a short-lived stabilization branch or signed candidate from a known revision.
- Fix on trunk first, then cherry-pick to the release branch when necessary.
- Keep incomplete behavior dark behind flags.
- Automate versioning, signing and release notes.
- Maintain a clearly supported patch policy.
The direction matters. A fix that lands only on the release branch ships this week and regresses next week, because the next cut comes from trunk. The book's failure catalog lists the inverse as its own release failure: a hotfix branch that never merges back.
Keep the branch short-lived. The book says it directly: "Long release branches accumulate merge risk." A stabilization branch that lives for weeks is a second product line, and every cherry-pick becomes a merge conflict someone resolves under deadline.
What may board the train after the cut
Late boarding is where trains break, because every request sounds urgent to the person making it. Write the rule before the first train, so the answer does not depend on who asks:
- Boards: a fix for a regression introduced in this train, a crash or data-loss fix, a security fix, a fix for a store rejection of this candidate.
- Waits: new features, copy and polish, refactors, "small" improvements, anything whose risk is not covered by a focused test.
Then make the evidence and the approver part of the rule. A late change has landed on trunk, has a focused test that passes, and has one named approver. Written as code, the policy is short enough to put in the release checklist:
enum class Reason { REGRESSION_IN_THIS_TRAIN, CRASH_OR_DATA_LOSS, SECURITY, STORE_REJECTION, NEW_FEATURE, POLISH }
data class LateBoardingRequest(
val reason: Reason,
val mergedOnTrunk: Boolean, // fix trunk first, then cherry-pick
val focusedTestsPass: Boolean,
val approvedByCaptain: Boolean, // one named person decides, not a thread
)
sealed interface Boarding {
data object Board : Boarding
data class NextTrain(val why: String) : Boarding
}
fun decide(request: LateBoardingRequest): Boarding = when {
request.reason == Reason.NEW_FEATURE || request.reason == Reason.POLISH ->
Boarding.NextTrain("not a fix for this train; it ships dark next time")
!request.mergedOnTrunk ->
Boarding.NextTrain("land it on trunk first, or the next train regresses")
!request.focusedTestsPass ->
Boarding.NextTrain("no evidence yet")
!request.approvedByCaptain ->
Boarding.NextTrain("the release captain has not approved it")
else -> Boarding.Board
}Nobody needs this as a library. The point is that every branch has a reason a person can read back to the requester, and that a feature never reaches the approval question. Count the late boardings per train; the number is one of the metrics below.
Flags let the train leave without the feature
The rule that unfinished work waits only works if waiting is cheap. Flags make it cheap: the code merges, rides the train dark, and the feature is turned on later. The Blueprint's checkout scenario shows the split cleanly: on the weekly train the binary was cut and submitted on schedule regardless of the feature's readiness, "because deployment runs on a cadence", and the feature's exposure was a separate dial the team turned the following week once the backend was healthy.
That is the decoupling in one sentence. Deployment is the binary reaching devices, on the train's schedule. Release is users reaching the new code path, on the feature team's schedule. Controlled Change lists the stages that come after the binary ships: internal dogfood, a trusted cohort, a small production percentage, wider cohorts, default on, old path disabled, then flag and code removed.
Two consequences for the train. First, a release flag defaults to the old, proven path, so a device that never fetched configuration behaves like the previous release. Second, every flag that rides the train adds a branch that someone must delete. The cleanup is its own discipline; see how to clean up mobile feature flags. How the stores themselves stage a binary, and what halting a store rollout does and does not do, is a separate topic covered in the phased release vs staged rollout article; the short version is that no store control takes a binary back from a device that has it.
A hotfix while the next train is in flight
The hard case is two releases at once: version N is live and broken, and version N+1 is already cut, in stabilization or in review. Start with what does not need a binary. Controlled Change lists what rollback can mean once a build is installed: disable a flag, route to the old implementation, change a server response, block an unsafe operation server-side, and only then release a patched binary or, as a last resort, raise a minimum version. If a remote control can contain the harm, use it first; see how to design a mobile app kill switch.
When a binary is still needed, run it as its own small release:
- Branch from the tag of the build users have, not from trunk, so the hotfix carries nothing the shipped build did not.
- Make the minimal change with focused validation.
- Contain the harm with server and flag controls while the store review runs.
- Check migration safety: the hotfix must read whatever data version N already wrote.
- Number it so it sorts above the shipped build, and decide how the in-flight train sorts above the hotfix.
- Merge the fix back to trunk, and cherry-pick it onto the in-flight release branch if that train has not yet shipped.
- Close with cleanup and an incident review.
Numbering is where two concurrent releases bite. On Android, versionCode decides which build is newer, and the system will not install an APK with a lower versionCode over a higher one. Reserve the scheme before you need it, so a hotfix never collides with the build already in review.
Then decide what happens to the in-flight train: carry the fix and continue, or hold that departure. Either is defensible. What matters is that it is a recorded decision by the captain, and that the fix lands in the next build either way. The book's test for the whole path is plain: "A hotfix should not require reconstructing a discontinued toolchain or finding the only developer with signing access."
Who can stop the train
"Stop" means three different things, and each needs an owner named in advance:
- Hold a departure. The release captain can decline to submit a candidate. This is the go/no-go decision.
- Halt a store rollout. Release engineering pauses the staged distribution of a build already approved, so fewer devices receive it.
- Halt exposure. The feature owner or on-call turns a flag to zero or pulls a kill switch. This is the fastest stop and the only one that works on devices that already have the build.
The third is the one teams forget to assign. The Blueprint names the failure: "A train that departs weekly with no one empowered to pull the emergency brake between departures means every problem waits for next Thursday." In its checkout scenario, either the squad lead or the on-call could set exposure to zero acting alone, with no meeting. Controlled Change puts the same requirement in its release owner's readiness list: decision authority to pause.
Write the names into the release checklist, not into memory. For how an incident with these controls reads in an interview, see the production incident interview question.
A go/no-go checklist that fits on one screen
Go/no-go is a decision, not a ceremony. Controlled Change warns against "ceremony so large that teams bypass it" and asks for automated evidence with human review focused on changed risk. A checklist the captain can read in two minutes:
- The candidate is a known revision with reproducible provenance and symbols.
- Every late boarding since the cut is on trunk, tested and approved.
- New features in this build are dark by default; each flag has an owner, a safe default and a removal date.
- Backend support for this build is deployed and compatible with the supported old clients.
- Local database and pending-operation migrations pass from every supported version.
- Stop conditions for the rollout are written: crash or hang change, critical journey completion, contract errors, support signals.
- Kill switches and flag reversals for this build have been exercised, not assumed.
- The hotfix path, the halt owner and the on-call are named for the release window.
Each item is a yes or a no. Any no is a no-go or an explicit, recorded exception with an owner. The book's line on the stop conditions applies here: "A release without measurable stop criteria is an uncontrolled experiment."
Measure the train, not only the app
Crash rate tells you about the build. It does not tell you whether the train is healthy. Track the process too:
| Signal | What it tells you | Action when it drifts |
|---|---|---|
| On-time departures | Whether the cut holds against pressure | Find who moved it and why; tighten late-boarding rules |
| Cut-to-store lead time | Whether the pipeline fits the cadence | Automate or shorten stabilization before changing the cadence |
| Late boardings per train | How much risk enters after the cut | Review what each fixed; a spike means work is merging unready |
| Hotfix rate | Whether the train carries fixes fast enough | Review what the train missed; consider the cadence |
| Time to contain | How fast a halt reaches devices | Drill the flag and kill switch path |
| Flag age | Whether dark launches are being cleaned up | Open removal work; old flags are debt |
The Blueprint's checkout register gives example budgets for two of these, cut to store under five days and under one out-of-band build per month, with the actions attached. Treat such numbers as starting points from a composite scenario, not standards. Controlled Change's release budgets add the rest: rollout stages and minimum observation windows, configuration propagation and rollback time, and the age of flags and compatibility bridges.
The short version for an interview
If an interviewer asks how you would run releases for a large app, the strong answer fits in five sentences:
- "The train departs on a fixed cadence; anything not ready rides the next one, dark behind a flag."
- "We cut a short-lived branch from a known revision, fix on trunk first and cherry-pick only regressions, crashes, security and store fixes, approved by the release captain."
- "The captain owns go or no-go; feature teams own exposure through their flags; on-call can halt exposure alone at any hour."
- "A hotfix branches from the shipped tag, is numbered above it, is contained by flags while in review, and merges back to trunk."
- "We measure on-time departures, cut-to-store time, late boardings, hotfix rate and time to contain, and change the cadence when those say so."
That answer shows the interviewer the thing they are scoring: that you see a release as decisions with owners, not as a date on a calendar.
Questions engineers ask about mobile release trains
What is a mobile release train?
A release process with a fixed departure schedule. On a known date the team cuts a release candidate from a known revision, stabilizes it for a short, bounded period, and submits it to the stores. Work that is not ready at the cut does not delay the train; it waits for the next one, usually shipped dark behind a flag so the code can travel without users seeing it.
Should a mobile app release weekly or every two weeks?
Pick the shortest cadence your stabilization and review time can sustain without skipping evidence. A shorter period means smaller batches and a shorter wait for ordinary fixes; a longer one buys more test time and more divergence between the release branch and trunk. Two signals tell you the cadence is wrong: a rising hotfix rate says the train is too slow to carry fixes, and a cut-to-store lead time that eats most of the period says it is too fast for your pipeline.
What happens to a feature that misses the code cut-off?
It rides the next train. If its code is already merged, it ships in the binary with its flag off by default, and exposure is turned on later as a separate decision. A feature is not a reason to hold the train or to cherry-pick after the cut; the late-boarding lane is for fixes to this train, not for new scope.
Can you ship a hotfix while the next release is already in review?
Yes, as a second, separate release: branch from the tag of the build users actually have, make the smallest change, validate it narrowly, and merge the fix back to trunk so the next train carries it. Plan the build numbers in advance. On Android, versionCode decides which build is newer, and the system will not install an APK with a lower versionCode over a higher one, so the hotfix has to sort above the shipped build and the next train above the hotfix.
