Skip to content
All essays

Mobile architecture

Phased release vs staged rollout: what halting does

Apple's phased release and Google Play's staged rollout both stop the spread of a build, and neither takes it back. What each store documents, what pausing and halting reach, why a crash-free number at 1 percent proves little, and how to write the halt rule before day one.

14 min read

I have run 500+ technical interviews from the hiring side over 15 years in mobile. When a system design round reaches "how would you ship this?", the common answer is "phased release, start at 1 percent, watch crashes." The follow-up that separates candidates is short: what does pausing actually stop, and what number makes you press it?

This article answers both, from the release-control side. It sets Apple's phased release next to Google Play's staged rollout using only what Apple and Google document, then treats the crash-free rate as what it should be: one gate among several, with a denominator, a baseline and a rule written before day one.

  • The mechanisms: what each store does, and what pause and halt can and cannot reach.
  • The gate: why a fixed crash-free number at a 1 percent stage proves little, and what to compare instead.
  • The rule: a halt policy you can write on one page, and the levers that still work once a build is installed.

What Apple's phased release does

Apple's App Store Connect documentation, Release a version update in phases, describes the mechanism precisely. Read every line of it, because the details decide what a pause is worth:

  • The update goes out over 7 days: 1 percent on day 1, then 2, 5, 10, 20, 50 and 100 percent on day 7.
  • It reaches a random sample of users with automatic updates on eligible devices, who are not told they are part of it.
  • Anyone can manually download a version in phased release from the App Store at any time.
  • You can pause for a total of 30 days, as many times as you like within that total. A resumed release picks up on the day it left off.
  • You can choose Release to All Users at any point, which makes the version available to every user with automatic updates.
  • Removing the app from sale stops the phased release for that version, and when the app returns, that version is available to all users immediately.

Two consequences follow. The percentages describe the automatic-update population, not everyone who has the new build: people who open the App Store and tap Update are outside the schedule. And a pause stops the schedule from advancing. Apple's page does not describe it as blocking manual downloads, and it describes no step that serves the previous version instead of the current one.

What Google Play's staged rollout does

Google's Play Console help page, Release app updates with staged rollouts, describes a different shape of control:

  • You choose the percentage of users who receive the release, and it does not increase automatically. Every step is a decision someone makes.
  • New and existing users are eligible, chosen at random for each release rollout.
  • You can target countries, and once a staged rollout has started you cannot remove a country.
  • When you halt, no additional users receive that version. Users who already received it remain on it.
  • If you resume a halted rollout, you affect the same set of users. If the bundle is broken, Google's guidance is to roll out a new release with a fixed bundle.

Play also documents halting a fully rolled-out release. Halting a release at 100 percent stops new and existing users from installing or updating to that version, and the previous fully rolled-out version takes its place for new and eligible users. You cannot do it for the first release on a track, and a previous release with a policy violation cannot stand in. The page describes stopping installs and updates; it does not describe moving devices that already have the halted version back.

What halting actually does, side by side

Put the controls in one table and the pattern is plain. Every store control acts on the devices that do not have the build yet. None acts on the devices that do.

ControlWho already has the buildWho can still get itWhat it cannot do
App Store: pause phased releaseEveryone updated so far, automatic and manualAnyone who updates manuallyRemove the build from devices, or pause for more than 30 days in total
App Store: Release to All UsersSame, plus everyone with automatic updatesEveryoneApple documents no way back to the phased schedule for that version
Play: halt a staged rolloutUsers who received it stay on itNo additional usersDowngrade anyone; resume reaches the same set of users
Play: halt a fully rolled-out releaseUsers on the halted versionNew and eligible users get the previous versionWork for the first release on a track, or fall back to a release with a policy violation
Remote flag or kill switchBehavior changes on versions that read itNot applicable: acts on behavior, not installsReach versions that never read the switch
Server-side refusalEvery version that calls the serverNot applicableStop work that never leaves the device
New binaryEventually everyone who updatesEveryone, after review and adoptionHelp anyone this week

That is why my book Controlled Change refuses to let the store carry the whole plan: "Store phased rollout is one mechanism." The store limits how many devices hold the new code. A flag decides how many of them run it. For the design of that second control, see mobile app kill switch design.

Ship the code dark, gate exposure separately

Controlled Change separates progressive exposure into stages: developer and automated environments, internal dogfood, a trusted beta cohort, a small production percentage, broader percentages by platform, device or region, default on, old path disabled, and finally flag and code removed. The store rollout covers only the move of a binary onto devices. Exposure is the moment a user actually runs the new path.

In practice that means two dials. The store dial (phased release or staged rollout) spreads the binary with the new path present but off. The flag dial turns the path on for a percentage you control, and can turn it back to zero for every version that reads it, on both platforms, without a review queue. When a candidate tells me the plan is "1 percent phased release", the strong follow-up answer is that the 1 percent is a ceiling on exposure, and the flag is where the real gate lives.

The second dial has a cost: the old path must stay shippable inside the binary for the whole rollout, and the flag must be removed afterwards. Mobile feature flag cleanup covers that end of it.

Crash-free rate needs a denominator first

"Crash-free rate" names at least three different numbers, and the gate is meaningless until you say which:

  • Crashlytics crash-free users: the percentage of users who engaged with the app in the period without a crash, where a user is one installation on one device (Firebase documentation).
  • Crashlytics crash-free sessions: the percentage of sessions that did not end in a crash. A new session starts at cold start, or when the app is foregrounded after at least 30 minutes in the background. Both Crashlytics metrics count fatal events only.
  • Android vitals user-perceived crash rate: issues per daily active user. Google's own example: a user who opens a game three times in a day and crashes once shows as 100 percent in Android vitals and 33 percent in Crashlytics (Android vitals).

The sources also see different people. Android vitals counts only certified devices with the app installed from Google Play, and only users who agreed to share data. On iOS, Xcode's Crashes organizer shows reports from customers who share diagnostic and usage information (Apple documentation). A low crash count can mean a stable build or a population that does not report.

Controlled Change says it in four words: "Every budget needs a denominator." Its appendix on budgets adds the reason, that crash-free users and crash-free sessions answer different questions. Pick one per gate, write down the eligibility rule and the window, and never compare a sessions number from one tool with a users number from another.

One more trap: Google Play publishes bad behavior thresholds, a 1.09 percent user-perceived crash rate and a 0.47 percent user-perceived ANR rate overall, above which Play may reduce your app's visibility or show a warning on the listing. Those are store-quality thresholds. They are not your release gate, and a release that stays under them can still be much worse than the version it replaces.

The arithmetic of a 1 percent stage

Small stages are for catching large failures early. They cannot prove a build is as stable as the last one. Run the numbers with an illustrative case: the previous version crashes for 0.5 percent of daily users, and the new one is twice as bad, at 1 percent.

Daily users on the new buildExpected users with a crash, healthy buildExpected, build twice as badWhat you can conclude
40024Nothing. A healthy build shows 4 or more crashing users about 14 percent of the time, and the bad build shows 3 or fewer about 43 percent of the time.
2,0001020Something. With a rule that halts at 16 or more, a healthy build trips it about 5 percent of the time, and the bad build slips under it about 16 percent of the time.
20,000100200A doubling is unmistakable. Smaller regressions now become visible.

The probabilities come from treating crashing users as Poisson counts, which is close enough to set expectations. Play Console's release dashboard makes the same point visually: when installs are too low for normalization to be meaningful, the normalized crash and ANR lines are drawn dotted.

So a rule like "halt below 99.5 percent crash-free" at day 1 of a phased release is either noise or silence. Write the minimum sample into the rule, and treat "not enough data yet" as its own state, distinct from "healthy". Then size the stages from your own traffic: if 1 percent of your users is 400 people a day, the first stage catches crash-on-launch and little else, and that is fine as long as everyone knows it.

Compare with the previous version, not a fixed number

An absolute threshold ignores what normal looks like for your app. A media app with heavy native code and a forms app have different baselines, and the same app has different baselines on different OS versions and device tiers. The question a gate should answer is narrower: is the new version worse than the one it replaces, for the same kind of user, over the same window?

  • Same window. Compare day 2 of the new version with the previous version over the same calendar days, not with last month's average.
  • Same segment. Compare by OS version, device tier, region and app version. Aggregate health can improve while one segment collapses.
  • Same metric definition. One tool, one denominator, on both sides.
  • Mind who updates first. On iOS, the early population mixes a random sample of automatic updaters with people who chose to update manually. They are not guaranteed to resemble the rest of your users, so treat early numbers as directional.

Keep an absolute ceiling as well, for the case where both versions are bad. The ratio catches regressions; the ceiling catches a baseline nobody should accept.

Crash-free is one gate of several

A build can be crash-free and still broken. Controlled Change says it directly: "Crash-free does not mean usable." Its release chapter lists the stop conditions to define before exposure, and only the first is about crashes:

  • crash, hang or ANR change;
  • critical journey completion;
  • operation unknown or rejection rate;
  • startup or frame regression;
  • authentication or contract errors;
  • accessibility or support signals;
  • AI quality or policy thresholds, and backend saturation or cost, where they apply.

The book's fictional reference system, Atlas, has an incident written for exactly this gap: a release retries booking with a new idempotency key after a timeout, a small network cohort creates duplicate reservations, crash-free metrics stay green, and support finds it. The lesson the book draws is that the release guardrail measured request errors, not duplicate reservations.

A halt rule therefore has two kinds of input. Rates, such as crashing users or failed checkouts, need a sample and a baseline. Invariant violations, such as a duplicate charge or a lost record, need neither: one is enough to stop.

Kotlin
// One guardrail: a rate where lower is better (users with a crash, ANRs,
// failed checkouts), measured for the new version and for the previous
// version over the same window and the same segment.
data class Guardrail(
    val name: String,
    val candidateUsers: Int,
    val candidateBad: Int,
    val baselineUsers: Int,
    val baselineBad: Int,
    val maxRatio: Double,     // 1.5 = new may be 50 percent worse, no more
    val ceiling: Double,      // absolute rate that halts regardless of baseline
) {
    val candidateRate: Double get() = candidateBad.toDouble() / candidateUsers
    val baselineRate: Double get() = baselineBad.toDouble() / baselineUsers
}

sealed interface StageDecision {
    data object WaitForSample : StageDecision
    data object Continue : StageDecision
    data class Halt(val reason: String) : StageDecision
}

class HaltRule(private val minimumUsers: Int) {
    fun evaluate(guardrails: List<Guardrail>, invariantViolations: Int): StageDecision {
        // A duplicate charge or a lost record needs no sample size.
        if (invariantViolations > 0) return StageDecision.Halt("invariant violated")
        var waiting = false
        for (g in guardrails) {
            // Below the minimum, a rate is noise. Waiting is not continuing.
            if (g.candidateUsers < minimumUsers || g.baselineUsers < minimumUsers) {
                waiting = true
                continue
            }
            if (g.candidateRate > g.ceiling) return StageDecision.Halt("${g.name} above ceiling")
            if (g.candidateRate > g.baselineRate * g.maxRatio) {
                return StageDecision.Halt("${g.name} worse than the previous version")
            }
        }
        return if (waiting) StageDecision.WaitForSample else StageDecision.Continue
    }
}

Three decisions are deliberate. Invariants halt before any sample check. A guardrail without enough users makes the stage wait, never continue. And each rate is judged against the previous version with a ratio and against an absolute ceiling, so a regression and a bad baseline both stop the rollout.

Write the halt rule before day one

Controlled Change puts the order in one line: "Write halt conditions before rollout." A rule written after the numbers arrive will be argued away by whoever wants to ship. Fit it on one page:

FieldWhat to write
GuardrailsEach metric, with its tool and denominator: crashing users, ANR rate, cold start p95, completion of the critical journey
ComparisonPrevious version, same window, same segments, with the segments listed
Minimum sampleUsers per segment before a rate counts; below it the stage waits
ThresholdsMaximum ratio to baseline and absolute ceiling, per guardrail
InvariantsEvents that halt on the first occurrence
Stages and windowsStore percentage, flag percentage and the minimum time at each step
Who may haltA named role that can halt without convening a meeting, and who may resume
What halt meansPause phased release, halt staged rollout, set the flag to zero, refuse at the server: which, in what order
After haltWho owns the fix, the new build, and the decision to resume or abandon

The book lists "decision authority to pause" among what a release owner needs, next to dashboards, cohort comparisons and a flag and kill-switch inventory. Most rollouts that go wrong had the dashboard. They did not have the sentence saying who may act on it.

After the halt: what still works

Halting is containment of spread, not recovery. The devices that already have the build keep it on both platforms, so the next steps act on behavior and data:

  • Turn the flag for the new path to zero, so versions that read it fall back to the old path.
  • Refuse unsafe operations at the server, which reaches every version, including the ones that never read the flag.
  • Repair data the bad build already wrote, if any.
  • Ship a fixed build. On Play, roll out a new release with the fixed bundle. On the App Store, submit a new version, with phased release enabled again if you want a controlled rollout.
  • Resume or abandon the halted rollout on purpose, and record why.

In an interview, the incident story around a halted rollout is where staff-level answers are separated from senior ones: the halt is step one, not the resolution. See the production incident interview question for how to tell that story.

The short version for an interview

If the question is "how would you roll this out?", an answer that holds up under follow-up fits in four sentences:

"I ship the binary through phased release on iOS and staged rollout on Play with the new path behind a flag that defaults off. The store stages limit how many devices could run it, the flag decides how many do, and only the flag can go back to zero on devices that already updated. Before day one I write the halt rule: crashing users, ANR and checkout completion against the previous version in the same window and segments, a minimum sample, and invariants that stop on first occurrence. The on-call lead can halt without a meeting, and halting means flag to zero first, store pause second, server refusal if data is at risk."

Questions engineers ask about phased release and staged rollout

What is the difference between App Store phased release and Google Play staged rollout?

Apple's phased release follows a fixed seven-day schedule (1, 2, 5, 10, 20, 50 and 100 percent) for users with automatic updates, and you can pause it for up to 30 days in total. Google Play's staged rollout uses a percentage you choose, does not increase on its own, and can be halted. Neither one removes the build from devices that already installed it.

Does pausing a phased release stop users from getting the update?

It stops the automatic-update schedule from advancing. Apple's documentation also says that apps and updates in phased release can be manually downloaded from the App Store by anyone at any time, so a pause is not a block on the version.

What happens to users when you halt a staged rollout on Google Play?

No additional users receive that version, and users who already received it remain on it. If you resume, Google says you affect the same set of users. For a release already at 100 percent, Play can halt it and make the previous fully rolled-out version available to new and eligible users, but not for the first release on a track.

What crash-free rate should block a mobile release?

There is no universal number. Write the rule as a comparison with the previous version over the same window and segment, with a minimum sample, a maximum allowed ratio and an absolute ceiling. Google Play's bad behavior thresholds (1.09 percent user-perceived crash rate, 0.47 percent ANR rate) are visibility thresholds for your listing, not a release gate.

Share this essay