Mobile architecture
Mobile feature flag cleanup: the flag outlives the release that removed it
On a server, cleaning up a feature flag is one motion. On a phone it is three, on three clocks, and the last one belongs to the installed fleet. Expiry by job, the order that survives old binaries, and how to price the flags you already have.
11 min read
Most advice on feature flag cleanup is written for web services. The rollout finished, so delete the branch, delete the key, deploy. On a server that is one motion. On a phone it is three, on three different clocks, and the last clock is not yours.
I have run 500+ technical interviews from the hiring side over 15 years in mobile. "We clean up the flags after the rollout" is a sentence I hear often in staff loops and architecture reviews, and it is rarely specified. A mobile flag is not gone when you delete the branch. It is gone when no supported binary still reads it.
This is the procedure from the day a flag is created to the day its last trace is removed. It follows my book Controlled Change: chapter 39 on feature flags, experiments and kill switches, and the flag cleanup playbook in Appendix D.
- Why mobile cleanup is different: the binary that evaluates the flag stays installed.
- Expiry by job: a release flag, an experiment and a kill switch end in different ways, or not at all.
- The two clocks: when the code can go, and when the remote key can go.
- The cleanup order, flag debt as architecture debt, and where to start in an old codebase.
Why mobile cleanup is different: the binary stays installed
Controlled Change opens its release chapter with the constraint every mobile engineer knows and few cleanup plans account for: a mobile release cannot be rolled back like a server deployment, and an installed binary may run for months against changing backends and configuration.
Every version you shipped with a flag check carries that flag's behavior inside it: the key it reads, the default it falls back to, the cache policy it applies. Deleting the check in your next release changes the next release. It changes nothing on the phones that have not updated.
The book's chapter on version skew lists the clocks a mobile system runs on, among them app releases, user adoption and feature-flag configuration, and says it plainly: "Compatibility is not a courtesy. It is a production topology." A flag lives on three of those clocks:
- The code clock: the branch in your source, removed in your next release. You control it.
- The fleet clock: the remote key, which old binaries keep reading until they update or fall out of support. Your users control it.
- The contract clock: the backend behavior, analytics dimensions and storage that serve the losing path. Your support window controls it.
Cleanup plans that run all three on the first clock are the ones that go wrong.
Expiry depends on the job the flag does
Before a flag can expire, someone has to say what kind of flag it is. The book's first rule for flags is not to put every boolean in one bag: the purpose decides the lifetime, the default, the audience and the evidence.
| Job | Expected lifetime | What the end looks like | Default when the value is missing |
|---|---|---|---|
| Release flag | Days to weeks | Winner made unconditional, flag deleted | The old, proven path |
| Experiment | Bounded by the analysis plan | A cleanup decision when the analysis closes | The control assignment |
| Migration flag | The migration window | Deleted with the old implementation | The path you can roll back to |
| Kill switch | Long-lived, rarely active | No end date: an owner, a review, a drill | The safest functional state |
| Configuration parameter | Ongoing, with a schema | Governed configuration, not cleanup | An embedded, validated value |
Two jobs on that table never expire, and that is deliberate. A kill switch earns its keep by being there when an incident starts; designing one well (fail open or fail closed, who may pull it, how you drill it) is a topic of its own. An entitlement is not on the table at all: the book is explicit that a client-side flag cannot authorize paid or privileged behavior, and the backend stays authoritative.
The book's ADR library adds the rule for the gray zone. An operational switch that proves permanent moves into governed configuration instead of pretending to be a temporary release flag. That rule keeps the cleanup list honest: everything on it is supposed to end.
Set the expiry when the flag is created
The cheapest moment to decide when a flag ends is the moment it is born, when the person creating it still knows why. Chapter 39 lists the metadata every temporary flag needs; the ones that make cleanup possible are the owner, the type, an expected expiry or revisit trigger, the safe default, the modules and contracts it touches, and a removal issue.
Metadata only helps if something reads it. The book's governance chapter gives the fitness function for this principle in one row: flags are temporary, checked by an owner and expiry check and a stale-flag report. Here is a small version that can run in CI against a flag registry:
import Foundation
enum FlagJob {
case release, experiment, migration // temporary: each one must end
case killSwitch, configuration // long-lived: reviewed, not expired
}
struct FlagDefinition {
let key: String
let job: FlagJob
let owner: String
let expires: Date?
let removalIssue: String?
var isTemporary: Bool {
switch job {
case .release, .experiment, .migration: return true
case .killSwitch, .configuration: return false
}
}
}
enum FlagProblem: Equatable {
case missingExpiry(key: String)
case missingRemovalIssue(key: String)
case expired(key: String, owner: String)
}
// Runs in CI on every change to the flag registry.
func audit(_ flags: [FlagDefinition], now: Date) -> [FlagProblem] {
flags.filter(\.isTemporary).flatMap { flag -> [FlagProblem] in
var problems: [FlagProblem] = []
if flag.removalIssue == nil {
problems.append(.missingRemovalIssue(key: flag.key))
}
guard let expires = flag.expires else {
return problems + [.missingExpiry(key: flag.key)]
}
if expires < now {
problems.append(.expired(key: flag.key, owner: flag.owner))
}
return problems
}
}Two choices in that check matter more than the code. Kill switches and configuration pass without an expiry, because their job is to last. And an expired flag reports its owner, because a warning addressed to nobody is a warning nobody acts on. The book also asks for one more rule: block new code from depending on a flag already marked for removal.
Chapter 47 says this about architecture exceptions, and it fits flags word for word: "An exception with no expiry is a changed standard." A temporary flag with no end date is a permanent branch that nobody decided to keep.
Two clocks: when the code can go, and when the key can go
The code clock is simple. Once the rollout outcome is confirmed, make the winning branch unconditional in the next release. Nothing about old binaries stops this: the new binary does not read the flag at all.
The fleet clock depends on a question most cleanup plans never ask: which way did the flag resolve? For a release flag, the embedded default is the old path. So when an old binary stops seeing the key, it falls back to the old path.
- The new path won. Deleting the key sends every installed version that still reads the flag back to the path you just retired. Keep serving the winning value until the minimum supported version is past the last release that evaluates the flag.
- The old path won. A missing key falls back to the default, and the default is the winner. One caveat: how an old binary's cache treats a missing key is decided by code you already shipped. Serve the off value until devices stop evaluating the new path, then delete.
struct AppVersion: Comparable {
let major: Int, minor: Int, patch: Int
static func < (a: AppVersion, b: AppVersion) -> Bool {
(a.major, a.minor, a.patch) < (b.major, b.minor, b.patch)
}
}
enum Winner { case oldPath, newPath }
enum KeyDecision: Equatable {
case deleteOnceOffHasPropagated
case keepServingWinner(untilMinimumSupportedPasses: AppVersion)
case delete
}
// lastReadingVersion: the last release that still evaluates the flag.
// The release after it made the winning branch unconditional.
func retireRemoteKey(winner: Winner,
lastReadingVersion: AppVersion,
minimumSupported: AppVersion) -> KeyDecision {
switch winner {
case .oldPath:
// A missing key means the embedded default, and for a release flag
// that default is the old path: the winner. How an old cache treats
// a missing key depends on code you already shipped, so serve "off"
// until devices stop evaluating "on", then delete.
return .deleteOnceOffHasPropagated
case .newPath:
// A missing key would send every binary that still reads the flag
// back to the path you just retired.
return minimumSupported > lastReadingVersion
? .delete
: .keepServingWinner(untilMinimumSupportedPasses: lastReadingVersion)
}
}Then the rule that keeps this from turning into a forced update: do not raise the minimum version to delete a flag. The book reserves a hard minimum version for narrowly justified risk, such as compromised credentials, unsafe policy, legal requirements or a server contract that cannot be maintained, because a block can lock out users who cannot update. A remote key that costs almost nothing to keep serving is none of those. Keep serving it, and let the support window close on its own schedule.
The order that survives old binaries
Appendix D of Controlled Change closes with a playbook for feature-flag and compatibility cleanup. In order, it asks you to:
- inventory the flag's type, owner, current value, clients and data dependencies;
- confirm the rollout outcome and the losing path;
- make the winning behavior unconditional in code;
- remove variant-specific tests and telemetry, but only after replacing the outcome metrics you still need;
- delete the remote key, its schema and the permissions to publish it;
- remove the compatibility adapter, the old API and the storage columns;
- verify that no runtime evaluation or active old client still requires it;
- close the ADR and the migration record with evidence.
The appendix says to adapt each playbook rather than copy it mechanically, and on mobile I adapt one thing: the check in step 7 also runs before step 5. Evaluation telemetry broken down by app version is the evidence that the fleet clock has run out, and it should exist before the key goes, not only after.
Step 4 is the one teams skip. The dashboard that proved the rollout is usually built on the variant dimension, so deleting the variant deletes the evidence. Step 5 has a detail worth noticing too: the permissions go with the key, so nobody can quietly recreate it for a binary that still reads it.
Step 6 runs on the contract clock. The backend support for the losing path follows the same rule the book gives for API migrations: remove old behavior only after the measured support window. The playbook's warning names the failure at the end of this list: the key is deleted while dead branches and bridges stay in the code. That half-finished cleanup leaves a flag that can no longer be turned on and a code path that can no longer be removed with confidence.
Flag debt is architecture debt, so price it
Chapter 39 puts its cleanup section under a blunt heading, flag debt is architecture debt, and chapter 47 gives the vocabulary to make that case to people who fund the work. Two of its debt categories describe a stale flag exactly: deprecation debt, where old paths remain after replacement, and compatibility debt, where old clients or contracts have no closure plan.
The book asks every debt to carry an interest, the cost paid per change or period, and a principal, the cost of removing it. For a flag, the interest is the state space. Flags multiply the release compatibility matrix, and the scenarios that need explicit tests are the ones a stale flag keeps alive: an old client with a new backend and the flag off, an operation created under old policy and resolved under new policy, a rollback after the local schema has already migrated. The book's answer is not more tests. It is a smaller state space: scope flags tightly, snapshot policy, and delete flags promptly.
Turn that into numbers before the rollout starts. The release budgets in the book's Appendix F include a maximum number of active temporary flags per feature and a maximum age for flags and compatibility bridges. A breach is then a decision with an owner, not a vague sense that the codebase is getting heavy.
Every live flag is also one more branch to rule out during an incident. How interviewers test that instinct is in the production incident interview answer.
Starting from a codebase full of old flags
If the registry does not exist yet, the book's migration path for flags is short and in a sensible order: inventory every flag and classify its job, owner, default and age; wrap one critical domain in a typed policy object; delete the oldest safe flag; then add expiry enforcement, configuration-failure tests and a kill-switch drill.
"Oldest safe" has a precise meaning after the two clocks above. It is a flag whose code clock and fleet clock have both run out: no supported binary evaluates it. Flags whose winner was the old path are often the quickest first deletions, because a missing key already sends every client to the winning behavior once the off value has propagated.
Resist the cleanup quarter. The governance chapter warns that a separate tech debt quarter often produces a temporary cleanup followed by ordinary neglect, and asks for architecture work to travel with the feature work that caused it. For flags, that means the removal issue is part of the feature's definition of done. The book's release chapter makes the same point in its exposure stages: the last two are "old path disabled" and "flag and code removed". A rollout that stops at default on is not finished.
What I listen for when a team says the flags are under control
| Topic | Mid-level | Senior | Staff |
|---|---|---|---|
| Expiry | "We clean up after launch." | An owner and an end date on each flag. | Expiry by job, enforced in CI; kill switches reviewed, not expired. |
| Code | "Delete the if." | Winner unconditional in the next release. | Replaces outcome metrics first; removes tests, analytics and docs with the branch. |
| Remote key | "Delete it in the console." | Waits for old versions to update. | Knows which way the flag resolved; ties deletion to the minimum supported version, never the reverse. |
| Debt | "We will do a cleanup sprint." | Keeps a stale-flag list. | Prices flags as debt with a budget per feature and an owner per breach. |
The staff column is the same cleanup, with the fleet in it. The same deletion discipline is what closes a migration, which is why it decides the architecture migration interview question as well.
Questions engineers ask about feature flag cleanup
When should you remove a feature flag from a mobile app?
Remove the code as soon as the rollout outcome is confirmed: make the winning branch unconditional in the next release. Remove the remote key later, when no supported app version still evaluates it. Those are two different dates, and on mobile the second one is set by your oldest supported binary, not by your release calendar.
Can I delete the remote config key once the rollout is at 100 percent?
Not if the new path won. Every installed version that still reads the flag falls back to its embedded default when the key disappears, and for a release flag that default is the old path. Keep serving the winning value until the minimum supported version is past the last release that reads the flag.
Should every feature flag have an expiry date?
Every temporary one: release flags, experiments and migration flags. Kill switches and configuration parameters are long-lived by design, so they get an owner and a review instead of an end date. A switch that turns out to be permanent should move into governed configuration, not keep pretending to be a temporary release flag.
How many feature flags is too many?
There is no universal number. Set a budget per feature for active temporary flags and a maximum age for flags and compatibility bridges, write it down before the rollout, and treat a breach as debt with an owner. The cost of a flag is the state space it adds to every test and every incident.
