Mobile architecture
Mobile app kill switch: how to design one
A remote kill switch is the one control you still hold after a binary is installed. How to design it so it works on the day you need it: what off means, safe defaults, caching, independence, reach, audit and drills.
12 min read
This article is about remote kill switches inside apps: the engineering control a team uses to turn off a broken or harmful capability in iOS and Android apps that are already installed. It is not about anti-theft features that lock a stolen phone, which share the name and most of the search results.
I have run 500+ technical interviews from the hiring side over 15 years in mobile. In system design rounds, "and how would you turn it off?" is a follow-up most candidates answer with one word: a flag. The word is right. The design behind it is usually missing, and that design decides whether the switch works on the day it is needed.
My book Controlled Change defines the control in its glossary as "A privileged control that disables or degrades risky behavior without waiting for a new binary." Every part of that sentence carries weight. Privileged: someone has to be allowed to pull it, and someone else must not be. Disables or degrades: off has to mean something specific. Without waiting for a new binary: it has to reach phones that will not update this week.
- The problem: a mobile binary cannot be pulled back from devices, and a fix takes a review queue plus an update cycle.
- The design: what off means, a safe default per capability, a cached last known value, independence from the failing path, reach, audit and a drill.
- The limits: what no client-side switch can undo, and what you pair it with.
A kill switch is not a release flag pointed the other way
Teams tend to keep every remote boolean in one bag. Controlled Change starts by splitting them by job, because the job decides the lifetime and the safe default:
| Control | Job | Lifetime | Safe default |
|---|---|---|---|
| Release flag | Separate deployment from exposure | Days to weeks | The old, proven path |
| Experiment | Compare product hypotheses | Bounded by the analysis plan | Control assignment |
| Operational kill switch | Disable a harmful capability | Long-lived, rarely active | The safest functional state |
| Entitlement | Represent server-authorized access | Ongoing | Deny until authorized |
The kill switch is the odd one out. A release flag is busy for a few weeks and then deleted. A kill switch sits unused for months and has to work perfectly the one time it is pulled, usually during an incident, usually by someone who did not build it. That is why the book's heading for it is blunt: kill switches must be boring and proven.
It also follows that a release flag is not automatically a kill switch. If the rollout flag for a new checkout is deleted the week the rollout ends, the new checkout no longer has a remote off. Decide, before the cleanup, which capabilities keep a standing switch.
Decide what off means before you build the switch
"Turn it off" is not a behavior. Write down, per switch, what the user sees and what happens to work in progress:
- Hide an optional entry point, for example a new sharing sheet.
- Route back to the old implementation, for example new checkout attempts return to the previous flow.
- Make it read-only: the user can see their data but not change it.
- Degrade: a smaller image size, no live updates, a manual refresh instead of a socket.
- Block one operation while the rest of the feature keeps working.
Whatever off means, it must not strand or discard work the user already committed. A draft, a queued upload or a pending payment created before the switch was pulled needs a defined path: finish under the protocol it started with, pause and wait, or move to a state support can see. A switch that silently drops accepted work has turned one incident into two.
Also say what off must not touch. A switch that takes out login, account recovery or the route to support to stop a broken photo filter has a blast radius larger than the bug.
Fail open or fail closed: choose per capability
The device reads the switch from configuration, and configuration fails. Before shipping, answer each of these for every switch:
- no configuration has ever been fetched (a fresh install on a train);
- the cached value is old;
- the value fails validation or signature checks;
- the user is offline, or the configuration service is degraded;
- the account on the device has changed;
- the value is missing, unknown or out of range.
The rule that keeps most teams safe: absence means safe. If the code reads "on unless the config says off", every failed fetch serves the risky path, and an outage of the configuration service turns into a full launch. Write it as "off unless a validated value says on" for anything new, destructive or privileged. For an established critical journey, the safe state is not nothing: it is the last known safe local implementation, so checkout falls back to the old checkout instead of disappearing.
Fail open is a legitimate choice for some capabilities, for example a cosmetic feature where an outage of the config service should not change what users see. The mistake is not choosing fail open. The mistake is never choosing, and inheriting whatever the SDK's default happens to be.
Cache the last known value, and let staleness move only toward safe
A switch read only from the network is useless at cold start, on a plane, and during the very outage you are trying to contain. Persist the last validated value and read it from disk at launch. Then decide how long an old value counts as evidence:
import Foundation
enum Capability: Sendable { case enabled, disabled }
// What the device last received from a successful, validated fetch.
struct CachedSwitch: Codable, Sendable {
let disabled: Bool
let fetchedAt: Date
}
struct KillSwitch: Sendable {
let key: String
let safeDefault: Capability // decided per capability and written down
let maxAge: TimeInterval // how long an "enabled" value may be trusted
func resolve(cached: CachedSwitch?, now: Date = Date()) -> Capability {
// Fresh install, wiped cache or a value that failed validation.
guard let cached else { return safeDefault }
// A kill never expires into an enable. Only a newer fetch lifts it.
if cached.disabled { return .disabled }
// An old "enabled" stops being evidence after maxAge.
if now.timeIntervalSince(cached.fetchedAt) > maxAge { return safeDefault }
return .enabled
}
}
// Read from disk at launch, before the protected feature or its SDK starts.
let rebuiltCheckout = KillSwitch(key: "checkout_rebuilt", safeDefault: .disabled, maxAge: 24 * 60 * 60)Two choices in that code are deliberate. A cached kill does not expire: a device that went offline right after the switch was pulled stays safe until it fetches a newer value, instead of quietly re-enabling the feature after a timer. And an old enable does expire, back to the safe default, because a value fetched a month ago says little about today. The time to live is a product decision per capability, not a constant shared by all of them.
If values are targeted by account or cohort, scope the cache to the account too. A cached decision made for the previous user of a shared tablet is not evidence for the current one.
Read it outside the code path it protects
This is the requirement most switches fail. Controlled Change puts it in one line: "A kill switch that depends on the broken code path or an unavailable config service is not containment."
- A crash at launch. If the app crashes before the configuration fetch completes, a newly fetched value never takes effect. Read the cached value before you initialize the feature or its SDK, and keep that read out of any code the feature owns.
- A broken network path. A switch fetched over the same host, certificate or pinning setup that is failing cannot arrive. Know which path your configuration travels on.
- The feature module itself. A switch evaluated inside the screen it protects only works if the screen gets far enough to ask.
- Background work. Turning off the UI does nothing for an upload queue, a sync job or a scheduled task that was already enqueued.
The last one appears in the book's failure catalog as its own entry: a kill switch that disables the UI but not the backend or background action. Check the switch where the work runs, not only where it starts:
enum class Capability { ENABLED, DISABLED }
data class PendingUpload(val operationId: String, val payload: ByteArray)
interface UploadStore {
fun nextPending(): PendingUpload?
fun markSent(operationId: String)
}
class UploadDrainer(
private val store: UploadStore,
private val send: (PendingUpload) -> Boolean,
// Read from the local cache, not from the upload code it protects.
private val killSwitch: () -> Capability,
) {
// Checked before every item, because a background drain can outlive the
// screen that started it. A kill pauses the queue; it never deletes it.
fun drain(maxItems: Int): Int {
var sent = 0
while (sent < maxItems) {
if (killSwitch() == Capability.DISABLED) break
val next = store.nextPending() ?: break
if (!send(next)) break // stays pending, retried later with the same id
store.markSent(next.operationId)
sent++
}
return sent
}
}The pending items survive the kill with their operation ids, so when the switch is lifted they resume without duplicates, provided the backend treats the id as an idempotency key.
Snapshot per journey, check per action only when it is safe
Configuration refreshes in the background. If the switch is read live on every screen, a refresh can move a user from the new checkout to the old one between the payment sheet and the confirmation. The book's answer is an evaluation scope per decision: per launch for broad experience settings, per journey for anything transactional, per action only when immediate operational control is required and safe.
A kill switch sits in tension with that rule, because the point of a kill is to act now. Resolve it explicitly: new journeys read the latest value at start; journeys already in progress either finish under the snapshot they started with, or reach a defined safe stopping point. Which of the two is right depends on the harm. A broken coupon display can finish its journey. A flow that is charging twice should stop at the next step, and the server should refuse the second charge regardless of what the client decided.
Reach: the switch only stops the versions that read it
A client switch has two limits on reach: how fast a running app picks up a new value, and which installed versions read the switch at all.
| Lever | Reaches | Does not reach |
|---|---|---|
| Client kill switch, polled | Versions that read it, after their next fetch | Versions shipped before the switch existed |
| Client kill switch, real-time channel or fetch on foreground | The same versions, in minutes instead of hours | Devices that are offline |
| Server-side refusal of the operation | Every version that calls the server | Work that never leaves the device |
| Halting the phased or staged rollout | Devices that have not updated yet | Devices that already have the build |
| Patched binary, then minimum version | Eventually, everyone who updates | Anyone this week |
Polling intervals are often longer than people assume. Firebase Remote Config, for example, defaults to a minimum fetch interval of twelve hours; a team relying on the default has a switch that takes up to half a day. Shorten the interval for a risky rollout, or use a real-time update channel, and accept the cost in requests and battery as the price of a fast stop.
For anything that touches the backend, pair the client switch with a server-side check. The server can refuse an operation from every version, including the ones that shipped before the switch existed. The client switch then becomes the polite version of that refusal: it stops the attempt before the user spends time on it.
Who may pull it, and how you know they did
A kill switch concentrates privileged control, so treat it like one:
- Least privilege. A named owner and the on-call role can pull it. A product dashboard used for experiments is not where an incident control should live next to growth tests.
- Audit. Every change records who, when, which value, which cohort and why.
- Observable. Release and incident teams can see the current state of every switch without asking the team that built it, ideally on the same dashboard as the release.
- Recorded with the decision. The app logs which policy version it evaluated with the operation, so support can explain what a user saw after the value changed.
- Inventoried. A release owner keeps a list of flags and kill switches, with the owner and the safe default of each.
Treat the value itself as untrusted input. Validate its type and range, and never let a remote value supply arbitrary class names, URLs or code paths. A kill switch is a boolean or a small enum, and it should stay one.
Drill it before the incident
The most common failure is the simplest: the switch was never exercised. Controlled Change says why that matters: "A control used only during an incident is likely to be misunderstood when it is needed."
A drill answers the questions an incident will ask: who has the permission today, where the control lives, how long until a running app picks up the change, what the user sees, what happens to work in progress, and whether the feature comes back cleanly when the switch is lifted. Run it in non-production and in a controlled production cohort, and record the time from decision to containment.
Give that time a budget. The book's fictional reference app, Atlas, writes release safety as a characteristic with a budget: a harmful capability can be stopped within ten minutes of the decision, with a kill-switch drill as the evidence. Your number will differ. The point is that it is a number, it is measured, and it is measured again after the configuration vendor, the networking stack or the launch sequence changes.
What a kill switch cannot undo
Be precise about the limits, in design reviews and in incidents. A client kill switch cannot:
- remove a binary from devices that installed it;
- undo data the feature already wrote, locally or on the server;
- recall requests already sent or charges already made;
- reach a version that never read the switch, or read it after the crash point;
- fix a local database migration that already ran.
Each of those needs its own lever: a server-side refusal, a data repair procedure, a patched binary, a migration that old code can still read. A team that says "we have a kill switch" as its whole containment plan has one lever and a list of incidents it cannot touch. For how containment reads as an incident story, see the production incident interview question; for kill switches inside a larger rollout, see the mobile architecture migration interview question.
A checklist before the switch ships
- Off is defined: hide, route back, read-only, degrade or block, and what happens to work in progress.
- The safe default is chosen per capability and written down, and absence of a value never serves the risky path.
- The last validated value is cached, read at launch, scoped to the account when targeted, and only an enable expires.
- The switch is read outside the code path it protects, including background work.
- Propagation time is known, and the backend can refuse the operation for every version.
- Pulling it is least privilege, audited and visible to release and incident teams.
- It has been drilled, with a measured time from decision to containment.
If a reviewer can tick all seven, the switch is a control. If not, it is a boolean with a hopeful name. For how a reviewer reads the rest of the plan around it, see how to review a mobile architecture proposal.
Questions engineers ask about mobile kill switches
What is a kill switch in a mobile app?
A remote control that disables or degrades one risky capability in apps that are already installed, without waiting for a new binary to pass review and reach devices. It is usually a server-controlled configuration value that the app reads, backed by a server-side check for anything the app sends to the backend.
Should a kill switch fail open or fail closed?
It depends on the capability, and the choice has to be written down per switch. A destructive, privileged or newly launched capability fails closed: with no trustworthy value, it stays off. An established critical journey falls back to its last known safe implementation rather than disappearing. What a switch must never do is serve the risky path because a fetch failed.
Can a kill switch roll back a mobile app release?
No. It can stop a behavior on the versions that read it. It cannot take a binary back from devices, undo data the feature already wrote, recall requests already sent, or reach an old version that never read the switch. For those you need a server-side refusal, a data repair, a patched binary, or as a last resort a minimum version.
How fast does a kill switch reach devices?
As fast as your configuration path delivers a new value to a running app, and no faster. A polling client is bounded by its fetch interval; Firebase Remote Config, for example, defaults to a minimum fetch interval of twelve hours unless you change it. Real-time update channels and fetch on foreground shorten that. Measure it in a drill instead of assuming it.
