Interviews
Tell me about a production incident you owned: the senior mobile answer, scored
"Tell me about the worst bug you shipped" is not a debugging question. It is scored on what you did between the first signal and the last change, on a platform where you cannot take a binary back. The structure, the follow-ups and the recovery lines, from the hiring side.
13 min read
Ask a senior mobile engineer about a production incident and you usually get a good debugging story. The crash, the stack trace, the two days of reproduction, the one-line fix. Then I ask: how did you find out? Who was affected while you were looking? What did you do before the fix existed? And the story, which was going so well, stops.
I have run 500+ technical interviews from the hiring side over 15 years in mobile. This question, in all its versions ("the worst bug you shipped", "a time production broke", "an outage you were part of"), is rarely lost on the bug. It is lost on everything around the bug. The bug is the setting. The score is in what you did between the first signal and the last change.
This page is about the incident story itself: what the interviewer writes down, the order that works, and what to say when a follow-up exposes a gap. For the other stories a behavioral round asks for, see the guides to iOS behavioral interview questions and Android behavioral interview questions.
- The question: tell me about a production incident you owned, or the worst bug you shipped.
- The shape of a strong answer: detection, blast radius, containment, cause versus trigger, and what changed afterwards.
- The test underneath: are you calm, methodical and plain about what happened, and does your team work differently because of it.
What the interviewer is writing down
An incident story gives the interviewer five things to score, and they are listening for each one whether or not they ask. Most weak answers cover one of them in depth and skip three.
| Part of the story | What gets written down | What sinks it |
|---|---|---|
| Detection | How the problem became known, and how fast | "We noticed it" with no source, or no curiosity about why support saw it first |
| Blast radius | Who was hurt, how many, and whether it was spreading | No numbers at all, or numbers nobody measured |
| Containment | What stopped the harm before a fix existed | Going straight from "we found it" to "we shipped a fix" |
| Cause | Why the system allowed this, separate from what set it off | "The root cause was a force unwrap" |
| Change | What works differently now, and who uses it | "We added more tests" and "we were more careful" |
Notice what is not on the list: how hard the bug was. A clever diagnosis is pleasant to hear and earns very little on its own.
Choose the incident with your fingerprints on it
The story you pick decides most of the score before you speak. Three rules filter the bank.
- A real decision point. You chose between options under pressure. An incident where the fix was obvious and someone else decided everything is an anecdote.
- Your part is clear. You caused it, contained it, found the cause or built the change afterwards. If the failure was entirely someone else's and your role was to be annoyed, it is a conflict story wearing the wrong costume.
- Something changed. If nothing works differently because of the incident, the story has no ending you can score.
Size is not a rule. A login loop on one OS version, contained in an afternoon, beats a famous outage you watched from a channel.
The structure: from first signal to changed mechanism
The weak structure follows the order you experienced it as an engineer: the bug report, the hunt, the fix, a line about tests. The strong structure follows the order the incident happened to the users and the system. Seven parts, each one or two sentences in the short version.
| Part | What you say | Example sentence shape |
|---|---|---|
| Context | Team, product area, release state. Two sentences. | "I was on the checkout team; version 8 was at 20% of a phased release." |
| Stakes | Why it mattered, in user or business terms | "Orders were failing, and we did not yet know if anyone was charged twice." |
| Role | What you owned, in the first person | "I was on call for the app, so containment was my call." |
| Options | What you could have done, and why not | "A hotfix would take days through review; the server flag would take minutes." |
| Action | What you did first, and how you got others to move | "I paused the release, then asked the payments team to confirm charges." |
| Result | Only numbers you can source | "From the dashboard: new failures stopped within the hour." |
| Learning | A changed behavior or mechanism, not a moral | "Every payment change now ships behind a flag the server can turn off." |
This is STAR with the two parts senior loops care about most added: the options you weighed and the way you moved other people. When a story feels flat in rehearsal, one of those two is usually missing.
Prepare two lengths. About ninety seconds that covers all seven parts thinly and invites questions, and a four-minute version for rounds that are explicitly retrospectives. Never a ten-minute monologue: the interviewer wants to steer, and a story that cannot be interrupted sounds rehearsed.
Detection: how did you find out?
There is no shameful answer here, only an unexamined one. "Support tickets told us before any alert did" is a perfectly good sentence, if the next one is what you changed so that telemetry tells you first next time. What does not work is a vague "we noticed".
On mobile, add one thing. What the app records is frozen in a binary users run for months, so you cannot add a missing dimension during the incident. Name the ones that let you localize the problem: app version, OS version, device class, rollout cohort. If they were missing, say so, and say you added them.
Blast radius: who, how many, and is it getting worse
Before the cause, the interviewer wants to know whether you knew the size of the problem: which users and flows, how many, and whether it was still growing. On mobile the last one is precise: exposure is the share of users on the new version, and it grows every hour the release keeps rolling out.
Then the question that separates levels: was this visible failure, or damage to data or money? A screen that shows an error is a bad day. A checkout that charges twice, or a migration that quietly drops records, is harm that outlives the fix. Candidates who raise it early read as people who have run incidents.
A useful frame to say out loud: some actions limit new exposure, and some recover harm already done. Stopping the rollout is the first kind. It does nothing for the users who already have the broken version, or for the records it already wrote. A story that ends at "we stopped the rollout" has only finished half the incident.
Containment when you cannot roll back a binary
This is where interviewers hear whether you have shipped apps or only services. A bad server deploy is reverted in minutes. A bad app version stays on every device that installed it until a fix is built, reviewed, released and adopted. On mobile, "we rolled back" is almost always the wrong sentence.
What you can do, from cheapest and most reversible to most expensive:
| Lever | What it does | What it cannot do |
|---|---|---|
| Pause the phased release or halt the staged rollout | Stops the version reaching more users. App Store phased release spreads an update over seven days to users with automatic updates; it can be paused. | Remove the version from devices that already have it |
| Server-side change or remote flag | Turns off the broken behavior in minutes, for every installed version that reads the flag | Help a version that never shipped the flag, or one where the broken path runs before the flag is read |
| Fix forward on the normal release | Repairs the cause with full testing, once the flag has contained the harm | Arrive quickly |
| Hotfix with expedited review | A minimal fix, faster, when the harm justifies it | Guarantee a date; expedited review is a request, not a promise |
| Forced update | Blocks the app until the user updates | Exist unless the check shipped long before; it turns your emergency into every user's emergency |
You do not need to recite this table. Your story should show you walked it in that order: stop the spread, contain with the fastest safe lever, then choose the fix path as a decision rather than a reflex. One more sentence lifts a senior answer: any server-side change had to be safe for the broken version, the previous one, and older versions that still had real users.
Root cause versus trigger
"The root cause was a missing null check" is the most common weak line in this round. A missing null check is where the code broke. It is not why the incident happened.
Separate three things. The trigger is what happened that day: an OS update, a server change, users upgrading from a version three releases old, a user re-enrolling a fingerprint. The cause is why the system let that trigger hurt users: the old upgrade path was never tested, the code treated the error as impossible, no gate measured the failing flow early in the rollout. Sometimes there is also an amplifier: the part that turned a small failure into a large one, such as every client retrying at the same instant.
Triggers will happen again in a different form. Causes are what you can fix so the next trigger does nothing. A candidate who says "the trigger was the OS update; the reason it reached users was that our crash gate only looked at the overall rate, not at new crash clusters on one OS version" has given the interviewer a sentence they can quote in the debrief.
What changed afterwards: a mechanism, not a resolution
The fix is a resolution. It closes this incident. What the interviewer wants is the mechanism: something that keeps working when you are not in the room, so this class of incident cannot happen the same way again.
- A rollout gate that now watches the signal that would have caught it, with a written halt rule.
- A test fixture for the upgrade path that broke, not just the fresh install.
- A remote flag around every change of that kind, decided before release, not during the incident.
- An alert on the specific error class, so detection does not depend on support tickets.
- A review checklist item or a lint rule, so other teams stop writing the same bug.
This is where level shows most clearly. One person and one fix caps at senior, however hard the bug was. Other people and a mechanism read staff. For senior roles, do not inflate: a solo fix told plainly beats imaginary organizational influence, and the follow-ups expose inflation in under a minute.
And keep the retrospective blameless in the telling. "The engineer who wrote it was careless" costs you more than the incident did. Blame-first reviews teach teams to hide the next incident for longer.
Running the channel while the system is still failing, and a postmortem that ends at a control and its owner instead of at a person, is a chapter in Before the Room.
The same incident, told weak and strong
An illustrative incident, a composite written for this page: a new app version migrates its local database, and the migration fails on devices upgrading from a version several releases old. Those users open the app to an empty list.
| Part | Weak | Strong |
|---|---|---|
| Opening | "We had a nasty migration bug." | "Version 30 was at 15% of a staged rollout when support reported users with empty lists." |
| Detection | "Users complained." | "Support saw it before our alerts did, because migration failures were not counted anywhere." |
| Blast radius | "Some users." | "Only users upgrading from version 27 or older. I asked first whether the records were gone or only hidden." |
| Containment | "We fixed it and released." | "I halted the rollout and checked that the old records were still on the device before anyone touched a fix." |
| Cause | "An unhandled case in the migration." | "The trigger was an old record shape. The cause was that our upgrade tests started from the previous version only." |
| Change | "We added tests." | "Upgrade tests now start from every version with meaningful users, and the release gate counts migration failures." |
The strong column does not say more about the bug. It says something about every part of the incident, and read aloud it is still under two minutes.
Mid, senior and staff: what the answers sound like
| Topic | Mid-level | Senior | Staff |
|---|---|---|---|
| Containment | "We shipped a hotfix." | Stopped the rollout, contained from the server, chose the fix path. | Had the levers in place before the incident, and says who decided. |
| Cause | "A null pointer." | Separates trigger from cause. | Names the amplifier and the gate that should have caught it. |
| Change | "More tests." | A specific test, alert or flag. | A mechanism other teams now use, with an owner. |
The staff column is not a bigger incident. It is the same incident with owners, decision rights and a way to know it will not happen again.
The follow-ups interviewers use to push
- "How did you find out?" Tests detection, and whether you changed it.
- "How many users, and how do you know?" Tests whether your numbers have a source.
- "Why didn't you just roll back?" Tests whether you know what can and cannot be undone on mobile.
- "What did you deliberately not do in the first ten minutes?" Tests containment before investigation.
- "Was that the root cause, or what set it off?" Tests trigger versus cause.
- "What would you do differently?" Tests learning as a behavior, not a moral.
- "What changed for the teams that were not yours?" Tests staff scope.
Interviewers usually pull one thread three levels deep: the fact, then the other person's view, then your reflection. A true story survives that easily, because you remember the meeting and the objection. A polished story with no texture fails at the second level.
Recovery lines when the answer wobbles
Every incident story has a weak spot, and a calibrated interviewer will find it. What scores is how you handle it. A few lines that keep the answer honest without sinking it:
- When you said "we" for too long: "Let me separate what I did from what the team did. My part was the containment and the change afterwards."
- When you said "we rolled back": "To be precise, we could not recall the installed version. We paused the release and turned the feature off from the server."
- When asked for a number you did not measure: "I did not measure that at the time. What I can source is the support ticket count, and that is one thing I would instrument now."
- When the cause you named was only the trigger: "That was what set it off. The reason it reached users was that nothing in our release gate looked at that flow."
- When nothing changed afterwards: "The fix shipped and the follow-up did not. If I ran it again, I would leave with an owner and a date for the gate change."
None of these lines make the story better than it was. They make it accurate, and an accurate story with a gap beats a smooth one that falls apart at the second follow-up.
Questions engineers ask about the incident question
How do you answer "tell me about a production incident" in a mobile interview?
Tell it in the order the incident happened to the system, not the order you found the bug: how you found out, who was hurt and how many, what you did first to stop it spreading, what caused it and what only triggered it, and what changed afterwards so it cannot happen the same way again. Say "I" for what you did. Keep the first version to about ninety seconds and let the interviewer pull threads.
What if I have never caused a major incident?
Size does not score. Ownership and what changed afterwards do. A crash on one screen for one OS version, contained in an afternoon, is a strong story if you can describe the detection, the decision and the mechanism you added. An incident you did not cause but contained also works, as long as you say plainly which part was yours.
Should I use STAR for an incident story?
STAR is a fine skeleton, but it leaves out the two parts senior interviewers score most: the options you weighed and how you got other people to act. For an incident, also make sure detection, blast radius, containment, cause and the changed mechanism each get a sentence. A STAR answer that skips containment sounds like a debugging story.
Can I say we rolled back the app?
Be exact. You can pause a phased release on the App Store or halt a staged rollout on Google Play, revert a server change and turn a feature off with remote configuration. You cannot take a binary back from devices that already installed it. Interviewers for mobile roles listen for that difference, and "we rolled back the app" tells them you may not know it.
