Skip to content
All essays

Interviews

Tell me about a production incident you owned: the senior mobile answer, scored

"Tell me about the worst bug you shipped" is not a debugging question. It is scored on what you did between the first signal and the last change, on a platform where you cannot take a binary back. The structure, the follow-ups and the recovery lines, from the hiring side.

13 min read

Ask a senior mobile engineer about a production incident and you usually get a good debugging story. The crash, the stack trace, the two days of reproduction, the one-line fix. Then I ask: how did you find out? Who was affected while you were looking? What did you do before the fix existed? And the story, which was going so well, stops.

I have run 500+ technical interviews from the hiring side over 15 years in mobile. This question, in all its versions ("the worst bug you shipped", "a time production broke", "an outage you were part of"), is rarely lost on the bug. It is lost on everything around the bug. The bug is the setting. The score is in what you did between the first signal and the last change.

This page is about the incident story itself: what the interviewer writes down, the order that works, and what to say when a follow-up exposes a gap. For the other stories a behavioral round asks for, see the guides to iOS behavioral interview questions and Android behavioral interview questions.

  • The question: tell me about a production incident you owned, or the worst bug you shipped.
  • The shape of a strong answer: detection, blast radius, containment, cause versus trigger, and what changed afterwards.
  • The test underneath: are you calm, methodical and plain about what happened, and does your team work differently because of it.

What the interviewer is writing down

An incident story gives the interviewer five things to score, and they are listening for each one whether or not they ask. Most weak answers cover one of them in depth and skip three.

Part of the storyWhat gets written downWhat sinks it
DetectionHow the problem became known, and how fast"We noticed it" with no source, or no curiosity about why support saw it first
Blast radiusWho was hurt, how many, and whether it was spreadingNo numbers at all, or numbers nobody measured
ContainmentWhat stopped the harm before a fix existedGoing straight from "we found it" to "we shipped a fix"
CauseWhy the system allowed this, separate from what set it off"The root cause was a force unwrap"
ChangeWhat works differently now, and who uses it"We added more tests" and "we were more careful"

Notice what is not on the list: how hard the bug was. A clever diagnosis is pleasant to hear and earns very little on its own.

Choose the incident with your fingerprints on it

The story you pick decides most of the score before you speak. Three rules filter the bank.

  • A real decision point. You chose between options under pressure. An incident where the fix was obvious and someone else decided everything is an anecdote.
  • Your part is clear. You caused it, contained it, found the cause or built the change afterwards. If the failure was entirely someone else's and your role was to be annoyed, it is a conflict story wearing the wrong costume.
  • Something changed. If nothing works differently because of the incident, the story has no ending you can score.

Size is not a rule. A login loop on one OS version, contained in an afternoon, beats a famous outage you watched from a channel.

The structure: from first signal to changed mechanism

The weak structure follows the order you experienced it as an engineer: the bug report, the hunt, the fix, a line about tests. The strong structure follows the order the incident happened to the users and the system. Seven parts, each one or two sentences in the short version.

PartWhat you sayExample sentence shape
ContextTeam, product area, release state. Two sentences."I was on the checkout team; version 8 was at 20% of a phased release."
StakesWhy it mattered, in user or business terms"Orders were failing, and we did not yet know if anyone was charged twice."
RoleWhat you owned, in the first person"I was on call for the app, so containment was my call."
OptionsWhat you could have done, and why not"A hotfix would take days through review; the server flag would take minutes."
ActionWhat you did first, and how you got others to move"I paused the release, then asked the payments team to confirm charges."
ResultOnly numbers you can source"From the dashboard: new failures stopped within the hour."
LearningA changed behavior or mechanism, not a moral"Every payment change now ships behind a flag the server can turn off."

This is STAR with the two parts senior loops care about most added: the options you weighed and the way you moved other people. When a story feels flat in rehearsal, one of those two is usually missing.

Prepare two lengths. About ninety seconds that covers all seven parts thinly and invites questions, and a four-minute version for rounds that are explicitly retrospectives. Never a ten-minute monologue: the interviewer wants to steer, and a story that cannot be interrupted sounds rehearsed.

Detection: how did you find out?

There is no shameful answer here, only an unexamined one. "Support tickets told us before any alert did" is a perfectly good sentence, if the next one is what you changed so that telemetry tells you first next time. What does not work is a vague "we noticed".

On mobile, add one thing. What the app records is frozen in a binary users run for months, so you cannot add a missing dimension during the incident. Name the ones that let you localize the problem: app version, OS version, device class, rollout cohort. If they were missing, say so, and say you added them.

Blast radius: who, how many, and is it getting worse

Before the cause, the interviewer wants to know whether you knew the size of the problem: which users and flows, how many, and whether it was still growing. On mobile the last one is precise: exposure is the share of users on the new version, and it grows every hour the release keeps rolling out.

Then the question that separates levels: was this visible failure, or damage to data or money? A screen that shows an error is a bad day. A checkout that charges twice, or a migration that quietly drops records, is harm that outlives the fix. Candidates who raise it early read as people who have run incidents.

A useful frame to say out loud: some actions limit new exposure, and some recover harm already done. Stopping the rollout is the first kind. It does nothing for the users who already have the broken version, or for the records it already wrote. A story that ends at "we stopped the rollout" has only finished half the incident.

Containment when you cannot roll back a binary

This is where interviewers hear whether you have shipped apps or only services. A bad server deploy is reverted in minutes. A bad app version stays on every device that installed it until a fix is built, reviewed, released and adopted. On mobile, "we rolled back" is almost always the wrong sentence.

What you can do, from cheapest and most reversible to most expensive:

LeverWhat it doesWhat it cannot do
Pause the phased release or halt the staged rolloutStops the version reaching more users. App Store phased release spreads an update over seven days to users with automatic updates; it can be paused.Remove the version from devices that already have it
Server-side change or remote flagTurns off the broken behavior in minutes, for every installed version that reads the flagHelp a version that never shipped the flag, or one where the broken path runs before the flag is read
Fix forward on the normal releaseRepairs the cause with full testing, once the flag has contained the harmArrive quickly
Hotfix with expedited reviewA minimal fix, faster, when the harm justifies itGuarantee a date; expedited review is a request, not a promise
Forced updateBlocks the app until the user updatesExist unless the check shipped long before; it turns your emergency into every user's emergency

You do not need to recite this table. Your story should show you walked it in that order: stop the spread, contain with the fastest safe lever, then choose the fix path as a decision rather than a reflex. One more sentence lifts a senior answer: any server-side change had to be safe for the broken version, the previous one, and older versions that still had real users.

Root cause versus trigger

"The root cause was a missing null check" is the most common weak line in this round. A missing null check is where the code broke. It is not why the incident happened.

Separate three things. The trigger is what happened that day: an OS update, a server change, users upgrading from a version three releases old, a user re-enrolling a fingerprint. The cause is why the system let that trigger hurt users: the old upgrade path was never tested, the code treated the error as impossible, no gate measured the failing flow early in the rollout. Sometimes there is also an amplifier: the part that turned a small failure into a large one, such as every client retrying at the same instant.

Triggers will happen again in a different form. Causes are what you can fix so the next trigger does nothing. A candidate who says "the trigger was the OS update; the reason it reached users was that our crash gate only looked at the overall rate, not at new crash clusters on one OS version" has given the interviewer a sentence they can quote in the debrief.

What changed afterwards: a mechanism, not a resolution

The fix is a resolution. It closes this incident. What the interviewer wants is the mechanism: something that keeps working when you are not in the room, so this class of incident cannot happen the same way again.

  • A rollout gate that now watches the signal that would have caught it, with a written halt rule.
  • A test fixture for the upgrade path that broke, not just the fresh install.
  • A remote flag around every change of that kind, decided before release, not during the incident.
  • An alert on the specific error class, so detection does not depend on support tickets.
  • A review checklist item or a lint rule, so other teams stop writing the same bug.

This is where level shows most clearly. One person and one fix caps at senior, however hard the bug was. Other people and a mechanism read staff. For senior roles, do not inflate: a solo fix told plainly beats imaginary organizational influence, and the follow-ups expose inflation in under a minute.

And keep the retrospective blameless in the telling. "The engineer who wrote it was careless" costs you more than the incident did. Blame-first reviews teach teams to hide the next incident for longer.

Running the channel while the system is still failing, and a postmortem that ends at a control and its owner instead of at a person, is a chapter in Before the Room.

The same incident, told weak and strong

An illustrative incident, a composite written for this page: a new app version migrates its local database, and the migration fails on devices upgrading from a version several releases old. Those users open the app to an empty list.

PartWeakStrong
Opening"We had a nasty migration bug.""Version 30 was at 15% of a staged rollout when support reported users with empty lists."
Detection"Users complained.""Support saw it before our alerts did, because migration failures were not counted anywhere."
Blast radius"Some users.""Only users upgrading from version 27 or older. I asked first whether the records were gone or only hidden."
Containment"We fixed it and released.""I halted the rollout and checked that the old records were still on the device before anyone touched a fix."
Cause"An unhandled case in the migration.""The trigger was an old record shape. The cause was that our upgrade tests started from the previous version only."
Change"We added tests.""Upgrade tests now start from every version with meaningful users, and the release gate counts migration failures."

The strong column does not say more about the bug. It says something about every part of the incident, and read aloud it is still under two minutes.

Mid, senior and staff: what the answers sound like

TopicMid-levelSeniorStaff
Containment"We shipped a hotfix."Stopped the rollout, contained from the server, chose the fix path.Had the levers in place before the incident, and says who decided.
Cause"A null pointer."Separates trigger from cause.Names the amplifier and the gate that should have caught it.
Change"More tests."A specific test, alert or flag.A mechanism other teams now use, with an owner.

The staff column is not a bigger incident. It is the same incident with owners, decision rights and a way to know it will not happen again.

The follow-ups interviewers use to push

  • "How did you find out?" Tests detection, and whether you changed it.
  • "How many users, and how do you know?" Tests whether your numbers have a source.
  • "Why didn't you just roll back?" Tests whether you know what can and cannot be undone on mobile.
  • "What did you deliberately not do in the first ten minutes?" Tests containment before investigation.
  • "Was that the root cause, or what set it off?" Tests trigger versus cause.
  • "What would you do differently?" Tests learning as a behavior, not a moral.
  • "What changed for the teams that were not yours?" Tests staff scope.

Interviewers usually pull one thread three levels deep: the fact, then the other person's view, then your reflection. A true story survives that easily, because you remember the meeting and the objection. A polished story with no texture fails at the second level.

Recovery lines when the answer wobbles

Every incident story has a weak spot, and a calibrated interviewer will find it. What scores is how you handle it. A few lines that keep the answer honest without sinking it:

  • When you said "we" for too long: "Let me separate what I did from what the team did. My part was the containment and the change afterwards."
  • When you said "we rolled back": "To be precise, we could not recall the installed version. We paused the release and turned the feature off from the server."
  • When asked for a number you did not measure: "I did not measure that at the time. What I can source is the support ticket count, and that is one thing I would instrument now."
  • When the cause you named was only the trigger: "That was what set it off. The reason it reached users was that nothing in our release gate looked at that flow."
  • When nothing changed afterwards: "The fix shipped and the follow-up did not. If I ran it again, I would leave with an owner and a date for the gate change."

None of these lines make the story better than it was. They make it accurate, and an accurate story with a gap beats a smooth one that falls apart at the second follow-up.

Questions engineers ask about the incident question

How do you answer "tell me about a production incident" in a mobile interview?

Tell it in the order the incident happened to the system, not the order you found the bug: how you found out, who was hurt and how many, what you did first to stop it spreading, what caused it and what only triggered it, and what changed afterwards so it cannot happen the same way again. Say "I" for what you did. Keep the first version to about ninety seconds and let the interviewer pull threads.

What if I have never caused a major incident?

Size does not score. Ownership and what changed afterwards do. A crash on one screen for one OS version, contained in an afternoon, is a strong story if you can describe the detection, the decision and the mechanism you added. An incident you did not cause but contained also works, as long as you say plainly which part was yours.

Should I use STAR for an incident story?

STAR is a fine skeleton, but it leaves out the two parts senior interviewers score most: the options you weighed and how you got other people to act. For an incident, also make sure detection, blast radius, containment, cause and the changed mechanism each get a sentence. A STAR answer that skips containment sounds like a debugging story.

Can I say we rolled back the app?

Be exact. You can pause a phased release on the App Store or halt a staged rollout on Google Play, revert a server change and turn a feature off with remote configuration. You cannot take a binary back from devices that already installed it. Interviewers for mobile roles listen for that difference, and "we rolled back the app" tells them you may not know it.

Related resources

Free resources that put this essay to work: short enough to use this week, and each one works on its own, with or without a book.

Related books

These books take the subject of this essay further, written from the hiring side from 500+ technical interviews. Each book page shows what is inside, who it is for, and a free sample where there is one.

Share this essay