Skip to content
All essays

Staff engineering

Staff engineer promotion packet: the evidence a calibration committee needs

The review that decides your level is mostly over before you walk into it. What a calibration committee can actually read, the evidence ledger that turns staff work into proof while it is happening, and the one field that keeps a promotion packet from reading as inflation.

13 min read

The last chapter of my book Before the Room opens with the sentence engineers argue with most: "In the reviews I have seen, the one that decides your level is mostly over before you walk into it." The evidence that settles a level usually existed, or failed to exist, well before anyone opened the packet.

I have run 500+ technical interviews from the hiring side over 15 years in mobile. The book's preface names the rooms that seat includes: "Debriefs, hiring committees, promotion calibrations, the quiet conversation after the candidate leaves the room." A calibration is a debrief about someone who already works there, and in my reading it runs on the same thing a hiring debrief does: what the people deciding can read.

So this is not a post about writing a better packet in review season. A staff promotion packet is mostly built before anyone asks for it, as a by-product of the work, in a format the book calls an evidence ledger.

  • The problem: six months after the work, memory is thinner than the work it stands for.
  • The tool: one ledger row per claim, with the artifact, the outcome, the scope and the proof you cannot show yet.
  • The habit: an entry when an artifact ships or a decision closes, and a level conversation with your manager while the work is still running, not at review time.

Memory is weak proof

Memory is weak even when everyone means well. Managers forget details. Peers remember the emotional peak of a project and not the mechanism that made it work. Sponsors remember the result and not the decisions that produced it. You remember how hard the work felt more than what changed because of it.

Before the Room puts the cost in one line:

Six months on, "I led the migration" is thinner than the work it stands for, and that thinness is one of the more common reasons I have seen a strong engineer get a flat review they did not see coming.

Staff work makes this worse, not better. Leadership work happens in the seams between teams, where no single manager is watching, and the book observes that seam work is often what a staff case rests on. The work most likely to carry a staff case is the work least likely to be remembered by any one person.

Self-promotion does not fix that, and engineers are right to recoil from it. What fixes it is an engineering habit pointed at leadership work: leave a trace of the mechanism while the work is alive.

What a calibration committee is reading for

Companies run reviews differently, and level definitions differ between them. The lens Before the Room uses is about the scope of a move, not a universal ladder: a senior engineer solves the problem, a staff engineer changes the decision around the problem, and a principal engineer changes the system that keeps producing that decision. The evidence chapter turns that lens on the ledger itself:

LevelWhat the ledger entries record
SeniorProblems solved, and the outcomes that prove it
StaffDecisions changed across teams, with scope and missing proof stated so the claim survives calibration
PrincipalMechanisms other teams now reuse, in seams no single manager can see

Three observations in the chapter come straight from calibration rooms, and they are worth reading as a description of the audience. Lists of activities get discounted quickly, because plenty of plateaued engineers could write the same list. Entries that do not claim what they have not yet earned get believed. And the people in the room often know who did what, so absorbed credit is visible.

If you are still deciding whether your work is staff-shaped at all, start with staff iOS engineer is not senior plus. This post assumes the work exists and asks whether anyone can read it.

The evidence ledger: one row per claim

The trace is a small ledger. Each entry names the situation, the leadership problem inside it, the action you took, the artifact that shows the mechanism, and the outcome that proves something changed. The template in the book's appendix has eleven fields:

FieldWhat it holds
DateWhen the artifact shipped or the decision closed
SituationThe project or context
Leadership problemThe problem inside the project, not the project's name
ActionWhat you did
ArtifactThe object that shows the mechanism: a plan, a record, a table, a postmortem
OutcomeWhat changed
ScopeLocal, team, cross-team, organizational or people, with other people's parts named
Evidence linkSomething a reviewer can open
Level evidence and its conditionThe level this might support, and what has to happen for it to count
Missing proofWhat you cannot show yet
Follow-up dateWhen you will check

The second field is where most packets go wrong. "Retiring the legacy client platform" is a situation. The leadership problem inside it is that six teams have to leave a legacy platform, each team pays the migration cost locally, and the platform owns the long-term benefit. The first is a project anyone on those teams could list. The second is the decision a staff engineer had to change.

Under pressure, the minimum version is one line each time an artifact ships or a decision closes: the problem, the artifact, what changed, and the missing proof.

The field that keeps it honest

Two of the fields keep the ledger from becoming a brag sheet. Scope calibrates the claim. Missing proof names what you cannot show yet, and the book is plain about which one carries more weight: "The missing proof matters most, because it separates evidence from inflation."

It feels backwards to write down what you have not proved in a document meant to get you promoted. From the debrief side it reads the other way. A claim that names its own gap tells the reader where the evidence ends, so the rest of the entry can be taken at face value. The chapter says it directly: "In the calibrations I have sat in, that is the kind of entry people believed, because it did not claim what it had not yet earned."

The level field works the same way. It states a condition instead of a title. The completed example in the appendix reads: "Staff-level evidence, on the condition that the plan is reused and adoption behavior changes without direct push." In the chapter's terms, treat level evidence as a hypothesis you are testing, not a title you are asserting.

Keep it on a cadence, not a mood: add an entry when an artifact ships or a decision closes, and fill the missing-proof field back in when the follow-up signal arrives.

A worked entry: the migration nobody wanted to fund

The book's example comes from its migration chapter, a composite: a staff engineer drives adoption of a new shared client across teams that do not report to her, and each team's arithmetic says wait. The version that becomes ledger evidence is the one where the sponsor funded the waves. Written up, the entry reads:

Six teams needed to leave the legacy client and each paid the cost locally. I wrote the wave plan and, with my director, took the funding question to the portfolio sponsor [link], who funded adoption as portfolio work. Three teams committed to wave one with named capacity. Scope: cross-team. Missing proof: old-path usage after one month.

Look at what it claims and what it does not. It claims a decision: adoption funded as portfolio work, by the person who owns that decision. It claims an outcome a reviewer can check: three teams with named capacity. It does not claim the migration worked, because the one-month old-path usage data does not exist yet. In the appendix version, the scope field also says that the platform lead built the migration tooling and the engineer's part was the plan and the funding decision.

That last line costs nothing and buys a lot. The book's warning is blunt: "A ledger that absorbs a colleague's work into your own is the brag sheet with better formatting".

Two more mobile entries, from the book's other scenes

The book's scenes are composites, and so are these two entries. They are my rewrites of two other scenes from Before the Room into ledger rows, to show what staff evidence looks like when the mobile work is a release path or an incident rather than a migration.

A release path another team owns. Release builds keep failing on a shared pipeline, and each failure waits days for a fix. For a mobile team a failed release build is not just a red badge: it can push a store submission, and every date that sits behind it.

FieldEntry
Leadership problemWhether the platform team treats the mobile build path as a supported product. Its reliability number measures backend deploys, where mobile build failures do not show up
ActionPut a month of release-build failures into an outcome evidence table with a two-option ask, walked it past the platform on-call engineer and the product manager, then asked the platform engineering manager to decide
OutcomeAn owner named for the mobile build path, with a same-day response in release weeks; the next failure was picked up the same day without a chase
ScopeCross-team. The platform on-call engineer showed that some failures came from the mobile team's own signing configuration; those fixes stayed with the mobile team
Missing proofThe same-day response slipped once, in the second release week. Two more release weeks before calling it closed

An incident review that would have ended at a person. A config change for checkout eligibility left eligible users unable to complete saved-card checkout for a sustained window. Checkout owned the customer impact, platform owned the config service, and nobody owned the fallback behavior.

FieldEntry
Leadership problemThe review was tempted to stop at the last visible action, the config change, instead of the conditions that let it reach every customer
ActionRan the systems postmortem so it ended at a condition, a primary control and its owner, with the follow-on gaps ranked beneath it
OutcomePrimary control funded: automated validation that rejects an eligibility config which removes an eligible segment, before rollout
ScopeCross-team. The platform lead owns and ships the validation, because the config service is the platform team's to change
Missing proofThe pipeline rejecting a test config that drops an eligible segment; and a month on, whether the same trigger would still cause the same damage

Notice what the incident entry does not claim: the fix. The fix belongs to the platform lead. What it claims is a changed decision, a review that would have ended at a person ending at a control with an owner. That is the staff line from the book applied to an incident, and it is a claim nobody else in the room can make on your behalf.

Three habits that pretend to be a packet

The chapter names three habits that feel like building a case and are not.

  • Collecting praise. A thank-you message can support a case but should not be its center. It records that people felt good, not that anything changed.
  • Listing activities. "Led meetings, coordinated teams, improved reliability" is easy to inflate and hard to evaluate. Plenty of engineers who stopped growing could write the same line.
  • Waiting for the manager to notice. Good managers notice a great deal. They do not watch the seams between teams, and that is where most staff work happens.

In my experience the third is the most common among strong engineers, and the hardest to see from the inside, because it feels like modesty. The short version of why it fails is in an earlier brief: the committee can only promote what it can read. The book's version is the one I would keep: "The ledger does not make the work louder. It makes it legible across the six months between the outcome and the review."

Calibrate with your manager in month two

The ledger has a second use that matters as much as the packet. Say the level question out loud while the work is happening. The book's wording, for the migration: this may be staff scope if the migration plan becomes a mechanism other teams reuse, so you are tracking the decision artifacts, the adoption evidence and the follow-up outcomes, and you would rather calibrate the claim against that evidence than against memory.

That sentence does two things. It builds the evidence as a by-product of the work. And it surfaces disagreement about your level early, when it is useful, instead of at review time, when it becomes disappointment. If your manager reads the same evidence and sees senior scope, you want to hear that in the second month, while there is still time to change the work.

From the hiring side, this is the same discipline a staff interview tests: separating your contribution from the team's outcome, and naming the artifact that made a decision durable. An engineer who has kept a ledger walks into either room with the answer already written. The line to keep: "I would rather calibrate the claim against the evidence than against memory."

From ledger to packet

When review season arrives, your company's packet form decides the headings. The ledger decides whether there is anything true to put under them. What follows is my own advice for moving from one to the other; it applies the book's fields, but the order is mine.

  • Sort by scope, not by date. Cross-team and organizational entries first. A reader skimming for staff evidence should meet it in the first entries, not the last.
  • One claim per entry, with the artifact linked. The reviewer should be able to open the wave plan, the decision record or the postmortem, not take your word for it.
  • Keep the missing proof in. If the follow-up arrived, the entry is stronger. If it did not, the stated gap is what lets the rest be believed.
  • Keep the scope line. Name who built what. The room often knows already.
  • Cut the rows that only record activity. In the appendix's words, "A ledger of small wins does not add up to scope."

What staff scope looks like across platforms, whatever your stack, is covered in staff engineers change the decisions a team can make. The packet is where that change has to be readable by people who never saw it happen.

Where the ledger stops

The book draws the boundary itself: "A ledger strengthens an honest calibration; it cannot win a closed one." If there is no open level to promote into, or the process is decided on politics, the evidence helps only at the margin, and it is no substitute for a fair review.

It is also not a substitute for the work: the appendix says not to use the ledger that way. A careful ledger of senior-scope work is an accurate senior packet, and that is worth knowing in month two rather than in review season.

Questions engineers ask about staff promotion packets

What goes in a staff engineer promotion packet?

Claims a reader can check. For each one: the leadership problem (not just the project), the action you took, the artifact that shows the mechanism, the outcome that proves something changed, the scope with other people's parts named, a link the reviewer can open, and the proof you cannot show yet. Your company's packet form decides the headings. Whether there is anything true to put under them is decided months earlier, while the work is happening.

Is an evidence ledger the same as a brag document?

No, and the difference is two fields. Before the Room adds a scope line, which calibrates the claim and records what other people did, and a missing-proof line, which names what you cannot show yet. Without those, a running list of wins is the brag sheet the book warns about. With them, it is evidence a committee can believe.

Should I include work I did with other people?

Yes, with the split written down. The book's example entry says the platform lead built the migration tooling and the engineer's part was the plan and the funding decision. A ledger that absorbs a colleague's work into your own reads as inflation, and the people in the calibration room often know who did what.

When should I start collecting promotion evidence?

When the work starts, not in review season. Add an entry when an artifact ships or a decision closes, fill in the missing proof when the follow-up signal arrives, and walk your manager through the ledger in the second month, so any disagreement about your level surfaces while there is still time to change the work.

Related resources

Free resources that put this essay to work: short enough to use this week, and each one works on its own, with or without a book.

Related books

These books take the subject of this essay further, written from the hiring side from 500+ technical interviews. Each book page shows what is inside, who it is for, and a free sample where there is one.

Share this essay