Mobile architecture
Image loading system design interview: a scrolling feed, answered from the hiring side
"Design the image loading for a scrolling feed" sounds like a library question. It is a memory question: what an image costs once it is decoded, who still wants it, and what you give back when the system asks. A worked answer from the hiring side, for iOS and Android.
12 min read
Ask a mobile engineer to design image loading for a feed and the first answer is usually a library name and "a memory cache and a disk cache". Then I ask: your memory cache holds 100 images. How many bytes is that? Most candidates have never had to say.
I have run 500+ technical interviews from the hiring side over 15 years in mobile. This prompt looks like the easy one in the set, and it is where answers that sound complete turn out to have no numbers in them. A limit of 100 images is not a budget. A hundred thumbnails and a hundred full-resolution photos are different apps.
This is a worked answer for iOS and Android together, because the decisions are the same and only the APIs differ. For the wider platform rounds, see the iOS messaging answer and the Android field app answer.
- The prompt: design the image loading for a social feed with avatars, photo thumbnails and a full-screen viewer, on iOS and Android.
- The shape: requirements per kind of image, decoded size, cache tiers and keys, byte budgets, one load per key, cancellation, cell reuse, prefetch, memory pressure, observability.
- The test underneath: can you put a number of bytes on every decision, and say who still wants the work you are doing.
The 45-minute round, minute by minute
| Minutes | What you do | What the interviewer is checking |
|---|---|---|
| 0 to 5 | Clarify the kinds of image, scroll speed, device range | Do you ask before you build? |
| 5 to 12 | Requirements per kind of image, and the decoded size | Do you know what an image costs in memory? |
| 12 to 20 | Memory and disk tiers, keys, budgets in bytes | Is the cache bounded by something real? |
| 20 to 30 | One load per key, cancellation, cell reuse | What happens during a fast scroll? |
| 30 to 37 | Prefetch and memory pressure | What do you give back, and when? |
| 37 to 45 | Metrics per device class, rollout | Have you run this in production? |
Most candidates spend the first fifteen minutes on HTTP caching headers and the choice of library. Both matter, and both are decisions the round assumes you can make. The points are in the bytes.
Requirements: three kinds of image, three policies
The question that earns the most in the first ten minutes: which images are in this product, and does each one deserve the same treatment? A feed carries at least three, and one rule for all of them is the most common weak answer.
| Image | What the user expects | Policy |
|---|---|---|
| Avatar | Always there, tiny, the same few faces again and again | Small target size, high memory hit rate, long disk lifetime |
| Feed thumbnail | Appears before the row settles during a scroll | Decode at cell size off the main thread, cancel when the row leaves |
| Full-screen photo | Sharp after a tap, zoomable | Load on demand at screen size, keep at most one or two in memory |
Then write down what you will not build: the upload path, video, the CDN itself. And say one sentence about the server: the cheapest decode is the one where the server already sent the size you need. If the CDN can resize by URL parameter, the client asks for its display size and most of this design gets cheaper.
Decoded size is the number that matters
A JPEG is compressed on the wire and on disk. On screen it is a bitmap, and a bitmap costs width times height times bytes per pixel, usually four. The file size tells you nothing about the memory.
| Image | Pixels | Decoded at 4 bytes per pixel |
|---|---|---|
| Camera photo, decoded as is | 4000 by 3000 | 48,000,000 bytes, about 48 MB |
| Same photo for a 120 point cell at 3x | 360 by 360 | 518,400 bytes, about 0.5 MB |
That is close to a hundred times the memory for the same row. So the pipeline decodes at display size: on iOS with ImageIO thumbnails (the code is in why iOS apps feel slow), on Android with ImageDecoder and a target size, or inSampleSize on older paths. Coil and Glide do this for you when they can resolve the target size, so make sure they can: a view or composable with a known size, or an explicit size on the request.
And decode off the main thread. On iOS a UIImage made from data may defer its decode until the first draw, which lands on the main thread in the middle of a scroll. Decoding into the final bitmap on a background task moves that cost to where nobody is waiting for a frame.
Memory and disk: what each tier holds
Two tiers, two different things stored. The memory tier holds decoded bitmaps at display size, ready to draw. The disk tier holds the downloaded bytes, compressed, so one file can be decoded at several sizes. Both sit behind the network, which is the slowest and most expensive tier.
The keys follow from that. The disk key is the URL. The memory key is the URL plus the target pixel size, because a 360 pixel thumbnail and a 1170 pixel full-screen image of the same photo are different bitmaps. Key the memory tier by URL alone and the viewer shows a blurry thumbnail stretched to the screen, or the feed holds a full-screen bitmap in a small cell.
One more decision that removes a whole class of bugs: treat image URLs as immutable. A new avatar gets a new URL. Then a cached entry is never stale, only old, and invalidation becomes eviction. If the server cannot promise that, use the HTTP validators it sends, and say so.
Budgets in bytes, not in images
Now the number. The memory tier gets a budget in bytes, and the budget depends on the device, because a phone with 2 GB of RAM and a flagship with 12 GB cannot afford the same cache. On Android:
class DecodedImageCache(context: Context) {
// A budget in bytes, scaled to the device class. A starting point you measure, not a constant.
private val maxBytes: Int = run {
val am = context.getSystemService(ActivityManager::class.java)
val classBytes = am.memoryClass * 1024 * 1024
if (am.isLowRamDevice) classBytes / 16 else classBytes / 8
}
private val cache = object : LruCache<String, Bitmap>(maxBytes) {
// Count what a bitmap costs, not that it exists.
override fun sizeOf(key: String, value: Bitmap): Int = value.allocationByteCount
}
fun get(key: String): Bitmap? = cache.get(key)
fun put(key: String, bitmap: Bitmap) { cache.put(key, bitmap) }
// Called from Application.onTrimMemory.
fun trim(level: Int) {
when {
level >= ComponentCallbacks2.TRIM_MEMORY_BACKGROUND -> cache.evictAll() // rebuild from disk later
level >= ComponentCallbacks2.TRIM_MEMORY_UI_HIDDEN -> cache.trimToSize(maxBytes / 2) // the user may come back
}
}
}Three decisions are in there. sizeOf counts bytes, so the cache holds many thumbnails or a few large images, whatever fits. The budget scales with the device and is smaller on low-RAM devices. And the cache gives memory back when the app leaves the screen. The other trim levels are deprecated and current Android versions no longer deliver them to apps, so the snippet uses only TRIM_MEMORY_UI_HIDDEN and TRIM_MEMORY_BACKGROUND.
On iOS, the memory tier is usually NSCache with totalCostLimit in bytes and each image's cost set to its decoded size, bytesPerRow times height. Its limit is not strict, and it evicts on its own under memory pressure, which is what a memory tier should do. The disk tier gets its own budget in bytes, with eviction by last access, and a cap the user would accept on a 64 GB phone.
A senior candidate says the budget is a starting point. A staff candidate says how it will be checked: memory high-water mark per device class, measured on the cheapest device the product supports.
One load per image, however many cells ask
Ten cells show the same avatar. If each one starts a download and a decode, you pay ten times for one image. The pipeline keeps a map of work in flight, keyed like the memory tier, and every caller for the same key joins the same task.
The Swift version of this has a known trap: an actor alone does not prevent the duplicate, because it is reentrant across the await. That follow-up is worked in the senior iOS follow-up questions. Here the interesting part is the next one.
Cancel the interest, not the download
The user flings past 200 rows in a second. Every row that appeared asked for an image, and almost none of them will be on screen when it arrives. Without cancellation, the network and the decoder work through a queue of images nobody will see, and the rows the user stopped on wait behind them.
So a cell cancels its request when it is reused. On iOS that is prepareForReuse cancelling the cell's task; on Android it is the request bound to the view or composable lifecycle, which Coil and Glide do for you. But deduplication makes this harder: if three cells share one load and one of them is reused, cancelling the load breaks the other two. The load must stop only when the last interested cell leaves.
struct ImageKey: Hashable, Sendable {
let url: URL
let maxPixelSize: Int // the display size is part of the key
}
actor ImageLoader {
private struct Entry {
let task: Task<CGImage, Error>
var waiters: Set<UUID>
}
private var inFlight: [ImageKey: Entry] = [:]
private let load: @Sendable (ImageKey) async throws -> CGImage
init(load: @escaping @Sendable (ImageKey) async throws -> CGImage) {
self.load = load
}
func image(for key: ImageKey) async throws -> CGImage {
let id = UUID()
let task = join(key, id)
defer { leave(key, id) }
return try await withTaskCancellationHandler {
try await task.value
} onCancel: {
Task { await self.leave(key, id) } // this cell stopped caring
}
}
private func join(_ key: ImageKey, _ id: UUID) -> Task<CGImage, Error> {
if var entry = inFlight[key] {
entry.waiters.insert(id)
inFlight[key] = entry
return entry.task
}
let load = self.load
let task = Task { try await load(key) }
inFlight[key] = Entry(task: task, waiters: [id])
return task
}
private func leave(_ key: ImageKey, _ id: UUID) {
guard var entry = inFlight[key], entry.waiters.remove(id) != nil else { return }
if entry.waiters.isEmpty {
inFlight[key] = nil
entry.task.cancel() // nobody on screen wants it any more
} else {
inFlight[key] = entry
}
}
}Every caller registers an id. Cancelling a caller removes its id; the shared task is cancelled only when no id is left. A cancelled cell may still wait for a result it no longer needs, which is fine, because of the next rule.
Cell reuse: the wrong image in the right cell
A slow load for row 12 completes after the cell has been reused for row 40. If the completion writes into the cell, the user sees someone else's photo for a moment, or for good. Cancellation alone does not prevent this, because a result can arrive in the instant between the reuse and the cancel.
The rule: the cell checks that the result is still for the key it is showing before it draws. Store the current key on the cell, compare on arrival, drop the result if they differ. The memory tier still keeps the image, so the work is not wasted if row 12 comes back.
Prefetch: a bet you must be able to call off
Prefetch starts loads for rows just below the screen so they arrive before the user does. On iOS, UICollectionViewDataSourcePrefetching tells you which index paths are coming and which ones to cancel. On Android, RecyclerView and LazyColumn prefetch item views ahead, and the image loaders offer a preload call you trigger from the list position.
Prefetch is a bet, and a bad bet costs data, battery and memory. So it runs at lower priority than visible images, it covers a distance you measured rather than guessed, and it is cancelled when the user changes direction. The number to watch is wasted prefetch: images loaded ahead that were never displayed.
Memory pressure: shed the tier you can rebuild
When the system asks for memory, give back what is cheap to recreate. Decoded bitmaps are the first to go, because the disk tier can rebuild them. Bytes on disk stay. Work in flight for rows the user is looking at stays.
On iOS, listen for the memory warning and clear the memory tier, and keep the footprint small in the background, because a large backgrounded app is a likely candidate for termination. On Android, onTrimMemory as in the snippet, and remember that bitmap pixels have lived in native memory since Android 8, so a Java heap that looks flat does not mean the images are under control.
Say what you would refuse, too: android:largeHeap as the fix for an image OOM. It moves the limit without bounding the growth.
Observability and rollout
Name the numbers, each split by device class and app version: memory tier hit rate and disk tier hit rate; time from a row becoming visible to its image appearing, at p50 and p95; decode time off the main thread; bytes downloaded per session; cancelled loads and wasted prefetch as shares of all loads; memory high-water mark. Then the one that matters most: terminations for memory, from MetricKit's exit metrics on iOS and ApplicationExitInfo on Android, because your crash reporter never sees them.
Roll a new decode size or budget out behind a remote flag, cheapest devices first, and compare memory terminations against the old path before widening it.
Mid, senior and staff: what the answers sound like
| Topic | Mid-level | Senior | Staff |
|---|---|---|---|
| Size | "Load the image in the background." | Decode at display size; key memory by URL and size. | Asks the server for the display size, and measures decoded bytes. |
| Budget | "A cache of 100 images." | Budgets in bytes, scaled to the device. | Names the memory owner and the high-water mark per device class. |
| Scroll | "Cancel on reuse." | One load per key, cancelled when the last cell leaves; result checked against the key. | Wasted prefetch and cancelled loads as metrics with a target. |
| Pressure | "Clear the cache when memory is low." | Drop decoded bitmaps, keep disk, keep visible work. | Tracks memory terminations by device class and gates rollout on them. |
The staff column is not more components. It is the same pipeline with numbers attached, and a way to know when one of them moved.
The follow-ups interviewers use to push
- "How much memory does your cache use?" Tests byte budgets and decoded size.
- "The server starts sending 4000 pixel originals." Tests decode at display size and the server contract.
- "Ten cells show the same avatar." Tests one load per key.
- "The user flings through 500 rows." Tests cancellation that respects shared loads.
- "A row shows someone else's photo for a moment." Tests the key check on arrival.
- "It crashes only on 2 GB phones." Tests budgets per device class, memory pressure and how you would see it.
Prepare one sentence for each. If a follow-up makes you add a box, your design was missing a decision. If it makes you point at a number you already said, you are having the conversation the round is for.
Questions engineers ask about image loading interviews
How do you answer "design an image loading pipeline" in a mobile system design interview?
Start from the three kinds of image in the product and the policy each one needs. Then decode at the size the image is displayed, keep the memory tier in decoded bytes and the disk tier in downloaded bytes, set both budgets in bytes per device class, share one load among every cell that wants the same image, cancel it when the last one leaves, and say what you drop first when the system asks for memory back.
Should I just say I would use Coil, Glide, Kingfisher or Nuke?
Name the library you would ship, then design as if you had to build it. The interviewer is scoring whether you know what the library decides for you: the decode size, the cache keys, the budgets, the cancellation and the eviction under pressure. A library with the wrong target size still decodes the wrong number of pixels.
Is NSCache an LRU cache?
Do not promise that. Apple documents NSCache as evicting objects automatically when memory is tight and treats totalCostLimit as a limit it may not enforce strictly. It is a good memory tier precisely because it gives memory back under pressure. If you need a strict order of eviction, you build it yourself.
How should I practise this round?
Take one feed with avatars, thumbnails and a full-screen viewer, answer out loud against a 45-minute timer, then change one fact: the device has 2 GB of RAM, the user flings through 500 rows, or the server starts sending 4000-pixel originals. Answer again. The second pass shows which parts of your design were budgets and which were guesses.