Most backend interview prep you find online gives you definitions. Definitions are not what gets you hired. In a real interview you have somewhere between forty seconds and two minutes per question, you are speaking out loud, and the interviewer is quietly deciding whether you have actually built something or only read about it.
So every answer below is written as something you can say, not something you recite. Roughly interview length. Concrete where a concrete detail proves you have shipped code. Where an example genuinely helps — a query, a snippet, a real failure mode — it is there; where it would only pad the answer, it is not.
Click any question to open it. Use Expand all when you are revising, and keep them closed when you want to test yourself.
How to use this properly: read the answer once, close the toggle, then say it out loud in your own words. If you cannot get through it without opening the toggle again, you do not know it yet. The gap between recognising an answer and producing one under pressure is the entire game.
Tier 1 — Language and Web Fundamentals
This is the screening round. Nobody is trying to trick you here; they are checking that you understand the runtime you write code in every day.
Q1 Walk me through what happens when an HTTP request hits your Python web application.
The request first lands on a reverse proxy — usually Nginx — which terminates TLS, serves any static files itself, and forwards everything else to the application server over a socket. That application server is Gunicorn or Uvicorn, and it is the piece that actually manages worker processes. The worker translates the raw HTTP request into a standard Python structure through WSGI or ASGI and hands it to the framework. The framework runs it through middleware — authentication, CORS, request logging — matches the URL against the router, and calls my view function. My view does the real work: hits the database, maybe reads a cache, serialises the result. The response travels back out through the same middleware chain in reverse, and Nginx writes it to the client.
Why they ask it. They want to know whether you see the whole path or only your view function. Naming Nginx, the worker model, and the middleware chain signals that you have deployed something, not only run manage.py runserver.
If they dig deeper. Expect "so why do you need Nginx at all if Gunicorn already speaks HTTP?" The answer is slow-client protection, TLS termination, static file serving and buffering — so your Python workers are never tied up waiting on a slow mobile connection.
Q2 What is the difference between WSGI and ASGI, and when do you actually need ASGI?
WSGI is the older synchronous interface. One worker handles exactly one request at a time, start to finish — so if that request is waiting three hundred milliseconds on a database call, the worker is simply blocked and doing nothing useful. ASGI is the asynchronous successor. It is built around an event loop, so a single worker can hold hundreds of requests in flight and switch between them whenever one is waiting on I/O. It also supports long-lived connections, which WSGI fundamentally cannot do — WebSockets, server-sent events, streaming responses.
I reach for ASGI when the workload is I/O-bound and high-concurrency: a service that mostly calls other APIs, anything with WebSockets, anything streaming tokens back to a browser. For an ordinary CRUD service that talks to one database, WSGI with a sensible worker count is simpler to reason about and simpler to debug, and I would not migrate just for the label.
Example. Django runs on either. FastAPI and Starlette are ASGI-native. Flask is WSGI, though it can be adapted.
Q3 What is the GIL, and how does it affect the backend code you write?
The Global Interpreter Lock is a mutex inside CPython that allows only one thread to execute Python bytecode at a time. So threads in Python never give you true parallel CPU execution — no matter how many cores the machine has, only one thread is running Python at any instant.
The practical consequence is that it matters far less than people assume for typical backend work. The GIL is released during I/O — network calls, disk reads, database queries — so if the service spends most of its time waiting, threads and async both scale fine. It only bites on CPU-bound work: image processing, large serialisation loops, numeric crunching in pure Python. For those I use multiprocessing, push the work to a Celery worker, or drop into a library like NumPy whose C extensions release the GIL.
If they dig deeper. Mention that Python 3.13 shipped an experimental free-threaded build without the GIL. Saying you are aware of it but would not put it in production yet reads as current and level-headed rather than as showing off.
Q4 List, tuple, set, dictionary — when do you choose each in backend code?
A list when order matters and the contents will change, like a page of results I am assembling before returning. A tuple when the collection is fixed and I want it hashable, so it can be a dictionary key or go into a set; I also use tuples to signal that the value is not meant to be modified. A set when I need membership checks or deduplication, because in is average constant time against a list which is linear — checking whether a user ID sits in a permitted set is always a set. A dictionary whenever I need keyed lookup, which in practice is most of the time: mapping IDs to objects so I can join two result sets in memory instead of issuing one query per row.
Example. The most common real use is collapsing an N+1 into a single query:
users = {u.id: u for u in User.objects.filter(id__in=order_user_ids)}
for order in orders:
order.user = users[order.user_id] # dict lookup, no extra query
Q5 What are decorators, and where have you used one in a real service?
A decorator is a function that takes a function and returns a wrapped version of it, so you can add behaviour around a function without editing the function itself. The @ syntax is sugar for reassigning the name to that wrapped version.
In real services I use them for exactly the cross-cutting concerns you would expect: authentication and permission checks on views, retry-with-backoff around flaky third-party calls, caching a computed result, and timing instrumentation that pushes latency into metrics. The value is that the business logic inside the function stays readable — the concern lives in one place and gets applied by name.
Example.
import functools, time
def retry(times=3, delay=0.5):
def decorator(fn):
@functools.wraps(fn) # keeps __name__ and __doc__ intact
def wrapper(*args, **kwargs):
for attempt in range(times):
try:
return fn(*args, **kwargs)
except TransientError:
if attempt == times - 1:
raise
time.sleep(delay * (2 ** attempt))
return wrapper
return decorator
Mentioning functools.wraps without being prompted is a small detail that lands well — it shows you have debugged a stack trace where the function name had been swallowed.
Q6 Why is a mutable default argument a bug in Python?
Because default arguments are evaluated once, when the function is defined — not on each call. So a mutable default like a list or dictionary is a single shared object living on the function, and every call that mutates it mutates that same object. The symptom is a function that appears to remember data from previous calls, which inside a long-running web worker means one user's data leaking into another user's response.
Example.
def add_item(item, basket=[]): # bug — one shared list, forever
basket.append(item)
return basket
add_item("a") # ['a']
add_item("b") # ['a', 'b'] <- not what anyone expected
The fix is to default to None and build the real value inside the function:
def add_item(item, basket=None):
basket = [] if basket is None else basket
basket.append(item)
return basket
Q7 What is the difference between is and ==?
== compares values by calling __eq__, so it asks whether two objects are equivalent. is compares identity — whether two names point at the exact same object in memory. They answer different questions and are not interchangeable.
In practice I only use is for singletons: is None, is True, is False. Everything else uses ==. The trap is that CPython interns small integers and short strings, so a is b sometimes returns True for equal values and lulls you into thinking it works — then it silently fails for the same code with a larger number or a string built at runtime.
Example. 256 is 256 is True in CPython, while comparing two separately computed values of 1000 can be False. That inconsistency is exactly why the rule is identity only for singletons.
Q8 Explain shallow versus deep copy, with a bug it could cause.
A shallow copy creates a new outer container but the inner objects are still the same references, so mutating something nested changes both copies. A deep copy recursively duplicates everything, so the two are fully independent.
The bug this causes in backend code is usually a configuration or default payload. Say I have a template dictionary for an API request body with a nested options dictionary. I shallow-copy it per request, set a per-user field inside options, and now every subsequent request carries the previous user's option — because all those copies share one nested dictionary.
Example.
import copy
template = {"retries": 3, "options": {"locale": "en"}}
a = template.copy() # shallow
a["options"]["locale"] = "hi"
template["options"]["locale"] # 'hi' <- the original mutated
b = copy.deepcopy(template) # fully independent
The working rule: shallow copy is fine for flat structures; anything nested and shared gets a deep copy, or better, gets rebuilt from a factory function each time.
Q9 How does Python manage memory and garbage collection?
CPython is primarily reference counted. Every object tracks how many references point at it, and the moment that count drops to zero the object is freed immediately. That gives very predictable, prompt cleanup, which is why with blocks and file handles behave so cleanly.
Reference counting alone cannot free reference cycles — object A holding B while B holds A never reaches zero — so there is a second mechanism, a generational cycle collector, that periodically walks tracked container objects and reclaims unreachable cycles. It is generational because most objects die young, so newer objects are scanned more often than long-lived ones.
Where this matters in a backend. Long-running worker processes. A module-level cache that only grows, or objects held alive by a logging handler, will not be collected because those references are legitimately still there — that is a leak in the practical sense even though the collector is working correctly. Every time I have chased growing worker memory, the cause was a reference we had forgotten about, not a GC failure. tracemalloc is the tool for finding it.
Q10 What are context managers and why do you use them?
A context manager guarantees that setup and teardown happen as a pair, even if the code between them raises. It is the with statement, and underneath it is any object implementing __enter__ and __exit__.
The reason I care is that in a backend, the resources that hurt you are the ones that leak slowly: database connections, file handles, locks. A try/finally does the same job, but a context manager makes the guarantee reusable and impossible to forget at the call site. I use them constantly for database transactions, so a failure rolls back rather than leaving a half-written record, and for temporarily acquiring a distributed lock.
Example.
from contextlib import contextmanager
@contextmanager
def timed(label):
start = time.perf_counter()
try:
yield
finally:
metrics.timing(label, time.perf_counter() - start) # runs even on exception
with timed("checkout.total"):
process_checkout(order)
Q11 What is a generator, and when would you use one instead of returning a list?
A generator produces values one at a time and only when asked, holding just its current state in memory rather than the whole sequence. Any function containing yield becomes one.
I use them whenever the dataset is large or unbounded and I do not need random access. Exporting half a million rows to CSV is the standard case: building a list of half a million dictionaries can push the worker into swap or get it OOM-killed, whereas a generator streams rows at roughly constant memory and lets me start writing the response before the query has finished draining. The trade-off is that a generator is single-pass — once consumed it is exhausted — and you cannot take its length or index into it. If I need the data twice, a list is the right answer and I would say so rather than reaching for a generator reflexively.
Example.
def rows(queryset):
for obj in queryset.iterator(chunk_size=2000): # never loads the full result set
yield [obj.id, obj.email, obj.created_at.isoformat()]
Tier 2 — APIs, Databases and Data Modelling
This is where most backend interviews are actually won or lost. The questions look basic; the follow-ups are not.
Q12 What makes an API RESTful?
REST is a set of architectural constraints rather than a specification. The ones that matter day to day: resources are identified by URLs and modelled as nouns, not verbs; HTTP methods carry the semantics, so GET reads and DELETE removes; the server is stateless, meaning every request carries everything needed to serve it and no session state lives in the worker; and responses declare their own cacheability.
In practice, being pragmatic here scores better than being doctrinaire. I aim for predictable resource URLs, correct methods and status codes, consistent pagination and error shapes. Where a genuine action does not map cleanly onto a resource — something like "retry this failed payment" — I will expose it as a sub-resource endpoint rather than contorting the model to stay pure.
Example. GET /orders/1023/items rather than GET /getOrderItems?id=1023. And POST /orders/1023/refunds rather than POST /refundOrder.
Q13 POST, PUT and PATCH — what is the difference, and which are idempotent?
POST creates a new subordinate resource and is not idempotent — calling it five times creates five records, which is exactly why double-clicked checkout buttons cause duplicate orders. PUT replaces a resource wholesale at a known URL and is idempotent: sending the same full representation ten times leaves the same final state. PATCH applies a partial modification and is idempotent only if the patch itself is absolute; setting status to shipped is idempotent, but a patch that says "increment quantity by one" is not.
The reason this matters beyond trivia is retries. Clients, mobile networks and load balancers all retry. Any operation that is not naturally idempotent needs protection — which in practice means an idempotency key.
Example. PUT /users/42 with the complete user object replaces it. PATCH /users/42 with {"email": "[email protected]"} changes only that field and leaves the rest untouched.
Q14 How do you version an API, and why?
You version because you cannot force every client to upgrade at once. Mobile apps in particular stay on old versions for months, so any breaking change without a version is an outage for a slice of your users.
My default is a version in the URL path — /api/v1/orders — because it is visible in logs, trivially routable at the proxy, and obvious to anyone reading a curl command. Header-based versioning is arguably cleaner as a design, but it is harder to debug and easier to get wrong in caches. Whichever you choose, the important discipline is defining what counts as breaking: removing a field, renaming one, tightening validation, or changing a type all break clients. Adding an optional field does not, so most changes should be additive and require no new version at all.
If they dig deeper. Talk about deprecation: announce, add a sunset header, monitor which clients still call the old version, and only then remove it. Interviewers like hearing that you plan the removal, not just the addition.
Q15 When do you return 400 versus 401, 403, 404, 409 and 422?
400 is a malformed request — bad JSON, a missing required parameter, something the server cannot even parse. 401 means not authenticated: there is no valid credential, so log in. 403 means authenticated but not authorised: we know who you are and you still cannot do this. 404 is the resource does not exist — and I will deliberately return 404 instead of 403 when even revealing existence would leak information, for example another tenant's record. 409 is a conflict with current state, like trying to cancel an order that has already shipped. 422 is well-formed but semantically invalid — the JSON parsed, the fields are there, but the email is not a valid email.
The distinction that matters most in practice is 401 versus 403, because getting it wrong sends clients into infinite re-login loops.
Q16 How would you authenticate an API — sessions, JWT, or OAuth?
Server-side sessions with an HttpOnly cookie are the safest default for a browser-based product with a single backend. State lives in Redis, logout is instant because you delete the record, and the token is never exposed to JavaScript so XSS cannot steal it.
JWTs are worth it when you genuinely need stateless verification across many services, or a mobile client where cookies are awkward. The honest trade-off is revocation: a signed token is valid until it expires, so a compromised or logged-out token keeps working. The standard mitigation is short-lived access tokens, around fifteen minutes, plus a long-lived refresh token that is stored server-side and can be revoked.
OAuth2 and OIDC are for delegated access — third-party sign-in, or letting another application act on a user's behalf. I would not build my own OAuth server for a first-party login flow; that is complexity with no payoff.
A line worth saying. Never put anything secret in a JWT payload — it is base64, not encryption. Anyone holding the token can read it.
Q17 What is the N+1 query problem and how do you fix it?
You run one query to fetch a list, then the ORM lazily issues one more query per row when you touch a related object. Fifty orders becomes fifty-one queries. It never shows up locally with ten seed rows and it destroys the endpoint in production.
The fix is to tell the ORM upfront which relations you need. In Django, select_related performs a SQL join and is right for forward one-to-one and foreign-key relations; prefetch_related runs a second query and stitches in Python, which is what you need for many-to-many and reverse relations. SQLAlchemy has the same split with joinedload and selectinload.
Example.
# 1 + N queries
for order in Order.objects.all():
print(order.customer.name) # one query per order
# 2 queries total
for order in Order.objects.select_related("customer").prefetch_related("items"):
print(order.customer.name)
How I catch it. Django Debug Toolbar locally, and a query-count assertion in the test for any list endpoint so a regression fails CI rather than production.
Q18 Explain database indexes. When does an index hurt you?
An index is a separate sorted data structure, usually a B-tree, that lets the database find rows without scanning the whole table. It turns a linear scan into a logarithmic lookup, which is the difference between two hundred milliseconds and two milliseconds on a large table.
They are not free. Every insert, update and delete has to maintain every index on the table, so write throughput drops as you add them. They consume disk and memory. And an index the planner never chooses is pure cost — very low-cardinality columns like a boolean flag usually fall into that bucket, because reading half the table through an index is slower than just scanning it.
What I actually do. Index foreign keys and anything in a frequent WHERE, JOIN or ORDER BY. For multi-column indexes, remember the leftmost-prefix rule: an index on (tenant_id, created_at) serves a query filtering on tenant_id alone, but not one filtering only on created_at. And I confirm with EXPLAIN ANALYZE rather than guessing — the planner is the only opinion that counts.
Q19 What does ACID mean, and what is a transaction isolation level?
Atomicity means the transaction happens completely or not at all. Consistency means it moves the database from one valid state to another, respecting constraints. Isolation means concurrent transactions do not observe each other's partial work. Durability means once it commits, it survives a crash.
Isolation is the one with dials on it. Read Committed, the PostgreSQL default, means you only ever read committed data but the same query can return different results twice within one transaction. Repeatable Read gives you a stable snapshot for the whole transaction. Serializable behaves as if transactions ran one after another, which is the strongest guarantee and the most likely to abort under contention, so your code must be ready to retry.
Example that makes it concrete. Two people book the last seat at the same moment. Under Read Committed both can read one seat remaining and both can book it. The fix is either a stricter isolation level with a retry, or a SELECT ... FOR UPDATE that takes a row lock, or a database constraint that makes the bad state impossible.
Q20 SQL or NoSQL — how do you actually choose?
I start from the data and the access pattern rather than from the technology. If the data is relational, if I need multi-row transactions, and if I will query it in ways I cannot fully predict today, a relational database is the right default. PostgreSQL covers an enormous range of workloads and gives you JSONB when part of the data really is schemaless, so you rarely have to choose one or the other.
I would pick a document store when documents are genuinely self-contained and read as a whole, a key-value store for cache and session data, a wide-column store for very high write volume with well-known access patterns, and a time-series database for metrics.
The honest framing to give. Most teams reaching for NoSQL early are solving a scale problem they do not have yet, and paying for it with joins reimplemented in application code. I would rather start relational, measure, and shard or move a specific hot workload out later when there is evidence.
Q21 How do you run a schema migration with zero downtime?
The rule is that the old code and the new schema must be able to coexist, because during a rolling deploy both versions are serving traffic at once. That forces every breaking change into multiple steps.
Renaming a column is the classic example. Instead of one rename, I add the new column, deploy code that writes to both and reads from the old, backfill in batches, deploy code that reads from the new, and only in a later release drop the old column. Adding a NOT NULL column follows the same shape: add it nullable with a default, backfill, then add the constraint.
The operational detail worth mentioning. On PostgreSQL, adding an index locks writes unless you use CREATE INDEX CONCURRENTLY, and a long migration holding a lock behind a queue of waiting queries will take the site down even though the migration itself looked fast. I also set a lock_timeout so a migration fails quickly rather than stalling the whole table.
Q22 How do you prevent SQL injection?
Use parameterised queries, always. The driver sends the SQL and the values separately, so user input is never parsed as SQL and there is nothing to escape. An ORM does this for you by default, which is most of the reason ORMs are safe.
The place it goes wrong is raw SQL built by string formatting, usually because someone needed a dynamic ORDER BY or a complex report the ORM made awkward. Values always go in as parameters. Identifiers — table and column names — cannot be parameterised, so if a column name really must be dynamic, it gets validated against an allowlist of known columns rather than passed through.
Example.
# unsafe
cursor.execute(f"SELECT * FROM users WHERE email = '{email}'")
# safe — the value never becomes SQL
cursor.execute("SELECT * FROM users WHERE email = %s", [email])
Beyond that: least-privilege database users, so the application role cannot DROP anything, and validation at the edge with something like Pydantic so malformed input never reaches the query layer.
Tier 3 — Scaling, Concurrency and Production
Senior signal lives here. These answers should sound like someone who has been on call.
Q23 Threads, processes or async — how do you choose in Python?
I choose based on what the work is waiting for. If it is I/O-bound — network calls, database queries, file reads — async is the most efficient option, because thousands of waiting coroutines cost almost nothing and there is no thread-switching overhead. Threads also work for I/O and are the pragmatic choice when the libraries I depend on are synchronous and I do not want to rewrite them. If the work is CPU-bound, neither helps because of the GIL, so I use multiprocessing or move it out of the request path entirely into a task queue.
The trap to avoid saying you would fall into. One blocking call inside an async handler stalls the entire event loop for every concurrent request, not just that one. If I have to call a synchronous library from async code, it goes through run_in_executor or asyncio.to_thread.
Concrete numbers help. A service that fans out to five upstream APIs per request goes from roughly the sum of their latencies to roughly the slowest of them, just by switching sequential awaits to asyncio.gather.
Q24 How would you design a caching layer, and how do you handle invalidation?
I start by finding what is actually expensive and frequently read, because caching something cheap just adds a consistency problem for no gain. The common pattern is cache-aside with Redis: the application checks the cache, and on a miss it reads the database, writes the result back with a TTL, and returns it.
Invalidation is the hard half. My default is a short TTL, because it bounds staleness without any coordination and is very hard to get wrong. Where the data must be fresh immediately, I invalidate explicitly on write — delete the key rather than update it, so there is no race between two writers. Where a key is expensive and hot, I use a lock or a single-flight guard so that when it expires, one request rebuilds it instead of a thousand hitting the database at once.
Failure modes worth naming. Cache stampede, when a popular key expires under load. Cache penetration, when repeated misses for a nonexistent key pass straight through — cache the negative result briefly. And always design so that Redis going down degrades latency, never correctness.
Q25 A request needs to do slow work — generate a report, send emails, call a slow third party. How do you handle it?
I get it out of the request cycle. The endpoint validates the input, writes a job record with a status of pending, pushes a task onto a queue, and returns 202 Accepted with an ID the client can poll or a channel it can subscribe to. Celery with Redis or RabbitMQ is the usual choice, or a managed queue if we are already on a cloud provider.
The parts that matter beyond "use Celery": the task must be idempotent, because queues deliver at least once and a worker can die after doing the work but before acknowledging. So the task checks the job record first and exits early if it has already run. Retries use exponential backoff with a cap, and anything that exhausts its retries lands in a dead letter queue instead of vanishing. I also keep the payload small — pass an ID and let the worker load the row, rather than serialising a whole object into the queue where it can go stale.
Q26 How would you implement rate limiting?
At the smallest scale it belongs in the proxy — Nginx or an API gateway can do this without any application code, and that is the right answer when the limit is coarse. When the limit is per-user or per-plan, it has to be in the application, and it has to be in shared state, because with four workers a local counter gives you four times the intended limit.
I use Redis with a sliding window or token bucket. Token bucket is my default because it allows a legitimate short burst while still bounding the sustained rate, which matches how real clients behave. The counter increments atomically — a small Lua script or INCR with an expiry — so concurrent requests cannot race past the limit.
The details that show experience. Return 429 with a Retry-After header so well-behaved clients back off instead of hammering. Key by user ID for authenticated traffic and by IP only as a fallback, since IPs are shared behind NAT. And fail open rather than closed if Redis is unavailable — a rate limiter should not become the thing that takes down the API.
Q27 A payment request times out and the client retries. How do you stop a double charge?
With an idempotency key. The client generates a unique key per logical operation and sends it as a header. On arrival the server tries to insert that key into a table with a unique constraint, in the same transaction as the payment. If the insert succeeds, this is the first attempt and we process. If it violates the constraint, we have seen this operation before, so we look up the stored response and return it unchanged rather than charging again.
The important part is that the uniqueness is enforced by the database, not by a check-then-act in application code. A "does this key exist" query followed by an insert has a race window that two concurrent retries will find, and the retries are simultaneous precisely because a timeout caused them.
Extra credit. Handle the in-flight case explicitly: if the key exists but the original request has not finished, return 409 so the client waits rather than getting a half-formed answer. Idempotency keys should also expire after some window, typically twenty-four hours.
Q28 How do you scale a read-heavy service?
In roughly the order of cheapest to most invasive. First I make sure the queries themselves are not the problem — an EXPLAIN ANALYZE and a missing index have solved more scale problems than any architecture change. Then caching, at whichever layer buys the most: HTTP caching at a CDN for anonymous responses, and Redis for computed results.
After that, read replicas. Reads go to replicas and writes to the primary, which scales reads close to linearly. The catch to name is replication lag: a user who writes and immediately reads can see their own change missing. I handle that by routing reads to the primary for a short window after a write by that user, rather than pretending the lag does not exist.
Beyond that, denormalise the specific hot path — materialised views or a precomputed summary table — and only then consider sharding, which I would treat as a last resort because it makes joins and transactions genuinely painful.
Q29 What does it mean for a service to be stateless, and why does it matter?
Stateless means no request depends on anything held in the memory or on the disk of the specific instance that serves it. Session data, uploaded files, in-process caches and background timers all have to move to shared infrastructure — Redis, object storage, a queue.
It matters because everything else in modern deployment assumes it. Autoscaling adds and removes instances at will. Rolling deploys kill instances mid-traffic. A container can be rescheduled to another node at any moment. If a user's session lives in one worker's memory, all of that translates into random logouts.
The sticky-session point. Sticky sessions are the workaround, and I would treat them as a smell rather than a solution: they break even load distribution, they make deploys disruptive for whoever was pinned to the instance going down, and they hide the actual problem. Moving state out is usually a smaller change than people expect.
Q30 Latency on an endpoint jumps from 100ms to 4 seconds in production. Walk me through what you do.
First I establish scope and timing, because that alone eliminates most causes. Is it one endpoint or everything? All users or one tenant? And what changed at that timestamp — a deploy, a migration, a traffic spike, a third-party incident? I check the deploy log before I check anything else, because the answer is a recent deploy far more often than not.
Then I go layer by layer using what the dashboards already show. Application metrics tell me whether time is going in the database, an outbound call, or CPU. If it is the database, I look at slow query logs and active connections — a missing index on a table that recently grew, or connection pool exhaustion where requests are queueing for a connection rather than doing work. If it is an outbound call, I check whether a dependency is degraded and whether we have a timeout and a circuit breaker on it, because an upstream that is slow rather than down is what usually cascades.
Distributed tracing is what turns this from guesswork into reading a flame graph. Once it is stable, the follow-up is an alert on the signal we missed.
Tier 4 — Scenario, Design and Behavioural
The round where they find out what you are like to work with. These deserve rehearsal more than the technical ones, because most people improvise them and it shows.
Q31 Design a URL shortener.
I would start by clarifying scale and requirements, because that changes the design: roughly how many writes per second, what read-to-write ratio, do links expire, do we need analytics, do we need custom aliases. Assume read-heavy by a wide margin, which is the usual shape.
The core is two operations. To shorten, generate a short key and store the mapping. I would use a counter-based ID encoded in base62 rather than hashing, because it guarantees uniqueness without a collision-retry loop; the counter can come from a database sequence or from ranges handed out to each instance so they do not coordinate per request. To resolve, look up the key and return a 301 or 302 — 302 if we want analytics, since a 301 gets cached by the browser and we stop seeing the traffic.
Scaling is mostly caching. The key-to-URL mapping is immutable, so it caches perfectly; a Redis layer in front of the database absorbs almost all reads. The database itself is a simple key-value shape, so it shards cleanly on the short key if it ever needs to.
The parts they are listening for. Why base62 over a hash, the 301-versus-302 analytics trade-off, and that you noticed the workload is read-heavy and immutable.
Q32 Tell me about a difficult bug you fixed in production.
Answer this in four beats — situation, investigation, fix, and what changed afterwards — and keep the investigation as the longest part, because that is the part they are assessing.
A good shape: the symptom was intermittent and only under load, so it was not reproducible locally. I started from the data rather than a hypothesis — correlated the error timestamps against deploys and traffic, found the errors clustered on one endpoint, and reproduced it by replaying concurrent requests against staging. The cause turned out to be a race condition: two concurrent requests both read a balance, both decided it was sufficient, and both wrote. The fix was a SELECT ... FOR UPDATE inside the transaction so the second request waited. Afterwards I added a database constraint that made the invalid state impossible regardless of application code, plus a concurrency test that would have caught it.
What makes this answer land. Ending on the systemic fix rather than the patch. Anyone can fix a bug; the signal they are looking for is whether you close the class of bug.
Q33 How do you approach code review?
When reviewing, I read for correctness and intent first — does this do what the ticket says, what happens on the error path, what happens when this runs twice, are the edge cases tested. Then design: is this the right place for the logic, will the next person find it. Style comes last and mostly should not be a human's job, so I would rather the team run a formatter and a linter in CI than argue about it in comments.
I try to be explicit about severity, because an unlabelled comment reads as blocking. I will say "this is a blocker" or "non-blocking, take it or leave it", and I ask questions rather than issue instructions when I might be missing context — "what happens if this list is empty?" surfaces the same issue as "this will crash" without putting anyone on the defensive.
When receiving review, I try to treat every comment as information about how readable the code is. If someone misread it, the comment is usually right even when the objection is wrong.
Q34 How do you keep your skills current?
Be specific, because a generic answer here is instantly forgettable. The version that works names actual sources and, more importantly, something you built.
For example: I read release notes for the tools I actually use rather than chasing every new framework, so Python, PostgreSQL and Django release notes, plus a couple of engineering blogs where teams write up real incidents. When something looks genuinely useful I build a small thing with it instead of just reading — the last one was rewriting a slow batch job as an async pipeline, which taught me more about where async does and does not help than any article would have. And I read the source of libraries I depend on when something behaves unexpectedly, which has been the single highest-return habit.
The framing that works. Depth in the stack you use, curiosity about the rest, and at least one concrete thing you made. Avoid listing courses you started and did not finish — interviewers ask follow-up questions.
A note on delivery
The content above gets you through the technical filter. What separates candidates at that point is almost never knowledge, it is delivery:
- Answer the question asked, then stop. Rambling past the answer is the most common way strong candidates weaken a good response. If they want more, they will ask.
- Say "I do not know" cleanly when you do not. Follow it with how you would find out. Interviewers are trying to find the edge of your knowledge — pretending it is further out than it is fails you faster than the gap would have.
- Bring one real detail per answer. A number, a failure you saw, a tool you used. It is the difference between sounding like you read this and sounding like you lived it.
- Ask a clarifying question on anything open-ended. For design questions this is not optional; jumping straight to a solution without scoping is itself a negative signal.