Reliability and operations interview questions
Keeping a deployment running: failure modes, observability, incidents, capacity and safe changes in someone else’s environment.
- A customer reports that some requests to your API fail, and you have logs and nothing else. Walk through your first hour.
- A customer’s payment webhook was processed twice. Design the fix.
- A downstream service recovers from an outage and immediately falls over again. Why, and how do you prevent it?
- Define service level objectives for an AI assistant a customer depends on. What do you measure?
- How do you roll out a risky change to a system running in a customer’s environment?