- 5+ years of experience in Incident Operations, Site Reliability Engineering, Technical Operations, or a similar role
- Experience working in on-call environments with SLA-driven responsibilities
- Strong understanding of distributed systems and production environments
- Experience with monitoring, alerting, and incident management tools
- Familiarity with APIs, system integrations, and observability platforms
- Hands-on experience with Python or Kotlin
- Understanding of SDLC and production reliability principles
- Strong communication, stakeholder management, and decision-making skills
- Ability to work effectively in high-pressure environments and manage multiple priorities
- Strong ownership mindset and cross-functional collaboration skills
Nice to have
- Experience with Datadog or Chronosphere
- Experience with PagerDuty, Rootly, or Slack workflows
- Experience managing external status pages
- Experience with incident management automation and process improvements
- Experience contributing to reliability tooling and platform engineering