- Proven experience in IT operations, application support (2nd/3rd line), or a similar production-facing role with 5+ years of experience
- Demonstrated ownership of incidents end-to-end, from alert handling through RCA and prevention
- Practical experience working within an ITIL framework, including incident, problem, and change management, with 2+ years of experience
- Experience working in Agile delivery environments alongside development teams
- Excellent English communication skills with the ability to explain technical issues to both technical and non-technical stakeholders
- Hands-on experience with Splunk, Apica, and Sysdig for log analysis, monitoring, and alerting
- Strong understanding of Prometheus and Grafana for dashboard analysis and alert tuning
- Practical experience operating services on Kubernetes and Linux CLI, including pod health checks, log analysis, and service restarts
- Experience executing and troubleshooting Jenkins deployment pipelines
- Proficiency with Git version control systems
- Experience querying and troubleshooting relational databases including Oracle and DB2
- Working knowledge of Spring/Hibernate applications, Kafka message flows, and XML/JSON payload analysis
Nice to have:
- Experience with Helm deployments
- Java/J2EE development background
- Operational experience with IBM DataStage
- Scripting skills in Bash or Python for automation
- Experience using Ansible for controlled configuration changes