Job description
Purpose:
To drive the reliability, availability, scalability, and operational resilience of critical technology services by applying advanced software engineering, automation, observability, and reliability engineering practices.
Main Duties and Responsibilities:
- Define and implement advanced reliability engineering practices across critical technology services.
- Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets.
- Design automation to reduce manual operational activities and improve system resilience.
- Develop and enhance monitoring, observability, alerting, and incident detection capabilities.
- Lead technical analysis and resolution of complex production incidents.
- Conduct root-cause analysis and drive permanent corrective and preventive actions.
- Design solutions to improve system availability, scalability, capacity, and disaster resilience.
- Identify reliability risks and recommend architectural and engineering improvements.
- Drive performance engineering and capacity planning for critical services.
- Provide advanced technical guidance and mentorship on SRE practices.
- Promote automation and engineering approaches that reduce operational toil and improve service reliability.
Preferred candidate
Years of experience
No experience required
Degree
Bachelor's degree / higher diploma