On-site
--
Integrated Solutions Tawantech

Job Details

Job description

Purpose:


To drive the reliability, availability, scalability, and operational resilience of critical technology services by applying advanced software engineering, automation, observability, and reliability engineering practices. 


Main Duties and Responsibilities: 


  • Define and implement advanced reliability engineering practices across critical technology services. 
  • Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets. 
  • Design automation to reduce manual operational activities and improve system resilience. 
  • Develop and enhance monitoring, observability, alerting, and incident detection capabilities. 
  • Lead technical analysis and resolution of complex production incidents. 
  • Conduct root-cause analysis and drive permanent corrective and preventive actions. 
  • Design solutions to improve system availability, scalability, capacity, and disaster resilience. 
  • Identify reliability risks and recommend architectural and engineering improvements. 
  • Drive performance engineering and capacity planning for critical services. 
  • Provide advanced technical guidance and mentorship on SRE practices. 
  • Promote automation and engineering approaches that reduce operational toil and improve service reliability. 

Preferred candidate

Years of experience

No experience required

Degree

Bachelor's degree / higher diploma

Similar Jobs