{bc}
linkedin

Expert Site Reliability Engineer

TAWANTECH
Riyadh, KSA
Full-time
Mid-Senior
Onsite
Discovered 2 weeks ago
Site Reliability EngineeringDevOpsCloud platformsKubernetesMonitoring and observabilityService Level Indicators and Objectives
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site Reliability EngineeringDevOpsCloud platforms
Smart Apply

Full Job Posting

Purpose

Drive the reliability, availability, scalability, and operational resilience of critical technology services.

Apply software engineering, automation, observability, and reliability engineering practices.

Main Duties and Responsibilities

  • Define and implement advanced reliability engineering practices across critical technology services.
  • Establish and monitor SLIs, SLOs, and reliability targets.
  • Design automation to reduce manual operational activities and improve system resilience.
  • Develop monitoring, observability, alerting, and incident detection capabilities.
  • Lead analysis and resolution of complex production incidents.
  • Conduct root-cause analysis and drive corrective and preventive actions.
  • Improve system availability, scalability, capacity, and disaster resilience.
  • Identify reliability risks and recommend architectural and engineering improvements.
  • Drive performance engineering and capacity planning for critical services.
  • Provide technical guidance and mentorship on SRE practices.
  • Reduce operational toil through automation and engineering approaches.

Qualifications and Requirements

  • Bachelor's degree in Computer Science, Software Engineering, IT, or a related field.
  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles.
  • Strong experience in cloud platforms, Kubernetes, and production environments.
  • Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics.
  • Hands-on experience with automation, scripting, CI/CD, and Terraform-based Infrastructure as Code.
  • Experience in complex incident management, troubleshooting, and Root Cause Analysis.
  • Understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery.
  • Experience reducing operational toil through automation and driving reliability improvements.
  • Strong analytical, problem-solving, and technical leadership skills.
  • Experience in banking, FinTech, or payment environments is preferred.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at TAWANTECH