Expert Site Reliability Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Drive the reliability, availability, scalability, and operational resilience of critical technology services.
Apply software engineering, automation, observability, and reliability engineering practices.
Key Skills for This Role
Full Job Posting
Purpose
Drive the reliability, availability, scalability, and operational resilience of critical technology services.
Apply software engineering, automation, observability, and reliability engineering practices.
Main Duties and Responsibilities
- Define and implement advanced reliability engineering practices across critical technology services.
- Establish and monitor SLIs, SLOs, and reliability targets.
- Design automation to reduce manual operational activities and improve system resilience.
- Develop monitoring, observability, alerting, and incident detection capabilities.
- Lead analysis and resolution of complex production incidents.
- Conduct root-cause analysis and drive corrective and preventive actions.
- Improve system availability, scalability, capacity, and disaster resilience.
- Identify reliability risks and recommend architectural and engineering improvements.
- Drive performance engineering and capacity planning for critical services.
- Provide technical guidance and mentorship on SRE practices.
- Reduce operational toil through automation and engineering approaches.
Qualifications and Requirements
- Bachelor's degree in Computer Science, Software Engineering, IT, or a related field.
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles.
- Strong experience in cloud platforms, Kubernetes, and production environments.
- Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics.
- Hands-on experience with automation, scripting, CI/CD, and Terraform-based Infrastructure as Code.
- Experience in complex incident management, troubleshooting, and Root Cause Analysis.
- Understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery.
- Experience reducing operational toil through automation and driving reliability improvements.
- Strong analytical, problem-solving, and technical leadership skills.
- Experience in banking, FinTech, or payment environments is preferred.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at TAWANTECH
IT Help Desk
Riyadh, KSA
The employer is seeking an IT Help Desk Specialist to provide first-line technical support, manage incidents and service requests, and resolve user issues within service targets. Required qualifications include a relevan
Fixed Assets Accountant
, KSA
The employer is seeking a Fixed Assets Accountant to manage fixed asset accounting, records, capitalization, and financial controls. The role requires experience with asset registers, CAPEX, depreciation, reconciliations
Data Privacy Expert
Riyadh, KSA
The employer is seeking a Data Privacy Expert to lead privacy governance, PDPL compliance, risk management, data subject rights, and regulatory engagement for a bank. The role requires at least five years of dedicated pr
Project Manager - Data Projects
Riyadh, KSA
The employer is seeking an experienced Project Manager to deliver data projects from initiation through completion. The role requires end-to-end project management, stakeholder and technical-team coordination, scope and
IT Help Desk
Riyadh, KSA
Fixed Assets Accountant
, KSA
Data Privacy Expert
Riyadh, KSA
Data Privacy Expert
Riyadh, KSA
Project Manager - Data Projects
Riyadh, KSA
IT Help Desk
Riyadh, KSA
Project Manager - Data Projects
Riyadh, KSA
Fixed Assets Accountant-Banking
Riyadh, KSA