{bc}
indeed

Site Reliability Engineer (SRE)

DICETEK LLC
Abu Dhabi, UAE
Contract
Onsite
Discovered 4 days ago
Site Reliability Engineering (SRE)DevOpsSLIs, SLOs, and error budgetsObservabilityDynatrace and Davis AIPrometheus
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site Reliability Engineering (SRE)DevOpsSLIs, SLOs, and error budgets
Smart Apply

Full Job Posting

What You Will Be Doing

  • Define SLIs, SLOs, and error budgets for business-critical digital banking services.
  • Build observability across metrics, logs, traces, dashboards, and alerts using Dynatrace, Prometheus, Grafana, and ELK.
  • Use AI-driven insights and anomaly detection to predict and resolve reliability issues.
  • Lead on-call triage, root-cause analysis, and blameless postmortems.
  • Improve deployment safety with canary, blue-green, rollout, and rollback strategies.
  • Optimize microservices for reliability, scalability, and inter-service resilience.
  • Conduct capacity planning, performance tuning, and resilience testing.
  • Automate operational toil with runbooks, remediation scripts, health checks, and self-healing workflows.
  • Embed reliability gates and validations into CI/CD pipelines.
  • Own the observability and AIOps stack and maintain operational documentation.
  • Ensure operational compliance and security alignment with internal controls.
  • Support critical production systems and provide escalation guidance during major incidents.

Experience and Qualifications

  • 5+ years of experience in SRE or DevOps roles managing large-scale, high-availability systems.
  • Experience in banking, fintech, e-commerce, or other data-intensive digital ecosystems.
  • Bachelor's degree in Computer Science or equivalent technical experience.
  • Strong Linux and performance troubleshooting experience.
  • Proven expertise in Terraform and Infrastructure as Code.
  • Proficiency with Kubernetes and container orchestration in microservices environments.
  • AWS experience is preferred; Azure or GCP exposure is advantageous.
  • Deep knowledge of Dynatrace, Prometheus, Grafana, and the ELK stack.
  • Experience with AI/ML-driven reliability, AIOps, anomaly detection, or predictive alerting.
  • Practical understanding of CI/CD pipelines and related platforms.
  • Experience with Kafka, RabbitMQ, Redis, Aurora, and RDS databases.
  • Strong scripting or programming skills in Python, Bash, or Go.

Ideal Candidate

  • Organized, structured, and meticulous in approach.
  • Experienced in cross-functional collaboration with distributed teams.
  • Strong analytical and troubleshooting skills for complex production systems.
  • Calm and composed communicator who can lead during high-impact incidents.
  • Proactive problem-solver who anticipates issues and drives preventive improvements.
  • Collaborative and adaptable team player in a fast-paced, regulated environment.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at DICETEK LLC