{bc}
indeed

Site Reliability Engineer (SRE)

DICETEK LLC
Abu Dhabi, UAE
Contract
Onsite
Discovered 3 days ago
Site Reliability EngineeringDevOpsSLIs, SLOs, and error budgetsObservabilityDynatracePrometheus
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site Reliability EngineeringDevOpsSLIs, SLOs, and error budgets
Smart Apply

Full Job Posting

What You Will Be Doing

  • Define and implement SLIs, SLOs, and error budgets for business-critical digital banking services.
  • Build observability using metrics, logs, traces, dashboards, and alerts with Dynatrace, Prometheus, Grafana, and ELK.
  • Use AI-driven insights and anomaly detection to proactively predict and resolve reliability issues.
  • Lead on-call triage, root-cause analysis, blameless postmortems, and actionable incident follow-ups.
  • Improve deployment safety with canary, blue-green, rollout, rollback, and production readiness practices.
  • Optimize microservices reliability, scalability, resilience, capacity, performance, and cost efficiency.
  • Automate operational toil through runbooks, remediation scripts, health checks, and self-healing workflows.
  • Embed reliability gates in CI/CD pipelines and maintain observability, documentation, playbooks, and operational standards.
  • Ensure operational compliance and security alignment with internal controls and regulatory standards.

Experience and Qualifications

  • 5+ years of experience in SRE or DevOps roles supporting large-scale, high-availability systems.
  • Experience in banking, fintech, e-commerce, or other data-intensive digital ecosystems.
  • Bachelor’s degree in Computer Science or equivalent technical experience.
  • Strong Linux and performance troubleshooting experience.
  • Proven Terraform and Infrastructure as Code expertise.
  • Proficiency with Kubernetes and container orchestration in microservices environments.
  • AWS experience is preferred; Azure or GCP exposure is advantageous.
  • Deep knowledge of Dynatrace, Prometheus, Grafana, and the ELK stack.
  • Experience with AI/ML-driven reliability, AIOps, anomaly detection, or predictive alerting.
  • Practical CI/CD experience with GitHub Actions, Jenkins, GitLab CI/CD, or Azure DevOps.
  • Experience with Kafka, RabbitMQ, Redis, Aurora, and RDS databases.
  • Strong scripting or programming skills in Python, Bash, or Go.

The Ideal Candidate

  • Organized, structured, and meticulous in approach.
  • Experienced in cross-functional collaboration and distributed teams.
  • Strong analytical and troubleshooting skills for complex production systems.
  • Calm and composed communicator during high-impact incidents.
  • Proactive problem-solver who drives preventive improvements.
  • Collaborative and adaptable in a fast-paced, regulated environment.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at DICETEK LLC