{bc}
indeed

Site Reliability Engineer (SRE)

DICETEK LLC
Dubai, UAE
Contract
Onsite
Discovered 3 days ago
Site reliability engineeringCloud infrastructureMonitoring and alertingPrometheusGrafanaELK/Elastic Stack
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site reliability engineeringCloud infrastructureMonitoring and alerting
Smart Apply

Full Job Posting

Job Purpose

Ensure the reliability, availability, performance, and scalability of critical applications and infrastructure.

The role requires experience in monitoring, automation, cloud technologies, incident management, and DevOps practices.

Key Responsibilities

  • Monitor and maintain application and infrastructure availability, performance, and reliability.
  • Design and implement monitoring, logging, and alerting solutions.
  • Manage production incidents and perform root cause analysis.
  • Automate operational and deployment processes to improve reliability and efficiency.
  • Collaborate with development, infrastructure, and DevOps teams to improve performance and resilience.
  • Implement and maintain CI/CD pipelines and Infrastructure as Code.
  • Support containerized environments and cloud-based infrastructure.
  • Develop scripts and automation tools to reduce manual operations.
  • Implement observability solutions including monitoring, logging, and distributed tracing.
  • Document operational procedures, incidents, and system configurations.

Required Technical Skills

  • Monitoring tools include Prometheus, Grafana, and Zabbix.
  • Logging tools include ELK/Elastic Stack and Splunk.
  • APM tools include Dynatrace, AppDynamics, and New Relic.
  • Cloud platforms include AWS, Microsoft Azure, or GCP.
  • Container platforms include Docker, Kubernetes, and OpenShift.
  • CI/CD tools include Jenkins, GitLab CI/CD, and Azure DevOps.
  • Infrastructure as Code tools include Terraform and Ansible.
  • Version control tools include Git, GitHub, and GitLab.
  • Incident management tools include ServiceNow and PagerDuty.
  • Distributed tracing tools include OpenTelemetry and Jaeger.
  • Scripting technologies include Bash, Python, and PowerShell.
  • Relevant systems include PostgreSQL, Oracle, SQL Server, IIS, Nginx, Apache, and REST APIs.

Qualifications and Experience

  • Bachelor’s degree in Computer Science, Information Technology, or a related field.
  • Proven experience as a Site Reliability Engineer, DevOps Engineer, or Production Support Engineer.
  • Strong experience in cloud infrastructure, automation, monitoring, and incident management.
  • Hands-on experience with Kubernetes and containerized environments.
  • Strong troubleshooting and root cause analysis skills.
  • Experience in highly available, large-scale production environments.
  • Excellent communication and collaboration skills.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at DICETEK LLC