{bc}
indeed

Site Reliability Engineer (SRE)

DICETEK LLC
Dubai, UAE
Contract
Onsite
Discovered 1 weeks ago
Site reliability engineeringCloud infrastructureMonitoring and alertingIncident managementAutomationKubernetes
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site reliability engineeringCloud infrastructureMonitoring and alerting
Smart Apply

Full Job Posting

Job purpose

The Site Reliability Engineer will ensure the reliability, availability, performance, and scalability of critical applications and infrastructure.

The role requires experience in monitoring, automation, cloud technologies, incident management, and DevOps practices.

Key responsibilities

  • Monitor and maintain application and infrastructure availability, performance, and reliability.
  • Design and implement monitoring, logging, and alerting solutions.
  • Manage production incidents and perform root cause analysis.
  • Automate operational and deployment processes.
  • Collaborate with development, infrastructure, and DevOps teams to improve performance and resilience.
  • Implement and maintain CI/CD pipelines and Infrastructure as Code.
  • Support containerized environments and cloud infrastructure.
  • Develop scripts and automation tools to reduce manual operational work.
  • Implement observability solutions including monitoring, logging, and distributed tracing.
  • Document operational procedures, incidents, and system configurations.

Technical skills

  • Monitoring tools include Prometheus, Grafana, and Zabbix.
  • Logging tools include ELK or Elastic Stack and Splunk.
  • APM tools include Dynatrace, AppDynamics, and New Relic.
  • Cloud platforms include AWS, Microsoft Azure, and GCP.
  • Container technologies include Docker, Kubernetes, and OpenShift.
  • CI/CD tools include Jenkins, GitLab CI/CD, and Azure DevOps.
  • Infrastructure as Code tools include Terraform and Ansible.
  • Version control tools include Git, GitHub, and GitLab.
  • Incident management tools include ServiceNow and PagerDuty.
  • Distributed tracing tools include OpenTelemetry and Jaeger.
  • Scripting languages include Bash, Python, and PowerShell.
  • Relevant database and web technologies include PostgreSQL, Oracle, SQL Server, IIS, Nginx, Apache, and REST APIs.

Qualifications and experience

  • A bachelor's degree in Computer Science, Information Technology, or a related field.
  • Proven experience as a Site Reliability Engineer, DevOps Engineer, or Production Support Engineer.
  • Strong experience in cloud infrastructure, automation, monitoring, and incident management.
  • Hands-on experience with Kubernetes and containerized environments.
  • Strong troubleshooting and root cause analysis skills.
  • Experience in highly available, large-scale production environments.
  • Excellent communication and collaboration skills.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at DICETEK LLC