{bc}
indeed

Site Reliability Engineer (SRE)

DICETEK LLC
Dubai, UAE
Contract
Onsite
Discovered 1 weeks ago
Monitoring and alertingPrometheusGrafanaELK/Elastic StackSplunkCloud infrastructure
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Monitoring and alertingPrometheusGrafana
Smart Apply

Full Job Posting

Job Purpose

Ensure the reliability, availability, performance, and scalability of critical applications and infrastructure.

The role requires experience in monitoring, automation, cloud technologies, incident management, and DevOps practices.

Key Responsibilities

  • Monitor and maintain application and infrastructure availability, performance, and reliability.
  • Design and implement monitoring, logging, and alerting solutions.
  • Manage production incidents and perform root cause analysis.
  • Automate operational and deployment processes.
  • Collaborate with development, infrastructure, and DevOps teams to improve resilience.
  • Maintain CI/CD pipelines and Infrastructure as Code.
  • Support containerized and cloud-based infrastructure.
  • Develop scripts and automation tools.
  • Implement monitoring, logging, and distributed tracing solutions.
  • Document operational procedures, incidents, and system configurations.

Technical Skills

  • Monitoring tools include Prometheus, Grafana, and Zabbix.
  • Logging tools include ELK/Elastic Stack and Splunk.
  • APM tools include Dynatrace, AppDynamics, and New Relic.
  • Cloud platforms include AWS, Microsoft Azure, or GCP.
  • Container platforms include Docker, Kubernetes, or OpenShift.
  • CI/CD tools include Jenkins, GitLab CI/CD, or Azure DevOps.
  • Infrastructure as Code tools include Terraform and Ansible.
  • Version control tools include Git, GitHub, or GitLab.
  • Incident management tools include ServiceNow and PagerDuty.
  • Distributed tracing tools include OpenTelemetry and Jaeger.
  • Scripting technologies include Bash, Python, and PowerShell.
  • Relevant platforms include PostgreSQL, Oracle, SQL Server, IIS, Nginx, Apache, and REST APIs.

Qualifications and Experience

  • Bachelor's degree in Computer Science, Information Technology, or a related field.
  • Proven experience as a Site Reliability Engineer, DevOps Engineer, or Production Support Engineer.
  • Strong experience in cloud infrastructure, automation, monitoring, and incident management.
  • Hands-on experience with Kubernetes and containerized environments.
  • Strong troubleshooting and root cause analysis skills.
  • Experience in highly available, large-scale production environments.
  • Excellent communication and collaboration skills.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at DICETEK LLC