Own reliability outcomes for assigned services, ensuring strong instrumentation, actionable alerts, meaningful dashboards, and up‑to‑date runbooks.
Define and implement SLIs and SLOs in partnership with product and engineering teams, and surface reliability performance in regular Cloud Operations reviews.
Identify operational toil and design automation to eliminate repetitive manual work.
Drive continuous improvement initiatives that increase observability, automation coverage, and system resilience.
Participate in the SRE on‑call rotation, progressing from secondary to primary ownership as readiness increases.
Command Sev‑2 and Sev‑3 incidents independently over time, with pairing and coaching from a Senior SRE; act as technical lead during Sev‑1 incidents.
Lead blameless post‑incident reviews and own follow‑up actions through completion.
Partner closely with managed services providers (IBM, HCL) to ensure clean escalation paths from L1/L2 monitoring into SRE ownership.
Design, build, and operate CI/CD pipelines supporting cloud‑native application delivery using tools such as GitHub Actions, CodeFresh, TeamCity, and Octopus Deploy.
Automate infrastructure and platform services using Infrastructure as Code (Terraform preferred).
Contribute to the evolution of Omnicell’s observability platform, including intelligent alerting, ML‑based anomaly detection, and automated diagnostics.
Participate in architecture and launch readiness reviews, bringing a reliability lens to system design.
Help establish reference implementations and “golden paths” that enable product teams to launch services with reliability built in from day one.
Bachelor’s degree in Computer Science, Engineering, or a related technical field.
5+ years of experience in software or platform engineering, including 3+ years in an SRE, DevOps, or reliability‑focused role.
Strong hands‑on experience with at least one major public cloud platform (AWS, Azure, or GCP).
Proficiency in Python or another object‑oriented programming language for automation and tooling.
Production experience with Kubernetes, Docker, and Helm.
Experience implementing Infrastructure as Code using Terraform or similar frameworks.
Working knowledge of modern observability tools across metrics, logs, and tracing.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career