Associate Site Reliability Engineer II
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
MetLife is seeking a Site Reliability Engineer to ensure the reliability, availability, and performance of critical applications and platforms.
The role monitors production systems, responds to incidents, improves observability, maintains runbooks, and automates operational tasks.
The engineer works with engineering, cloud, and infrastructure teams on operational readiness and service reliability.
Key Skills for This Role
Full Job Posting
Role Overview
MetLife is seeking a Site Reliability Engineer to ensure the reliability, availability, and performance of critical applications and platforms.
The role monitors production systems, responds to incidents, improves observability, maintains runbooks, and automates operational tasks.
The engineer works with engineering, cloud, and infrastructure teams on operational readiness and service reliability.
Core Responsibilities
- Monitor service health, dashboards, alerts, and key reliability indicators.
- Respond to alerts, support bridge calls, execute runbooks, communicate status, and escalate when required.
- Create and maintain dashboards, log queries, telemetry checks, alert validation, and actionable monitoring signals.
- Document operational procedures, update recovery steps, validate readiness, and support knowledge sharing.
- Automate repetitive checks, data collection, remediation, reporting, and toil reduction.
- Support root cause analysis, postmortems, and corrective or preventive action closure.
- Identify alert noise, toil, monitoring gaps, and preventive improvements.
- Support SLOs, SLIs, SLAs, error budgets, operational readiness reviews, and production support standards.
- Use or improve AI-assisted tools for anomaly detection, incident correlation, root cause hints, and knowledge retrieval.
- Collaborate with engineering, infrastructure, cloud, and application teams on service performance.
Skills and Experience
- Foundational knowledge of Linux, networking, application support, cloud operations, production operations, and ITIL-style incident and change processes.
- Ability to use Python, PowerShell, Bash, or equivalent scripting for automation, diagnostics, evidence collection, and reporting.
- Experience with Git, CI/CD basics, ServiceNow or equivalent ticketing, and observability tools is relevant.
- Exposure to Azure services, Docker, Kubernetes, and hybrid cloud operations is required.
- Understanding of SLIs, SLOs, SLAs, error budgets, alerting, incident response, postmortems, and operational runbooks is required.
- Hands-on SQL skills are needed for operational diagnostics, data validation, and service health checks.
- Ability to use AI-assisted investigation and correlation tools responsibly while validating evidence.
Minimum Qualifications
- At least 2 years of experience in production support, DevOps, infrastructure, cloud operations, or software engineering.
- Experience supporting business-critical systems and working in incident, problem, and change management processes.
- Ability to automate standard operational tasks using Python, PowerShell, Bash, or equivalent.
- Bachelor’s degree in computer science, engineering, or equivalent practical experience.
- Exposure to regulated enterprise, insurance, banking, or financial services environments is preferred.
- Business proficiency in English is required; Japanese language skills are a plus.
Preferred Exposure
- Exposure to hybrid cloud platforms involving on-premises and Azure-hosted services.
- Exposure to ELK or Elastic, Grafana, Prometheus, Splunk, Azure Monitor, or Azure Application Insights.
- Exposure to GitHub, Azure DevOps, pipelines, repositories, and operational change controls.
- Exposure to Kubernetes-based production services and containerized application support.
- Exposure to operational readiness reviews, service health reviews, and toil reduction initiatives.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at MetLife
Assistant Manager- Operations
Noida, IND
The employer is seeking an Assistant Manager to manage operational governance, reporting, compliance monitoring, and business-continuity activities. The role requires stakeholder-management, data-analysis, reporting, com
Insurance Specialist
, UAE
General Information Location Dubai, United Arab Emirates Working Schedule Full-Time Work Arrangement In Office Relocation Assistance Available No Posted Date 15-Sep-2026 Job ID 20638 ### Description and Requirements The
Director AI Tools Deployment at Scale
New York City, USA
HR Business Partner
Cary, USA
Implementation System Analyst II
Tampa, USA
Director - Enterprise Architecture
Cary, USA
Assistant Manager- Operations
Noida, IND
Insurance Specialist
, UAE
Ops Readiness Lead Absence & Disability
, USA
Program Manager - Risk Technology
New York City, USA