Senior Staff Site Reliability Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Senior Staff Site Reliability Engineer role in India focused on AI-powered enterprise platforms.
The position combines real-time incident leadership, hands-on engineering, reliability direction, automation, observability, and AI-assisted tooling.
The role is a senior individual-contributor position that includes mentoring and growing SRE talent.
Key Skills for This Role
Full Job Posting
About NVIDIA
NVIDIA develops technology spanning computer graphics, accelerated computing, GPUs, AI, robotics, and self-driving systems.
The company describes a diverse and supportive environment focused on innovation and high-impact engineering.
Role Overview
Senior Staff Site Reliability Engineer role in India focused on AI-powered enterprise platforms.
The position combines real-time incident leadership, hands-on engineering, reliability direction, automation, observability, and AI-assisted tooling.
The role is a senior individual-contributor position that includes mentoring and growing SRE talent.
What You’ll Be Doing
- Lead major incidents through triage, coordination, decisions, and executive communication across global time zones.
- Set technical direction for reliability, scalability, and developer-efficiency initiatives across multiple teams.
- Design and operate distributed, Kubernetes-based, and cloud-native infrastructure.
- Build incident detection, triage, communication, remediation, and self-healing automation.
- Improve observability, reduce alert noise, and drive root-cause analysis and systemic prevention.
- Apply LLMs, anomaly detection, and signal correlation to incident workflows and decision support.
- Champion AI-assisted engineering practices and partner across Cloud, Platform, Security, and AI/ML teams.
- Define SLOs and error budgets, influence architecture, mentor engineers, and raise standards through design and code review.
What We Need To See
- 10 or more years in Site Reliability Engineering, Production Engineering, Platform Engineering, or Incident Management with technical leadership at scale.
- Bachelor’s or master’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- Incident Commander or major incident response experience in complex, high-availability environments.
- Deep knowledge of distributed systems, monitoring, reliability principles, SLIs, SLOs, error budgets, capacity planning, and graceful degradation.
- Strong proficiency in at least one programming language such as Python, Go, or Java.
- Hands-on experience with public cloud platforms, Docker, Kubernetes, infrastructure-as-code, and CI/CD.
- Strong Linux or Unix, networking, observability, relational database, SQL, indexing, and query optimization knowledge.
- Experience or familiarity with AI/ML concepts applied to operational workflows.
- Excellent written and verbal communication skills for executive briefings and senior stakeholder influence.
Ways To Stand Out
- Experience building incident management or automation platforms using ChatOps, workflow orchestration, alert intelligence, or AI.
- A demonstrable record of reducing mean time to detection or mean time to recovery with measurable results.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at NVIDIA
Senior System Software Engineer - Halos Core and Robotics Platform
Bengaluru, IND
Software Security Compiler Engineer
Bengaluru, IND
Senior System Software Engineer - Halos Core and Robotics Platform
Bengaluru, IND
Accounts Payable Accountant
Bengaluru, IND
Accounts Payable Accountant
Bengaluru, IND
Architect - GPU Performance
Bengaluru, IND
Software Platform Support Engineer - GPU Cloud
, USA
Senior Solution Architect, AI Infrastructure
, USA
Senior Software Engineer, AI Agent Compute
, USA