Senior Staff Site Reliability Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
NVIDIA is hiring a Senior Staff Site Reliability Engineer in India.
This senior individual-contributor role combines real-time incident leadership with hands-on reliability engineering for AI-powered enterprise platforms.
The role sets technical direction, builds automation and observability, and mentors SRE talent.
Key Skills for This Role
Full Job Posting
Role overview
NVIDIA is hiring a Senior Staff Site Reliability Engineer in India.
This senior individual-contributor role combines real-time incident leadership with hands-on reliability engineering for AI-powered enterprise platforms.
The role sets technical direction, builds automation and observability, and mentors SRE talent.
What you'll be doing
- Lead major incidents through triage, coordination, decisions, and executive communication across global time zones.
- Set technical direction for reliability, scalability, and developer-efficiency initiatives across enterprise systems.
- Design, build, and operate distributed Kubernetes-based and cloud-native systems.
- Automate incident detection, triage, communication, and remediation with self-healing systems.
- Improve observability and signal quality to detect issues earlier and reduce alert noise.
- Lead root cause analysis and turn learnings into systemic fixes, automation, and prevention.
- Apply LLMs, anomaly detection, and signal correlation to incident operations.
- Partner with Cloud, Platform, Security, and AI/ML teams on SLOs, error budgets, and architecture.
- Mentor engineers and raise standards through design reviews and code reviews.
What we need to see
- 10+ years in Site Reliability Engineering, Production Engineering, Platform Engineering, or Incident Management with technical leadership at scale.
- BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- Experience as an Incident Commander or major incident response leader in complex, high-availability environments.
- Deep knowledge of distributed systems, monitoring, SLIs, SLOs, error budgets, capacity planning, and graceful degradation.
- Strong proficiency in at least one programming language such as Python, Go, or Java.
- Hands-on experience with AWS, Azure, or GCP and Docker and Kubernetes.
- Experience with infrastructure as code, including Terraform, AWS CDK, or CloudFormation, and CI/CD pipelines.
- Linux/Unix, networking, and observability experience with OpenTelemetry, Prometheus, and Grafana.
- Working knowledge of relational databases, SQL, indexing, and query optimization.
- Experience or familiarity with AI/ML concepts applied to operational workflows.
- Strong written and verbal communication skills for executive briefings and technical influence.
Preferred experience
- Experience building AI-powered incident management or automation platforms.
- Experience scaling reliability across distributed teams and multi-region AI/ML infrastructure.
- Contributions to open-source infrastructure or observability projects, conference talks, or the wider SRE community.
About NVIDIA
NVIDIA develops technologies in artificial intelligence, high-performance computing, visualization, and accelerated computing.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at NVIDIA
Senior System Software Engineer - Halos Core and Robotics Platform
Bengaluru, IND
Software Security Compiler Engineer
Bengaluru, IND
Senior System Software Engineer - Halos Core and Robotics Platform
Bengaluru, IND
Accounts Payable Accountant
Bengaluru, IND
Accounts Payable Accountant
Bengaluru, IND
Architect - GPU Performance
Bengaluru, IND
Software Platform Support Engineer - GPU Cloud
, USA
Senior Solution Architect, AI Infrastructure
, USA
Senior Software Engineer, AI Agent Compute
, USA