{bc}
linkedin

Senior Staff Site Reliability Engineer

NVIDIA
Bengaluru, IND
Full-time
Mid-Senior
Onsite
Discovered 1 weeks ago
Site Reliability EngineeringIncident managementIncident CommanderDistributed systemsKubernetesDocker
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site Reliability EngineeringIncident managementIncident Commander
Smart Apply

Full Job Posting

Role overview

NVIDIA is hiring a Senior Staff Site Reliability Engineer in India.

This senior individual-contributor role combines real-time incident leadership with hands-on reliability engineering for AI-powered enterprise platforms.

The role sets technical direction, builds automation and observability, and mentors SRE talent.

What you'll be doing

  • Lead major incidents through triage, coordination, decisions, and executive communication across global time zones.
  • Set technical direction for reliability, scalability, and developer-efficiency initiatives across enterprise systems.
  • Design, build, and operate distributed Kubernetes-based and cloud-native systems.
  • Automate incident detection, triage, communication, and remediation with self-healing systems.
  • Improve observability and signal quality to detect issues earlier and reduce alert noise.
  • Lead root cause analysis and turn learnings into systemic fixes, automation, and prevention.
  • Apply LLMs, anomaly detection, and signal correlation to incident operations.
  • Partner with Cloud, Platform, Security, and AI/ML teams on SLOs, error budgets, and architecture.
  • Mentor engineers and raise standards through design reviews and code reviews.

What we need to see

  • 10+ years in Site Reliability Engineering, Production Engineering, Platform Engineering, or Incident Management with technical leadership at scale.
  • BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • Experience as an Incident Commander or major incident response leader in complex, high-availability environments.
  • Deep knowledge of distributed systems, monitoring, SLIs, SLOs, error budgets, capacity planning, and graceful degradation.
  • Strong proficiency in at least one programming language such as Python, Go, or Java.
  • Hands-on experience with AWS, Azure, or GCP and Docker and Kubernetes.
  • Experience with infrastructure as code, including Terraform, AWS CDK, or CloudFormation, and CI/CD pipelines.
  • Linux/Unix, networking, and observability experience with OpenTelemetry, Prometheus, and Grafana.
  • Working knowledge of relational databases, SQL, indexing, and query optimization.
  • Experience or familiarity with AI/ML concepts applied to operational workflows.
  • Strong written and verbal communication skills for executive briefings and technical influence.

Preferred experience

  • Experience building AI-powered incident management or automation platforms.
  • Experience scaling reliability across distributed teams and multi-region AI/ML infrastructure.
  • Contributions to open-source infrastructure or observability projects, conference talks, or the wider SRE community.

About NVIDIA

NVIDIA develops technologies in artificial intelligence, high-performance computing, visualization, and accelerated computing.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at NVIDIA