{bc}
indeed

Senior Staff Site Reliability Engineer

NVIDIA
Karnataka, IND
Full-time
Onsite
Discovered 1 weeks ago
Site Reliability EngineeringIncident managementDistributed systemsKubernetesDockerPublic cloud platforms
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site Reliability EngineeringIncident managementDistributed systems
Smart Apply

Full Job Posting

About NVIDIA

NVIDIA develops technology spanning computer graphics, accelerated computing, GPUs, AI, robotics, and self-driving systems.

The company describes a diverse and supportive environment focused on innovation and high-impact engineering.

Role Overview

Senior Staff Site Reliability Engineer role in India focused on AI-powered enterprise platforms.

The position combines real-time incident leadership, hands-on engineering, reliability direction, automation, observability, and AI-assisted tooling.

The role is a senior individual-contributor position that includes mentoring and growing SRE talent.

What You’ll Be Doing

  • Lead major incidents through triage, coordination, decisions, and executive communication across global time zones.
  • Set technical direction for reliability, scalability, and developer-efficiency initiatives across multiple teams.
  • Design and operate distributed, Kubernetes-based, and cloud-native infrastructure.
  • Build incident detection, triage, communication, remediation, and self-healing automation.
  • Improve observability, reduce alert noise, and drive root-cause analysis and systemic prevention.
  • Apply LLMs, anomaly detection, and signal correlation to incident workflows and decision support.
  • Champion AI-assisted engineering practices and partner across Cloud, Platform, Security, and AI/ML teams.
  • Define SLOs and error budgets, influence architecture, mentor engineers, and raise standards through design and code review.

What We Need To See

  • 10 or more years in Site Reliability Engineering, Production Engineering, Platform Engineering, or Incident Management with technical leadership at scale.
  • Bachelor’s or master’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • Incident Commander or major incident response experience in complex, high-availability environments.
  • Deep knowledge of distributed systems, monitoring, reliability principles, SLIs, SLOs, error budgets, capacity planning, and graceful degradation.
  • Strong proficiency in at least one programming language such as Python, Go, or Java.
  • Hands-on experience with public cloud platforms, Docker, Kubernetes, infrastructure-as-code, and CI/CD.
  • Strong Linux or Unix, networking, observability, relational database, SQL, indexing, and query optimization knowledge.
  • Experience or familiarity with AI/ML concepts applied to operational workflows.
  • Excellent written and verbal communication skills for executive briefings and senior stakeholder influence.

Ways To Stand Out

  • Experience building incident management or automation platforms using ChatOps, workflow orchestration, alert intelligence, or AI.
  • A demonstrable record of reducing mean time to detection or mean time to recovery with measurable results.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at NVIDIA