Senior Staff Site Reliability Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
NVIDIA is seeking a Senior Staff Site Reliability Engineer to improve the efficiency and performance of infrastructure across on-premises and cloud environments.
The role focuses on globally scalable compute infrastructure, automation, reliability, observability, and lifecycle management.
Key Skills for This Role
Full Job Posting
Role Overview
NVIDIA is seeking a Senior Staff Site Reliability Engineer to improve the efficiency and performance of infrastructure across on-premises and cloud environments.
The role focuses on globally scalable compute infrastructure, automation, reliability, observability, and lifecycle management.
Responsibilities
- Lead architecture transformation initiatives for the IT Compute Core Team and build new on-premises and cloud service offerings.
- Design, scale, and deploy DNS, NTP/PTP, DHCP, and LDAP services with automation, monitoring, high availability, capacity planning, and lifecycle management.
- Define efficiency metrics and drive software and hardware optimizations, including SR-IOV and DPU improvements.
- Use eBPF and XDP for observability and DDoS mitigation.
- Analyze system and capacity data and coordinate enterprise-wide implementation plans.
- Develop tools for data collection, analysis, visualization, reporting, alerting, and monitoring.
- Collaborate with technical and product stakeholders to develop IT products and services.
Required Qualifications
- Bachelor’s degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
- 15+ years of compute platform engineering experience focused on automation.
- Experience with containerization architectures and distributed systems infrastructure.
- Experience evaluating architectures and identifying containerization opportunities for scalability, reliability, and efficiency.
- Strong analytical skills for defining and tracking performance metrics.
- Experience with data analysis, performance profiling, Terraform, and configuration management tools.
- Proficiency in Go and/or Python.
- Linux proficiency with kernel internals knowledge.
- Experience with large bare-metal build infrastructure environments.
- Understanding of VLAN, VXLAN, SDN, BGP, and Anycast networking.
Preferred Expertise
- Deep understanding of DNS, LDAP, and security infrastructure components.
- Hands-on experience implementing containers at scale.
- Experience deploying and managing DNS and LDAP services at scale.
- Understanding of microservices architecture, infrastructure as code, and configuration management tools.
Workplace
- The posting indicates a hybrid work arrangement.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at NVIDIA
Senior System Software Engineer - Halos Core and Robotics Platform
Bengaluru, IND
Software Security Compiler Engineer
Bengaluru, IND
Senior System Software Engineer - Halos Core and Robotics Platform
Bengaluru, IND
Accounts Payable Accountant
Bengaluru, IND
Accounts Payable Accountant
Bengaluru, IND
Architect - GPU Performance
Bengaluru, IND
Software Platform Support Engineer - GPU Cloud
, USA
Senior Solution Architect, AI Infrastructure
, USA
Senior Software Engineer, AI Agent Compute
, USA