Senior Staff Site Reliability Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
NVIDIA is seeking a Senior Staff Site Reliability Engineer to improve the efficiency and performance of on-premises and cloud infrastructure.
The role focuses on globally scaled infrastructure services, reliability, automation, monitoring, capacity planning, and lifecycle management.
Key Skills for This Role
Full Job Posting
Role Overview
NVIDIA is seeking a Senior Staff Site Reliability Engineer to improve the efficiency and performance of on-premises and cloud infrastructure.
The role focuses on globally scaled infrastructure services, reliability, automation, monitoring, capacity planning, and lifecycle management.
What You Will Be Doing
- Lead initiatives to transform the IT Compute Core Team architecture and build new on-premises and cloud service offerings.
- Design, scale, and deploy DNS, NTP/PTP, DHCP, and LDAP infrastructure services for global performance and reliability.
- Define service efficiency metrics and drive software and hardware optimizations, including SR-IOV and DPU technologies.
- Use eBPF and XDP technologies for observability and DDoS mitigation.
- Collect and analyze system data for capacity planning and coordinate enterprise-wide infrastructure changes.
- Develop tools for data collection, analysis, visualization, reporting, alerting, and monitoring.
- Collaborate with leadership, engineers, program managers, and product managers to develop IT products and services.
What We Need To See
- Bachelor’s degree in engineering, computer science, mathematics, or a related field, or equivalent experience.
- At least 15 years of experience in compute platform engineering focused on automation.
- Experience designing and deploying containerization architectures and distributed systems infrastructure.
- Experience evaluating application architectures and identifying containerization opportunities.
- Strong analytical skills with the ability to define and track performance metrics.
- Experience with data analysis, performance profiling, Terraform, and configuration management tools.
- Proficiency in Go and/or Python, Linux operating systems, and kernel internals.
- Experience with bare-metal build infrastructure and network architectures including VLAN, VXLAN, SDN, BGP, and Anycast.
Ways To Stand Out
- Deep understanding of DNS, LDAP, and security infrastructure components.
- Hands-on experience implementing containers and managing DNS or LDAP services at scale.
- Strong understanding of microservices architecture, infrastructure as code, and configuration management tools.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at NVIDIA AI
Solutions Architect - NVIDIA Cloud Partners and Datacentre Infrastructure
Dubai, UAE
NVIDIA is seeking a Solutions Architect to guide cloud partners and strategic customers through the design, deployment, and operation of AI and HPC GPU infrastructure. The role focuses on datacentre power, cooling, MEP i
Senior Sales Account Manager, Government and Smart Spaces UAE
Dubai, UAE
NVIDIA is seeking a Senior Sales Account Manager to grow government and smart-spaces business in the UAE by developing strategic relationships, pipelines, and AI transformation opportunities. The role requires 10+ years
Senior Solutions Architect, InfiniBand and Networking Ethernet - NVIS
Sydney, AUS
NVIDIA is seeking a Senior Networking Solutions Architect to deliver large-scale Ethernet and InfiniBand projects for AI and high-performance computing customers. The role combines customer engagement, network and system
PhD Intern, AI ML in Wireless L1/L2 - Fall 2026
Bengaluru, IND
NVIDIA is seeking a PhD intern to develop and optimize AI/ML functions for wireless Layer 1 and Layer 2 software in its Aerial RAN team. The work includes signal-processing research, model architecture selection, trainin
System Software Engineer - Local AI
Pune, IND
NVIDIA is seeking a Systems Software Engineer to build and optimize high-performance local AI inference software for RTX and DGX systems. The role focuses on inference runtimes, model optimization, system debugging, and
Senior Account Manager Telecommunications
Dubai, UAE
The employer is seeking a senior account manager to grow NVIDIA’s telecommunications business through market strategies, executive relationships, strategic partnerships, and adoption of AI and accelerated-computing solut
Enterprise Account Manager
Dubai, UAE
NVIDIA AI is seeking a Senior Enterprise Account Manager to grow revenue and market share by managing strategic customers and ecosystem partners. The role requires enterprise and software sales experience, success agains
Senior System Software Engineer
Dubai, UAE
NVIDIA is seeking a Senior System Software Engineer to improve globally distributed authentication and storage services. The role covers backend microservices, distributed systems architecture, testing, API support, and
Solutions Architect - NVIDIA Cloud Partners and Datacentre Infrastructure
Dubai, UAE
Senior Sales Account Manager, Government and Smart Spaces UAE
Dubai, UAE
Senior Solutions Architect, InfiniBand and Networking Ethernet - NVIS
Sydney, AUS
PhD Intern, AI ML in Wireless L1/L2 - Fall 2026
Bengaluru, IND
System Software Engineer - Local AI
Pune, IND
Senior Account Manager Telecommunications
Dubai, UAE
Enterprise Account Manager
Dubai, UAE
Senior System Software Engineer
Dubai, UAE