{bc}
linkedin

Senior Staff Site Reliability Engineer

NVIDIA AI
Bengaluru, IND
Full-time
Mid-Senior
Onsite
Discovered 1 weeks ago
Site reliability engineeringCompute platform engineeringCloud infrastructureOn-premises infrastructureAutomationContainerization
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site reliability engineeringCompute platform engineeringCloud infrastructure
Smart Apply

Full Job Posting

Role Overview

NVIDIA is seeking a Senior Staff Site Reliability Engineer to improve the efficiency and performance of on-premises and cloud infrastructure.

The role focuses on globally scaled infrastructure services, reliability, automation, monitoring, capacity planning, and lifecycle management.

What You Will Be Doing

  • Lead initiatives to transform the IT Compute Core Team architecture and build new on-premises and cloud service offerings.
  • Design, scale, and deploy DNS, NTP/PTP, DHCP, and LDAP infrastructure services for global performance and reliability.
  • Define service efficiency metrics and drive software and hardware optimizations, including SR-IOV and DPU technologies.
  • Use eBPF and XDP technologies for observability and DDoS mitigation.
  • Collect and analyze system data for capacity planning and coordinate enterprise-wide infrastructure changes.
  • Develop tools for data collection, analysis, visualization, reporting, alerting, and monitoring.
  • Collaborate with leadership, engineers, program managers, and product managers to develop IT products and services.

What We Need To See

  • Bachelor’s degree in engineering, computer science, mathematics, or a related field, or equivalent experience.
  • At least 15 years of experience in compute platform engineering focused on automation.
  • Experience designing and deploying containerization architectures and distributed systems infrastructure.
  • Experience evaluating application architectures and identifying containerization opportunities.
  • Strong analytical skills with the ability to define and track performance metrics.
  • Experience with data analysis, performance profiling, Terraform, and configuration management tools.
  • Proficiency in Go and/or Python, Linux operating systems, and kernel internals.
  • Experience with bare-metal build infrastructure and network architectures including VLAN, VXLAN, SDN, BGP, and Anycast.

Ways To Stand Out

  • Deep understanding of DNS, LDAP, and security infrastructure components.
  • Hands-on experience implementing containers and managing DNS or LDAP services at scale.
  • Strong understanding of microservices architecture, infrastructure as code, and configuration management tools.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at NVIDIA AI