{bc}
indeed

Senior Staff Site Reliability Engineer

NVIDIA
Karnataka, IND
Full-time
Hybrid
Discovered 1 weeks ago
Compute platform engineeringInfrastructure automationDNS, NTP/PTP, DHCP, and LDAPContainerization architecturesDistributed systems infrastructureTerraform
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Compute platform engineeringInfrastructure automationDNS, NTP/PTP, DHCP, and LDAP
Smart Apply

Full Job Posting

Role Overview

NVIDIA is seeking a Senior Staff Site Reliability Engineer to improve the efficiency and performance of infrastructure across on-premises and cloud environments.

The role focuses on globally scalable compute infrastructure, automation, reliability, observability, and lifecycle management.

Responsibilities

  • Lead architecture transformation initiatives for the IT Compute Core Team and build new on-premises and cloud service offerings.
  • Design, scale, and deploy DNS, NTP/PTP, DHCP, and LDAP services with automation, monitoring, high availability, capacity planning, and lifecycle management.
  • Define efficiency metrics and drive software and hardware optimizations, including SR-IOV and DPU improvements.
  • Use eBPF and XDP for observability and DDoS mitigation.
  • Analyze system and capacity data and coordinate enterprise-wide implementation plans.
  • Develop tools for data collection, analysis, visualization, reporting, alerting, and monitoring.
  • Collaborate with technical and product stakeholders to develop IT products and services.

Required Qualifications

  • Bachelor’s degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
  • 15+ years of compute platform engineering experience focused on automation.
  • Experience with containerization architectures and distributed systems infrastructure.
  • Experience evaluating architectures and identifying containerization opportunities for scalability, reliability, and efficiency.
  • Strong analytical skills for defining and tracking performance metrics.
  • Experience with data analysis, performance profiling, Terraform, and configuration management tools.
  • Proficiency in Go and/or Python.
  • Linux proficiency with kernel internals knowledge.
  • Experience with large bare-metal build infrastructure environments.
  • Understanding of VLAN, VXLAN, SDN, BGP, and Anycast networking.

Preferred Expertise

  • Deep understanding of DNS, LDAP, and security infrastructure components.
  • Hands-on experience implementing containers at scale.
  • Experience deploying and managing DNS and LDAP services at scale.
  • Understanding of microservices architecture, infrastructure as code, and configuration management tools.

Workplace

  • The posting indicates a hybrid work arrangement.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at NVIDIA