Design, develop, test and operate networking systems to support large scale AI training jobs
Research, develop and deploy numerous technologies and network topologies in order to evolve and scale our AI networks
Work closely with our hardware, software and sourcing teams to develop new networking solutions and influence the future of networking and its associated infrastructure
Define and develop optimized network automation tools and systems, including configuration, provisioning, monitoring, alarming, auto-remediation and more
Be oncall to learn from real world production challenges and take the lessons to improve current and future generation products
Provide guidance on network architecture including scale-up and scale-out topologies, transport protocols, and performance optimization techniques
Minimum Qualifications
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
8+ years of experience in system performance engineering, network infrastructure engineering, or a related field within large-scale distributed computing or HPC environments
Experience coding in languages like Python, C++, Go
Experience in designing, deploying and operating datacenter networks at scale
Experience in network automation software leveraging software defined networking principles
Preferred Qualifications
Understanding of AI training workloads and demands they exert on networks
Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
4+ years of experience working on networks supporting large scale training workloads
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career