Base Career helps you apply smarter for this job.
Key skills for this role
Reliability: Ensure the reliability, scalability, and security of AI infrastructure supporting HPC & AI workloads. Incident Management: Lead incident response, root cause analysis, and continuous improvement to minimize downtime and optimize service availability. Performance Optimization: Identify and resolve bottlenecks in compute, storage, networking, and specialized hardware (GPUs, InfiniBand) to enhance AI system performance. Infrastructure Automation: Develop and maintain automation tools for deployment, monitoring, predictive analysis and management of AI infrastructure, including containerized environments (Kubernetes, Docker). Technical Leadership: Provide technical guidance in cloud and AI infrastructure technologies, collaborating with cross-functional teams to drive innovation and best practices. Master's Degree in Computer Science, Information Technology, or related field AND 1+ year(s) technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 6+ years technical experience in software engineering, network engineering, or systems administration 8+ years of professional software engineering experience, with 5+ years in service operations, monitoring, and reliability improvement for infrastructure. 1+ years experience with incident management and reliability engineering in cloud or AI environments. Master's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 5+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience. 2+ years technical experience working with large-scale cloud or distributed systems. 1+ years experience in distributed systems and/or cloud platforms (Azure, Kubernetes, Docker, containers ecosystem). 1+ years experience with GPUs, InfiniBand, or similar high-performance technologies.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Redmond, USA
Warrenton, USA
Redmond, USA
London, GBR
, USA
Hillsboro, USA
London, GBR
Bengaluru, IND
Redmond, USA
Microsoft is a global technology company that develops software, hardware, and cloud services, known for products like Windows, Office, Azure, and Xbox.
Full-time
Senior · 8+ years experience
Apply faster on company sites with our extension.