Troubleshoot and resolve complex network incidents affecting production site services, driving root cause analysis and implementing corrective actions to prevent recurrence
Formulate the right metrics and definitions of success to report and drive quality, efficiency, cost, and timeliness, and evolve these over time to match changes to the infrastructure and business requirements; manage related programs, project budget plans, and budgets
Develop operational process improvement plans and transform with partner teams the improvements to scalable and automated AI-powered workflows by writing and reviewing the code to improve operational efficiency
Participate in team oncall rotation and improve issue/event escalation and emergency/incident response, including detailed after-action reviews to prevent future recurrences
Build cross-functional relationships and deliver results with partner teams (internally and externally) associated with all coordination, colocation operations, and compliance issues, including vendors, contractors, and stakeholders/partners
Support real-world operations and production challenges that affect network capacity and reliability, and take the lessons to improve current and future-generation coordination products and business processes
Identify operational gaps between system components and contribute to the design and delivery of scalable solutions that improve network reliability and availability efficiency
Participate in an oncall rotation to provide 24/7 support for production site services network infrastructure
Engage stakeholders across the network engineering, facilities, and vendor teams to coordinate project execution and communicate progress against milestones
Up to 15 percent of travel (domestic and international)
Minimum Qualifications
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
8+ years of experience in designing, deploying, operating, and troubleshooting large-scale production network infrastructure
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
Experience working within a global team and collaborating with cross-functional teams in a fast-paced and dynamic environment with limited supervision
Experience in configuring and troubleshooting routing and switching protocols, including BGP, IS-IS, OSPF, MPLS, and spanning tree variants
Experience with network protocols including TCP/IP, DHCP, and DNS, and hands-on experience with both IPv4 and IPv6 environments
Experience working in multi-vendor network environments with hands-on familiarity with enterprise or hyperscale networking hardware
Strong organizational/multitasking skills and ability to adapt to quickly changing priorities with excellent communication skills (verbal and written)
Bachelor's degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience
Preferred Qualifications
Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
Understanding of AI methods and tools, as well as training workloads and the demands they exert on the network. Knowledge of data-driven analysis and analytics applied to a full project lifecycle
Familiarity with physical infrastructure design: rack elevations, cable types, connector types, optic types, patch panels, power/cooling, and facility infrastructure
Familiarity and experience with the creation and management of Standard Operating Procedures (SOPs), Methods of Procedure (MOPS), and policies, alongside comprehensive technical and operations guidance and knowledge management practices
Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
M.S. in Computer Science/Information Technology, Computer Engineering, Operations leadership, or a related discipline, or an equivalent mix of business process and technical experience
Some coding experience in Python, or a willingness to learn
Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy review)