Base Career helps you apply smarter for this job.
Key skills for this role
We are seeking a deeply technical, hardware-passionate Datacentre Operations Engineer to execute on-the-ground operations for our new East London deployment—Radiant's newest AI infrastructure site. This role focuses on delivering precise, repeatable physical practices—including advanced smart-hands support, complex cabling, and hands-on operation of high-density, air-cooled compute systems—to guarantee world-class SLAs on next-generation hardware architecture.
Working closely with Infrastructure (HPC) SRE, Network Engineering, and Datacentre Strategy teams, you will uphold uncompromising standards across East London's data centre floor, spanning multiple data halls. You will live and breathe the hardware, maintaining elite facility reliability through hands-on deployment, proactive maintenance, rapid incident response, and structured break/fix execution across high-density air cooling systems, busbar-based power distribution, and next-generation GPU compute platforms. The role centres on technical execution and optimised output.
You will turn global engineering standards into flawless, repeatable daily routines, continually honing on-the-ground practices to keep our most advanced hardware running at peak performance. East London's hardware platform is NVIDIA B300-class, fully air-cooled compute, delivered through a new OEM partner being onboarded to Radiant's fleet; while the interfaces and overall architecture closely align with other OEM platforms you may already know, you'll be one of the first teams operating it on the ground, so a fast learning curve and strong first-principles troubleshooting matter more than prior exposure to this specific vendor.
You will be part of the founding team standing up 24x7x365 on-site coverage at East London to meet tight customer SLAs, and as Radiant expands its EMEA footprint, your relentless drive for hardware perfection and proven field expertise with high-density environments will serve as the operational blueprint to scale execution models efficiently across the region.
Radiant is powered intelligence, delivering scalable, high-performance compute infrastructure purpose-built for AI and HPC workloads. Operating across global data centres, we run mission-critical environments where uptime, throughput, and ultra-low latency are non-negotiable.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
London, GBR
, GBR
London, GBR
London, GBR
London, GBR
Lewis Center, USA
Lewis Center, USA
Columbus, USA
We are seeking a deeply technical, hardware-passionate Datacentre Operations Engineer to execute on-the-ground operations for our new East London deployment—Radiant's newest AI infrastructure site. This role focuses on delivering precise, repeatable physical practices—including advanced smart-hands support, complex cabling, and hands-on operation of high-density, air-cooled compute systems—to guarantee world-class SLAs on next-generation hardware architecture.
Working closely with Infrastructure (HPC) SRE, Network Engineering, and Datacentre Strategy teams, you will uphold uncompromising standards across East London's data centre floor, spanning multiple data halls. You will live and breathe the hardware, maintaining elite facility reliability through hands-on deployment, proactive maintenance, rapid incident response, and structured break/fix execution across high-density air cooling systems, busbar-based power distribution, and next-generation GPU compute platforms. The role centres on technical execution and optimised output.
You will turn global engineering standards into flawless, repeatable daily routines, continually honing on-the-ground practices to keep our most advanced hardware running at peak performance. East London's hardware platform is NVIDIA B300-class, fully air-cooled compute, delivered through a new OEM partner being onboarded to Radiant's fleet; while the interfaces and overall architecture closely align with other OEM platforms you may already know, you'll be one of the first teams operating it on the ground, so a fast learning curve and strong first-principles troubleshooting matter more than prior exposure to this specific vendor.
You will be part of the founding team standing up 24x7x365 on-site coverage at East London to meet tight customer SLAs, and as Radiant expands its EMEA footprint, your relentless drive for hardware perfection and proven field expertise with high-density environments will serve as the operational blueprint to scale execution models efficiently across the region.
Join a team operating some of the world’s most advanced high-performance computing infrastructure. As a Datacentre Operations Engineer, you’ll work hands-on with cutting-edge GPU and CPU platforms — including the latest NVIDIA architectures — powering dense, large-scale compute environments used for AI, machine learning, and next-generation workloads.
This is an opportunity to build expertise at the forefront of modern infrastructure, where reliability, scale, and performance matter every day. You’ll collaborate with experienced engineers across a globally distributed organisation that values openness, inclusion, technical excellence, and continuous learning.
We move quickly, solve meaningful challenges, and give people the space to make an impact. If you thrive in fast-paced environments, enjoy working with advanced technology, and want to help shape the future of high-performance compute, you’ll find both challenge and opportunity here.
Exposure to industry-leading GPU and AI infrastructure
Opportunities to grow alongside a rapidly scaling global business
A collaborative, inclusive, and supportive engineering culture
Real ownership and the ability to influence operational excellence
Work that sits at the intersection of people, performance, and technology
A modern, flexible, globally connected workplace with ambitious goals
Quickly diagnose and resolve hardware and network issues to maximise uptime; execute structured fault isolation methodologies to drive rapid resolution
Respond to critical hardware alerts via our monitoring and observability platform; contribute to ongoing service improvement to improve monitoring capability and alert quality
Deploy and maintain HPC and AI hardware for uninterrupted operations, including hardware troubleshooting, firmware updates, and component replacement
Execute break/fix procedures for advanced hardware platforms, including GPU module exchange, component-level fault isolation, and firmware-level diagnostics
Execute or support break/fix operations on ultra-high-density compute systems including NVIDIA B300-class chassis or equivalent platforms, including GPU/fan module exchange, chassis-level fault isolation, and busbar connection/disconnection—under the direction of the Lead where qualification is in progress
Operate, monitor, and maintain high-density air cooling infrastructure in conjunction with our datacentre partner, including CRAC/CRAH units, in-row and containment cooling, and associated airflow management systems
Facilitate in conjunction with our datacentre partner, routine and corrective maintenance on air cooling systems: monitoring supply/return temperatures and airflow rates, maintaining hot-aisle/cold-aisle containment and blanking, and performing scheduled filter and component inspections
Follow and contribute to SOPs for safe working around high-density, air-cooled compute platforms
Monitor thermal performance and raise anomalies before they escalate into incidents
Contribute to site-level capacity management operations, maintaining accurate records of power, space, and cooling utilisation
Support capacity planning activities by providing accurate as-built data and flagging infrastructure changes to the Lead and relevant teams
Manage on-the-ground assets from point of purchase and delivery through lifecycle management and disposal, owning asset management within Radiant's CMDB system
Handle RMAs and support requests within Radiant's Service Level Objectives (SLOs) to meet customer contract SLAs
Contribute to ongoing maintenance, fostering compliance and leveraging strong vendor partnerships
Operate cooling, power distribution (including busbar and PDU infrastructure), and other critical data centre technologies to maintain high operational standards
Develop and maintain datacentre/hardware management SOPs, ensuring continual alignment with Radiant's governance and compliance requirements
Apply ITSM frameworks: Incident, Major Incident, Change Management, and service improvement
Operate and support services 24x7x365 for production environments as part of a structured on-site shift rotation—working 12-hour shifts across a 4-team pattern with two engineers on shift at all times—to meet tight customer SLAs
Prioritise and triage incident and smart-hands workload ensuring tight SLA coverage is maintained with a small on the ground team
Contribute to Incident postmortem analyses, root cause analysis, document learnings, and automate remediations
Communicate technical decisions clearly to stakeholders and customers
Champion a culture of: do, document, automate
Willing to cross train and upskill in Infrastructure/Platform SRE practices
Willing to travel across EMEA to support future datacentre onboarding and train in new technologies
Brookfield-backed AI infrastructure company building and operating integrated AI factories for enterprises, telecom providers, and sovereign institutions.
Visit company websiteJobs and hiring trendsFull-time
Mid · 3+ years experience
Hybrid
Apply faster on company sites with our extension.