Base Career helps you apply smarter for this job.
Key skills for this role
The Senior Platform Reliability Engineer is part of the team that operates Firmus AI FactoryOS in production: the GPU compute fleet, and the platform services it depends on, including exabyte-scale storage, the shared core services, the virtualisation hosting the management plane, and the observability infrastructure the estate is measured through. This is state-of-the-art AI infrastructure, among the largest deployments in Asia Pacific, built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation.
This is a hands-on senior role with deep technical expertise, working in a team that shares accountability for the compute fleet and the platform services it depends on. The team runs those to a declared service level and sets the acceptance requirements each service has to meet before it goes live. The team also builds the shared administrative infrastructure the estate is run from, in consultation with the AI Infrastructure team , and operates it as a shared service. Automation is a first-class part of this role: the team builds and maintains the guarded automation and remediation tooling that turns manual response into a self-healing capability.
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Launceston, AUS
, AUS
Sydney, AUS
San Francisco, USA
Launceston, AUS
Sydney, AUS
Melbourne, AUS
San Francisco, USA
Launceston, AUS
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system.
AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against.
The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability, and builds and operates the shared services the estate's own operation depends on. The function works closely with the engineering teams that build the platform, supplying the production evidence that shapes what they fix and what they build next.
The Senior Platform Reliability Engineer is part of the team that operates Firmus AI FactoryOS in production: the GPU compute fleet, and the platform services it depends on, including exabyte-scale storage, the shared core services, the virtualisation hosting the management plane, and the observability infrastructure the estate is measured through. This is state-of-the-art AI infrastructure, among the largest deployments in Asia Pacific, built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation.
This is a hands-on senior role with deep technical expertise, working in a team that shares accountability for the compute fleet and the platform services it depends on. The team runs those to a declared service level and sets the acceptance requirements each service has to meet before it goes live. The team also builds the shared administrative infrastructure the estate is run from, in consultation with the AI Infrastructure team , and operates it as a shared service. Automation is a first-class part of this role: the team builds and maintains the guarded automation and remediation tooling that turns manual response into a self-healing capability.
Permanent full-time
At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.
Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.
Australian private AI infrastructure company building energy-efficient AI factories and cloud services across Asia-Pacific.
Visit company websiteJobs and hiring trendsFull-time
Senior · 8+ years experience
Onsite
Apply faster on company sites with our extension.