Base Career helps you apply smarter for this job.
Key skills for this role
We're looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale. You understand what it takes to keep production ML and data workloads fast, observable, and resilient — from GPU-backed training and inference clusters to the pipelines that connect them to real-world industrial environments.
This role is ideal for candidates who want deep infrastructure ownership at a company where uptime, latency, and reliability directly affect physical operations — not just software. You'll partner closely with AI, data, and product engineering teams to make sure the systems they build can actually run in production, safely and at scale.
We're looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale. You understand what it takes to keep production ML and data workloads fast, observable, and resilient — from GPU-backed training and inference clusters to the pipelines that connect them to real-world industrial environments.
This role is ideal for candidates who want deep infrastructure ownership at a company where uptime, latency, and reliability directly affect physical operations — not just software. You'll partner closely with AI, data, and product engineering teams to make sure the systems they build can actually run in production, safely and at scale.
Own and evolve Nexxa's core infrastructure — compute, networking, storage, and deployment systems — end-to-end
Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams
Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments
Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
, USA
Sunnyvale, USA
Sunnyvale, USA
Toronto, CAN
, CAN
, CAN
Sunnyvale, USA
Sunnyvale, USA
Partner with data and AI teams to support the infrastructure behind: Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks) Feature stores, embedding indices, and retrieval pipelines Model training, evaluation, and serving infrastructure
Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks)
Feature stores, embedding indices, and retrieval pipelines
Model training, evaluation, and serving infrastructure
Define and drive observability practices — metrics, logging, tracing, and alerting — across distributed systems
Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations
Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations
Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity
Collaborate with engineering leadership to define infrastructure roadmap and platform strategy
Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org
You can own ambiguous, high-stakes infrastructure problems end-to-end
Systems you build stay reliable as usage and scale grow — you design for the next order of magnitude, not just today
You bring strong technical judgment on tradeoffs between reliability, cost, and speed
You raise the bar for operational rigor and engineering discipline across the team
You help define what's next for the platform, not just execute what's known
Innovative Environment: Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies
Collaborative Culture: Be part of a team that values innovation, discipline, and continuous improvement
Professional Growth: Benefit from significant opportunities for career development and advancement
Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions
If you're passionate about building the infrastructure that powers advanced AI solutions in the real world, we'd love to connect.
Verified company details for this employer are not available yet.
Full-time
Senior · 6+ years experience
Remote
Apply faster on company sites with our extension.