Base Career helps you apply smarter for this job.
Key skills for this role
Mirantis is looking for a Technical Product Manager to own observability for k0rdent AI, our control plane for GPU infrastructure and distributed AI workloads. In this role, you will define the observability strategy, roadmap, and feature priorities that determine how operators gain visibility into the health, performance, and resource utilization of GPU clusters running large-scale training and inference. You will shape how k0rdent AI handles everything from GPU-level metrics and distributed tracing across AI workloads, to multi-tenant log aggregation and intelligent alerting — powered by the OpenTelemetry ecosystem, and Prometheus-compatible metrics pipelines.
The ideal candidate brings strong technical fluency in observability tooling and the AI infrastructure stack. You will work directly with engineering to shape requirements, with marketing to define positioning, and with customers to help ensure their success.
Responsibilities
Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, east-west fabric (InfiniBand, RoCE), high-performance storage, DPU/SmartNIC telemetry, workload schedulers, inference serving, and data services
Translate requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise platform teams into clear product direction; partner with engineering to define requirements and evaluate trade-offs
Manage the observability backlog using feedback from production deployments and design partners to refine priorities
Track and shape our response to emerging observability standards and technologies, including OpenTelemetry (OTel), DCGM GPU metrics, InfiniBand/RoCE fabric counters, storage platform telemetry APIs, and AI workload profiling
Define integration strategies for vendor telemetry sources across the ecosystem — NVIDIA compute and BlueField DPUs, storage platforms (VAST, Weka, DDN), workload managers (SLURM), inference stacks, and vector and relational databases — into a unified, operator-facing observability plane
Partner with product marketing and field teams on positioning, technical briefs, and reference architectures; represent Mirantis with customers, analysts, and ecosystem partners
5+ years in product management or a senior technical role owning an observability product or operating large-scale monitoring infrastructure
Working knowledge of Prometheus, OpenTelemetry, distributed tracing (Jaeger, Tempo), and log aggregation (Loki, Elasticsearch/OpenSearch)
Fluency in Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architecture
Ability to work directly with engineering on technical trade-offs and with field teams in competitive GPU cloud and NeoCloud deals
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
, USA
About the Role & Mission AI infrastructure is undergoing a generational shift. Disaggregated, raw GPU hardware must be transformed into multi-tenant, production-ready, and sovereign AI clouds - without locking enterprise
, USA
Mirantis is repositioning from a recognized private cloud infrastructure leader into a full-stack AI Infrastructure company — the vendor-agnostic, open-standard platform for building and operating AI clouds for Neoclouds
, USA
Job Summary Mirantis is looking for a Technical Product Manager to own observability for k0rdent AI, our control plane for GPU infrastructure and distributed AI workloads. In this role, you will define the observability
, USA
Mirantis is seeking a Senior Technical Product Marketing Manager focused on k0rdent AI, our AI infrastructure platform. You will be responsible for the technical proof behind how we bring k0rdent AI to market: white pape
, USA
Mirantis is building a new enterprise AI infrastructure product that lets organizations run and govern large language models on their own Kubernetes clusters. You will join a small senior team early, with broad ownership
, USA
Mirantis is seeking a Senior Product Marketing Manager focused on k0rdent AI, our AI infrastructure platform. You will be responsible for how we bring k0rdent AI to market and how it is understood by the platform enginee
, USA
Mirantis is building a new enterprise AI infrastructure product that lets organizations run and govern large language models on their own Kubernetes clusters. You will join a small senior team early, with broad ownership
, USA
Overview Design and build the software that provisions, integrates, and operates high-performance storage for GPU-accelerated compute and AI platforms. You will develop the control-plane services, drivers, and tooling th
, USA
, USA
, USA
, USA
, USA
, USA
, USA
, USA
Strongly Preferred:
Exposure to GPU observability, including DCGM metrics, AI workload profiling and performance analysis
Familiarity with east-west fabric telemetry - InfiniBand counters, RoCEv2 congestion metrics (ECN, PFC, DCQCN), or switch-level fabric health
Experience with high-performance storage telemetry from platforms such as VAST Data, Weka, or DDN, including IOPS, latency, and throughput instrumentation at scale
Familiarity with NVIDIA BlueField DPU telemetry, SR-IOV, or offload pipeline observability
Exposure to workload-level visibility for SLURM job scheduling, inference serving stacks (vLLM, Triton, TensorRT-LLM), or data service telemetry from vector databases (Milvus, Qdrant) and relational databases in AI pipelines
Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. https://www.mirantis.com/
Mirantis is looking for a Technical Product Manager to own observability for k0rdent AI, our control plane for GPU infrastructure and distributed AI workloads. In this role, you will define the observability strategy, roadmap, and feature priorities that determine how operators gain visibility into the health, performance, and resource utilization of GPU clusters running large-scale training and inference. You will shape how k0rdent AI handles everything from GPU-level metrics and distributed tracing across AI workloads, to multi-tenant log aggregation and intelligent alerting — powered by the OpenTelemetry ecosystem, and Prometheus-compatible metrics pipelines.
The ideal candidate brings strong technical fluency in observability tooling and the AI infrastructure stack. You will work directly with engineering to shape requirements, with marketing to define positioning, and with customers to help ensure their success.
Responsibilities
Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, east-west fabric (InfiniBand, RoCE), high-performance storage, DPU/SmartNIC telemetry, workload schedulers, inference serving, and data services
Translate requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise platform teams into clear product direction; partner with engineering to define requirements and evaluate trade-offs
Manage the observability backlog using feedback from production deployments and design partners to refine priorities
Track and shape our response to emerging observability standards and technologies, including OpenTelemetry (OTel), DCGM GPU metrics, InfiniBand/RoCE fabric counters, storage platform telemetry APIs, and AI workload profiling
Define integration strategies for vendor telemetry sources across the ecosystem — NVIDIA compute and BlueField DPUs, storage platforms (VAST, Weka, DDN), workload managers (SLURM), inference stacks, and vector and relational databases — into a unified, operator-facing observability plane
Partner with product marketing and field teams on positioning, technical briefs, and reference architectures; represent Mirantis with customers, analysts, and ecosystem partners
5+ years in product management or a senior technical role owning an observability product or operating large-scale monitoring infrastructure
Working knowledge of Prometheus, OpenTelemetry, distributed tracing (Jaeger, Tempo), and log aggregation (Loki, Elasticsearch/OpenSearch)
Fluency in Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architecture
Ability to work directly with engineering on technical trade-offs and with field teams in competitive GPU cloud and NeoCloud deals
Strongly Preferred:
Exposure to GPU observability, including DCGM metrics, AI workload profiling and performance analysis
Familiarity with east-west fabric telemetry - InfiniBand counters, RoCEv2 congestion metrics (ECN, PFC, DCQCN), or switch-level fabric health
Experience with high-performance storage telemetry from platforms such as VAST Data, Weka, or DDN, including IOPS, latency, and throughput instrumentation at scale
Familiarity with NVIDIA BlueField DPU telemetry, SR-IOV, or offload pipeline observability
Exposure to workload-level visibility for SLURM job scheduling, inference serving stacks (vLLM, Triton, TensorRT-LLM), or data service telemetry from vector databases (Milvus, Qdrant) and relational databases in AI pipelines
Build the observability foundation for the AI cloud era, working directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises
Collaborate with a world-class, distributed team committed to openness and technical excellence
Shape the product narrative and influence go-to-market success
It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to isamoylova@mirantis.com
By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.
#remote
We are a Leader for Container Management in G2 (#2 after AWS)!
Kubernetes-native AI infrastructure and cloud software company serving enterprises with open-source orchestration, automation, security, and support.
Visit company websiteJobs and hiring trendsFull-time
Senior · 5+ years experience
Remote
Apply faster on company sites with our extension.