Base Career helps you apply smarter for this job.
Key skills for this role
As the first dedicated ML Ops Engineer , you’ll own the tooling and infrastructure that make our ml engineers wildly productive and ensure we are able to efficiently iterate on ML models, prompts, and datasets and deploy our AI systems into a predictable production environment. You’ll bridge the gap between research and DevOps—designing reproducible dataset pipelines, automated experiment workflows, and Terraform-based cloud deployments that scale.
As the first dedicated ML Ops Engineer , you’ll own the tooling and infrastructure that make our ml engineers wildly productive and ensure we are able to efficiently iterate on ML models, prompts, and datasets and deploy our AI systems into a predictable production environment. You’ll bridge the gap between research and DevOps—designing reproducible dataset pipelines, automated experiment workflows, and Terraform-based cloud deployments that scale.
• Design version-controlled data pipelines (feature stores, data registries) using tools such as Delta Lake, Apache Iceberg • Implement systems for data validation, lineage tracking, and automated quality checks (e.g., Great Expectations).
• Build and maintain experiment orchestration with platforms like MLflow, torchx, and Apache Airflow. • Provide templated systems and tools to ML Engineers that easily launch training/evaluation data processing systems • Automate hyper-parameter sweeps and A/B tests, exposing clear dashboards for results.
• workflows that package, test, and promote models and agents through staging to production. • Implement canary deployments and rollbacks for models/agents services
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
San Francisco, USA
Palo Alto, USA
Palo Alto, USA
San Francisco, USA
Los Altos, USA
Los Altos, USA
Terraform Infrastructure-as-Code •
• Author and maintain Terraform modules for all ML infra—networking, GPU/TPU clusters, object storage, secrets, monitoring. • Enforce best practices for state management, workspaces, and automated plan/apply stages via CI.
• Integrate logging, tracing, and metric collection (Prometheus, Grafana, Datadog) across data pipelines and model endpoints. • Set SLIs/SLOs for data freshness and model latency; implement alerts and runbooks.
Security & Compliance • Work with Security to implement IAM least-privilege, key rotation, and data-encryption policies. • Support audit requirements (SOC 2, GDPR, HIPAA where applicable).
Experience with GPU orchestration (NVIDIA DGX, Karpenter, or Ray).
Familiarity with IaC security scanning (Checkov, tfsec).
Exposure to policy-as-code (OPA/Gatekeeper).
Prior work in real-time streaming (Kafka, Flink) and online feature serving .
Contributions to open-source ML Ops projects.
Reports to: Director of Infra
Full-time
Senior · 5+ years experience
Hybrid
Apply faster on company sites with our extension.