:Senior MLOps + DevOps Engineer (On-Prem AI Platform)
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Architect, build, and scale AI and machine learning platforms in an on-premises enterprise environment.
Own ML systems, infrastructure, CI/CD, and production reliability for scalable machine learning and GenAI solutions.
Key Skills for This Role
Full Job Posting
Role Overview
Architect, build, and scale AI and machine learning platforms in an on-premises enterprise environment.
Own ML systems, infrastructure, CI/CD, and production reliability for scalable machine learning and GenAI solutions.
Platform Architecture and Ownership
- Design and own end-to-end ML platform architecture covering data, training, deployment, and monitoring.
- Define secure, scalable ML system practices and standardize MLOps and DevOps frameworks.
Model Deployment and Serving
- Deploy and manage ML and LLM models on GPU-based on-premises infrastructure.
- Optimize inference latency, throughput, and batching, and implement versioning, A/B testing, and rollback strategies.
CI/CD and Infrastructure
- Design CI/CD pipelines for ML models, APIs, and data workflows with automated testing, deployment, and release management.
- Manage Linux infrastructure, Docker containers, Kubernetes or OpenShift workloads, and restricted or air-gapped environments.
Data, Monitoring, and Reliability
- Build pipelines integrating structured databases with high-volume logs and streaming data for batch and real-time inference.
- Implement model and infrastructure observability with Prometheus, Grafana, and ELK stack.
- Ensure high availability, SLA adherence, incident response, and production reliability.
GenAI and Collaboration
- Deploy RAG pipelines and vector databases, manage LLM serving frameworks, and work with agent orchestration frameworks.
- Mentor engineers, collaborate with cross-functional teams, drive design reviews, and support production readiness.
Required Skills and Experience
- Strong Python and Bash scripting skills.
- Deep understanding of the ML lifecycle and productionization.
- Experience deploying ML and LLM systems in production.
- Experience with Linux, Docker, Kubernetes or OpenShift, CI/CD tools, SQL, and data pipelines.
- At least 8 years of MLOps, DevOps, or platform engineering experience and proven experience scaling production ML systems.
Good to Have
- GPU optimization knowledge.
- Experience with MLflow or Kubeflow.
- Experience with Terraform or Ansible.
- Experience in on-premises or restricted environments.
Ideal Candidate
A hands-on platform architect who can operate across ML systems and infrastructure while driving automation, scalability, and reliability.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career