Base Career helps you apply smarter for this job.
Key skills for this role
Lead and mentor the DevOps and Platform Engineering team.
Define and implement DevOps, Cloud, and AI infrastructure strategy, roadmap, and best practices.
Collaborate with Engineering, QA, Security, Product, Data Engineering, and AI/ML teams.
Promote a DevOps culture focused on automation, reliability, scalability, security, and continuous improvement.
Drive adoption of AI-powered operational practices (AIOps) to improve monitoring, incident management, and operational efficiency.
Design, implement, and manage scalable CI/CD pipelines.
Automate build, testing, deployment, and release management processes.
Implement Infrastructure as Code (IaC) and GitOps practices.
Build and maintain automated workflows for AI/ML model deployment and MLOps pipelines.
Integrate AI-assisted automation tools to improve deployment velocity and operational efficiency.
Collaborate with Data Science and AI teams to build and maintain scalable AI/ML infrastructure.
Design and support MLOps pipelines for model training, testing, deployment, monitoring, and lifecycle management.
Manage AI/ML workloads on cloud platforms and Kubernetes environments.
Implement model observability, drift detection, performance monitoring, and automated retraining workflows.
Support for GPU-based infrastructure and optimization for AI/ML workloads was required.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Sandy, GBR
Bristol, GBR
East Kilbride, GBR
Bedford, GBR
Peterlee, GBR
Glasgow, GBR
Cookstown, CAN
Cookstown, CAN
Architect, implement, and manage cloud infrastructure (AWS/Azure/GCP).
Ensure high availability, scalability, resilience, and disaster recovery capabilities.
Manage containerization and orchestration platforms (Docker, Kubernetes, OpenShift).
Optimize cloud infrastructure for performance, cost, and security.
Support hybrid and multi-cloud infrastructure environments.
Implement enterprise monitoring, logging, tracing, and alerting systems.
Ensure system uptime, performance optimization, and proactive incident response.
Conduct root cause analysis and implement preventive and self-healing mechanisms.
Leverage AI/ML-based monitoring and predictive analytics for anomaly detection and operational insights.
Define and track SRE metrics, including SLIs, SLOs, and SLAs.
Implement DevSecOps and security automation practices.
Manage IAM, secrets management, vulnerability management, and compliance standards.
Ensure infrastructure and AI platform security best practices.
Collaborate with Security teams to enforce governance, compliance, and audit requirements.
Global provider of online vehicle auction and remarketing services.
Visit company websiteJobs and hiring trendsFull-time
Senior · 8+ years experience
Onsite
Apply faster on company sites with our extension.