Base Career helps you apply smarter for this job.
Key skills for this role
As an AI Operations Engineer , you will help us run AI-powered workflows, applications and agents reliably and securely in production. You will take an operational engineering perspective across the service lifecycle, helping new AI capabilities move safely into production and ensuring they remain dependable once live.
You will work closely with the Hyperautomation Lead and AI Engineers , who design and build AI-enabled solutions. Your role is to make sure those solutions are production-ready and supportable, with appropriate deployment, environments, access, monitoring, resilience and recovery arrangements.
This is a hands-on engineering role . You will configure cloud services, automate deployments, manage environments and access, implement monitoring, troubleshoot production issues and improve service reliability over time. Because AI services often depend on multiple platforms, APIs and enterprise systems, you will also help coordinate operational dependencies across Technology, Security and other teams.
You do not need to be an AI model developer. We are looking primarily for strong production engineering skills combined with an interest in how AI-enabled applications behave in real-world environments.
Key Responsibilities
· Own the operational health of AI services. Help establish clear support arrangements, dependencies and operational standards for AI-powered workflows, applications and agents, and ensure services remain reliable, secure and supportable once live.
· Make new services production-ready. Work alongside the Hyperautomation Lead and AI Engineers to ensure new and changed services have appropriate deployment, monitoring, access controls, failure handling, recovery, support arrangements and documentation before go-live.
· Engineer deployment and cloud environments. Build and maintain repeatable deployment and environment patterns using CI/CD, infrastructure-as-code and configuration management, and configure the cloud services, identity, secrets and connectivity needed to operate AI services securely.
· Implement observability and improve reliability. Create and maintain useful logs, metrics, dashboards, alerts and health checks. Use telemetry, incidents and recurring operational issues to improve resilience, error handling, recovery and automation and to reduce manual operational effort.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
London, GBR
London, GBR
London, GBR
Haywards Heath, GBR
London, GBR
London, GBR
London, GBR
London, GBR
London, GBR
· Manage incidents and operational problems. Act as a technical responder for production incidents, troubleshoot issues across cloud, application, identity, integration and workflow layers, and coordinate service restoration where other teams or vendors are involved. Contribute to root-cause analysis and ensure appropriate corrective actions are identified and followed through.
· Operate integrations and technical dependencies. Maintain visibility over APIs, connectors, authentication and external platform dependencies, troubleshoot integration issues and work with internal teams to resolve problems that affect production services.
· Maintain strong operational documentation. Produce and keep up-to-date technical documentation, runbooks, support procedures, recovery processes and operational playbooks covering how services are deployed, monitored, supported and restored.
· Implement operational governance and controls. Ensure appropriate security, access, auditability, data-handling and AI operational controls are built into production services, including permissions boundaries, logging, approval controls and recovery mechanisms where appropriate.·
Coordinate technically with platform providers and vendors. Act as an operational and technical counterpart to providers such as AWS, Microsoft, Salesforce, Anthropic and Mindflow, coordinating technical support and escalations and assessing platform changes that may affect reliability, security or supportability.
· Continuously improve how AI services are operated. Feed operational learning back to the Hyperautomation Lead, AI Engineers and wider Technology teams, help improve reusable operational patterns and remove recurring sources of support effort or production risk.
Knowledge, skills and experience required
Must-have
We are looking for someone with strong production engineering and operational ownership experience. You do not need to be an AI model developer.
Production cloud engineering
Hands-on experience operating applications or services in AWS, Azure or a comparable cloud environment , including configuration, troubleshooting, access, monitoring and production support.
Deployment, infrastructure and automation
Practical experience with:
· CI/CD and Git-based delivery practices
· Infrastructure-as-code such as Terraform, CloudFormation, CDK or equivalent
· Environment and configuration management
· Automating repeatable operational tasks
Programming, APIs and integrations
Practical programming or scripting experience using Python, TypeScript/JavaScript, Bash, PowerShell or similar , with the ability to automate tasks and diagnose production issues.
Good working knowledge of APIs, HTTP, JSON, authentication and system integrations, including troubleshooting dependencies across different systems.
Observability, incident and problem management
Experience:
· Working with production logs, metrics, dashboards and alerts
· Troubleshooting live services
· Responding to and coordinating production incidents
· Contributing to root-cause analysis and corrective actions
· Improving monitoring and reliability based on operational experience
You should be comfortable working through incidents that span several systems, teams or external providers rather than only troubleshooting a single application.
Security, access and operational controls
Good understanding of:
· Identity and access management
· Secrets and credential management
· Least-privilege access
· Secure configuration
· Auditability and operational controls
You should be comfortable applying security and governance requirements as part of normal production engineering rather than treating them as a separate activity.
Operational ownership and documentation
Experience taking responsibility for how production services are operated and supported, including creating and maintaining:
· Runbooks and operational playbooks
· Recovery and troubleshooting procedures
· Deployment and support documentation
· Service dependencies and escalation paths
Strong written documentation skills are important for this role.
Cross-team and vendor technical coordination
Ability to work effectively with engineers, infrastructure and security teams, business stakeholders and external technology providers.
You should be comfortable:
· Coordinating technical resolution where several teams are involved
· Working with vendors during incidents or technical escalations
· Explaining technical risks and operational issues clearly
· Following issues and corrective actions through to resolution
Nice to have
Experience in any of the following would be useful, but is not required :
· Operating AI-enabled applications, LLM applications, agents or workflow automation in production
· Understanding AI-specific operational considerations such as model/API dependencies, rate limits, latency, retries, tool calls, structured outputs and failure handling
· Mindflow, n8n or similar workflow/orchestration platforms
· Amazon Bedrock, Anthropic, OpenAI or other model platforms
· Microsoft Entra, AWS IAM or Salesforce
· CloudWatch, DataDog or similar observability tooling
· Containers or Kubernetes
· Distributed tracing and SRE practices such as SLIs and SLOs
· MCP, RAG, agentic AI or tool/function calling
· Secure cloud networking for enterprise integrations
· Building reusable deployment templates, monitoring patterns, infrastructure modules or operational tooling
Candidates from Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, Production Engineering or MLOps backgrounds may be particularly well suited to the role.
We welcome strong production engineers who may not yet have extensive AI operations experience but are interested in developing that capability.
Collinson is an equal opportunity employer and welcomes differences in all their forms including: colour, race, ethnicity, gender identity, sexual orientation, neurodivergence, family status, age, individuals with disabilities and people from all backgrounds, cultures and experiences as we strongly believe this contributes to our on-going success.
We are focused on continually evolving our purpose driven, high performing culture, providing an environment where our people have the opportunity to achieve their full potential and do interesting and meaningful work. Our company values are: Take Action, Do the right thing, One team and Be insight led. These help guide everything we do internally in terms of how we think, act and interact, right through to how we deliver value to our customers and clients.
In your application, please feel free to note which pronouns you use (For example - she/her/hers, he/him/his, they/them/theirs, etc).
If you need any extra support throughout the interview process, then please email us at ukrecruitment@collinsongroup.com
Verified company details for this employer are not available yet.
Full-time
Senior
Hybrid
Apply faster on company sites with our extension.