Applied AI Site Reliability Engineer III
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Join a Product Engineering team building and operating reliable, scalable, secure, and cost-effective cloud-native platforms and AI-enabled products.
The role focuses on Site Reliability Engineering, cloud platform engineering, DevOps, observability, performance engineering, Kubernetes, infrastructure as code, and AI/ML production operations.
Key Skills for This Role
Full Job Posting
About the Role
Join a Product Engineering team building and operating reliable, scalable, secure, and cost-effective cloud-native platforms and AI-enabled products.
The role focuses on Site Reliability Engineering, cloud platform engineering, DevOps, observability, performance engineering, Kubernetes, infrastructure as code, and AI/ML production operations.
Key Responsibilities
- Own production reliability, performance, availability, scalability, and operational efficiency using SLIs, SLOs, SLAs, and error budgets.
- Implement observability with metrics, logs, tracing, dashboards, and actionable alerts.
- Manage cloud-native production environments across AWS, Azure, or GCP.
- Build infrastructure with Kubernetes, Docker, Terraform, CI/CD, and Git-based deployment practices.
- Perform performance testing, capacity planning, autoscaling, resilience testing, and chaos engineering.
- Participate in incident response, root-cause analysis, on-call operations, and blameless postmortems.
- Operate AI/ML, GenAI, and agentic workloads in production.
- Address model drift, train/serve skew, output variance, latency, and token or GPU cost anomalies.
- Develop runbooks, operational playbooks, technical specifications, and reliability standards.
Required Skills and Experience
- 5+ years of experience in software engineering, SRE, DevOps, or production engineering.
- 3+ years of experience operating large-scale, distributed, cloud-native production systems.
- Programming or scripting experience with Python, Go, Bash, Java, or C#/.NET.
- Strong Kubernetes, Docker, Terraform, CI/CD, and cloud infrastructure experience.
- Experience with cloud providers including AWS, Azure, or GCP.
- Experience with observability and performance or load-testing tools.
- Experience operating AI/ML or GenAI workloads in production.
- Understanding of MLOps, LLMOps, AI reliability, agentic workloads, DevSecOps, RBAC, secrets management, and least privilege.
- Strong software engineering fundamentals, analytical ability, troubleshooting, communication, and stakeholder-management skills.
Preferred Candidate Profile
- Experience with Azure OpenAI, AWS Bedrock, or Vertex AI is highly desirable.
- Familiarity with MLflow, LangSmith, LangFuse, or equivalent AI or agent orchestration tools is desirable.
- Experience with large-scale distributed systems, cloud platforms, SaaS products, AI/ML platforms, GenAI applications, or highly available enterprise applications is preferred.
Travel
- Travel of up to 10% may be required.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at Clarus Advisers
Senior Product Specialist – ServiceNow CMDB / Business Analyst
, IND
Clarus Advisers is seeking a Senior Product Specialist / ServiceNow Business Analyst to own requirements and functional delivery for CMDB, ITOM, and ITSM implementations. Candidates need 5+ years of ServiceNow business a
WIFI Device Driver Developer
Hyderabad, IND
The employer is seeking a Senior Driver Developer to build Wi-Fi and networking drivers for next-generation automotive and embedded platforms. The role requires C and Python programming, Linux and Windows internals knowl