{bc}
indeed

Applied AI Site Reliability Engineer III

Clarus Advisers
Karnataka, IND
Full-time
Onsite
Discovered 1 weeks ago
Site Reliability EngineeringCloud platform engineeringDevOpsObservabilityKubernetesDocker
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site Reliability EngineeringCloud platform engineeringDevOps
Smart Apply

Full Job Posting

About the Role

Join a Product Engineering team building and operating reliable, scalable, secure, and cost-effective cloud-native platforms and AI-enabled products.

The role focuses on Site Reliability Engineering, cloud platform engineering, DevOps, observability, performance engineering, Kubernetes, infrastructure as code, and AI/ML production operations.

Key Responsibilities

  • Own production reliability, performance, availability, scalability, and operational efficiency using SLIs, SLOs, SLAs, and error budgets.
  • Implement observability with metrics, logs, tracing, dashboards, and actionable alerts.
  • Manage cloud-native production environments across AWS, Azure, or GCP.
  • Build infrastructure with Kubernetes, Docker, Terraform, CI/CD, and Git-based deployment practices.
  • Perform performance testing, capacity planning, autoscaling, resilience testing, and chaos engineering.
  • Participate in incident response, root-cause analysis, on-call operations, and blameless postmortems.
  • Operate AI/ML, GenAI, and agentic workloads in production.
  • Address model drift, train/serve skew, output variance, latency, and token or GPU cost anomalies.
  • Develop runbooks, operational playbooks, technical specifications, and reliability standards.

Required Skills and Experience

  • 5+ years of experience in software engineering, SRE, DevOps, or production engineering.
  • 3+ years of experience operating large-scale, distributed, cloud-native production systems.
  • Programming or scripting experience with Python, Go, Bash, Java, or C#/.NET.
  • Strong Kubernetes, Docker, Terraform, CI/CD, and cloud infrastructure experience.
  • Experience with cloud providers including AWS, Azure, or GCP.
  • Experience with observability and performance or load-testing tools.
  • Experience operating AI/ML or GenAI workloads in production.
  • Understanding of MLOps, LLMOps, AI reliability, agentic workloads, DevSecOps, RBAC, secrets management, and least privilege.
  • Strong software engineering fundamentals, analytical ability, troubleshooting, communication, and stakeholder-management skills.

Preferred Candidate Profile

  • Experience with Azure OpenAI, AWS Bedrock, or Vertex AI is highly desirable.
  • Familiarity with MLflow, LangSmith, LangFuse, or equivalent AI or agent orchestration tools is desirable.
  • Experience with large-scale distributed systems, cloud platforms, SaaS products, AI/ML platforms, GenAI applications, or highly available enterprise applications is preferred.

Travel

  • Travel of up to 10% may be required.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Clarus Advisers