Base Career helps you apply smarter for this job.
Key skills for this role
FlexAI is looking for a Staff DevOps / SRE Engineer to define our infrastructure strategy, establish SRE best practices, and build systems capable of running large-scale AI workloads across distributed, multi-cloud environments.
You’ll work closely with developers to ensure our platform is reliable, performant, and scalable — without slowing down product velocity.
FlexAI is looking for a Staff DevOps / SRE Engineer to define our infrastructure strategy, establish SRE best practices, and build systems capable of running large-scale AI workloads across distributed, multi-cloud environments.
You’ll work closely with developers to ensure our platform is reliable, performant, and scalable — without slowing down product velocity.
Design and evolve the infrastructure backbone for our AI and PaaS platform
Build highly available, fault-tolerant, and scalable systems
Define and drive SRE practices (SLIs, SLOs, error budgets)
Lead Infrastructure as Code using Pulumi
Own and scale Kubernetes clusters and containerized workloads
Standardize and automate infrastructure for global deployments
Design and scale CI/CD pipelines for fast, reliable releases
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Santa Clara, USA
Santa Clara, USA
Santa Clara, USA
San Jose, USA
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Build self-healing systems and automated remediation workflows
Drive GitOps and platform engineering practices
Implement end-to-end observability using VictoriaMetrics and Grafana (metrics, logs, traces)
Identify and resolve performance bottlenecks (latency, throughput, cost)
Lead incident response, root cause analysis, and postmortems
Partner with backend, AI, runtime, and security teams
Guide infrastructure decisions and scaling strategy
Mentor engineers and raise the bar on reliability and engineering standards
Embed security into infrastructure and deployment workflows
Design for resilience (disaster recovery, chaos testing, capacity planning)
8+ years of experience in DevOps, SRE, or Infrastructure Engineering
Proven experience operating large-scale, distributed systems in production
Deep expertise in:
Kubernetes & container orchestration Pulumi (or similar IaC tools) Cloud or hybrid environments (AWS, GCP, Azure, or on-prem) Observability stacks (Prometheus, Grafana, OpenTelemetry)
Kubernetes & container orchestration
Pulumi (or similar IaC tools)
Cloud or hybrid environments (AWS, GCP, Azure, or on-prem)
Observability stacks (Prometheus, Grafana, OpenTelemetry)
Strong experience with CI/CD, automation, and release engineering
Proficiency in Python, Go, or Bash
Strong systems thinking and debugging skills in high-scale environments
Experience defining and operating with SLOs / SLAs
Experience in startup environments
Comfortable leveraging AI coding tools and agents to move faster
Experience with AI/ML infrastructure or GPU workloads
Familiarity with distributed or high-performance compute systems
Exposure to platform engineering / internal developer platforms
Experience scaling systems from Beta to production
Work on cutting-edge AI infrastructure
Build systems that power developers and enterprises
High ownership, fast execution, real impact
Collaborative, high-caliber team
Universal AI compute infrastructure for developers and enterprises.
Visit company websiteJobs and hiring trendsFull-time
Senior · 8+ years experience
Onsite
Apply faster on company sites with our extension.