Base Career helps you apply smarter for this job.
Key skills for this role
FlexAI is looking for a Senior DevOps / SRE Engineer to build and operate the infrastructure powering our AI and PaaS platform.
You’ll work closely with developers to ensure our systems are reliable, performant, and scalable , while enabling fast product iteration. This role is hands-on and execution-focused, with opportunities to contribute to system design and reliability practices as we scale.
FlexAI is looking for a Senior DevOps / SRE Engineer to build and operate the infrastructure powering our AI and PaaS platform.
You’ll work closely with developers to ensure our systems are reliable, performant, and scalable , while enabling fast product iteration. This role is hands-on and execution-focused, with opportunities to contribute to system design and reliability practices as we scale.
Build and maintain infrastructure for our AI and PaaS platform
Deploy and operate Kubernetes clusters and containerized services
Implement Infrastructure as Code using Pulumi (or similar tools)
Help define and implement SLIs, SLOs, and error budgets
Improve system reliability, availability, and performance
Participate in on-call rotations , incident response, and postmortems
Build and improve CI/CD pipelines for reliable and fast releases
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Santa Clara, USA
Santa Clara, USA
Santa Clara, USA
San Jose, USA
Bengaluru, IND
Bengaluru, IND
Bengaluru, IND
Automate operational workflows and reduce manual toil
Contribute to GitOps and platform engineering practices
Implement and maintain observability using VictoriaMetrics, Grafana (metrics, logs, traces)
Monitor systems and troubleshoot performance issues (latency, throughput, cost)
Work closely with developers, platform, and AI teams to support production systems
Help debug issues across infrastructure and application layers
Contribute to improving engineering productivity and developer experience
What You’ll Need to Be Successful
4+ years of experience in DevOps, SRE, or Infrastructure Engineering
Experience operating production systems at scale
Hands-on experience with:
Kubernetes & containers Infrastructure as Code (Pulumi, Terraform, etc.) Cloud or hybrid environments (AWS, GCP, Azure, or on-prem) Observability tools (Prometheus, Grafana, OpenTelemetry)
Kubernetes & containers
Infrastructure as Code (Pulumi, Terraform, etc.)
Cloud or hybrid environments (AWS, GCP, Azure, or on-prem)
Observability tools (Prometheus, Grafana, OpenTelemetry)
Experience with CI/CD systems and automation
Proficiency in Python, Go, or Bash
Strong debugging and problem-solving skills
Familiarity with SLOs and reliability practices
Experience working in startup or fast-paced environments
Comfortable leveraging AI coding tools and agents
Experience with AI/ML infrastructure or GPU workloads
Familiarity with distributed systems or compute platforms
Exposure to platform engineering concepts
Experience supporting systems from Beta to production
Work on cutting-edge AI infrastructure
Build systems that power developers and enterprises
High ownership, fast execution, real impact
Collaborative, high-caliber team
Universal AI compute infrastructure for developers and enterprises.
Visit company websiteJobs and hiring trendsFull-time
Senior · 4+ years experience
Onsite
Apply faster on company sites with our extension.