{bc}
linkedin

Senior Site Reliability Engineer

Cisco
Bengaluru, IND
Full-time
Mid-Senior
Onsite
Discovered 1 weeks ago
Site Reliability EngineeringAWSApache AirflowAmazon EKS and KubernetesTerraformPython
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site Reliability EngineeringAWSApache Airflow
Smart Apply

Full Job Posting

Team overview

The Network Assurance Data Platform team within Cisco ThousandEyes builds and operates data infrastructure for network assurance, analytics, and intelligence.

The team manages cloud infrastructure, big data workflows, ML platform components, and reliability engineering practices across AWS environments.

Role overview

The Senior Site Reliability Engineer will operate scalable infrastructure supporting data pipelines, ML workloads, cloud-native services, and cost-optimized AWS operations.

The role combines reliability engineering, data infrastructure operations, automation, FinOps, and technical leadership.

Core responsibilities

  • Own reliable, scalable, and cost-efficient infrastructure for the Network Assurance Data Platform.
  • Manage Airflow, EMR, Spark, Hadoop, EKS, containerized services, ML workloads, and production platform components.
  • Build Terraform automation and Python tooling for integrations, reporting, operations, and reliability improvements.
  • Drive FinOps through cost visibility, allocation, forecasting, anomaly detection, optimization, and governance.
  • Improve observability, alerting, incident response, capacity planning, autoscaling, storage optimization, and workload efficiency.
  • Partner across data, ML, platform, finance, and product teams and provide technical leadership and mentorship.

Minimum qualifications

  • A bachelor's degree or higher in Engineering, Computer Science, or equivalent practical experience.
  • 8-10 years of relevant experience in SRE, DevOps, cloud infrastructure, platform engineering, data infrastructure, or production engineering.
  • Strong production AWS, Apache Airflow, EKS or Kubernetes, Terraform, Python, FinOps, cloud cost optimization, Linux, networking, and distributed systems experience.
  • Experience with AWS EMR, Spark, Hadoop, containerized workloads, observability, incident management, production troubleshooting, and operational excellence.

Preferred qualifications

  • Experience with large-scale SaaS platforms, ML infrastructure, batch processing, and optimization of EMR, Spark, Airflow, EKS, storage, and compute workloads.
  • Experience with AWS cloud-native services, Savings Plans, Reserved Instances, Spot adoption, Graviton migration, lifecycle management, and right-sizing.
  • Experience with cost management tools, CI/CD systems, Puppet, Ansible, Helm, Argo CD, FinOps programs, stakeholder reporting, and executive updates.
  • FinOps certification or equivalent cloud financial management experience is a plus.

Success measures

Data and ML infrastructure is reliable, scalable, performant, and cost-efficient.

Airflow, EMR, Spark, and EKS workloads have strong observability, automation, and production maturity.

AWS cost visibility, forecasting, governance, and workload efficiency improve while cloud wastage decreases.

Engineering and data teams receive clear visibility into platform health, cost drivers, risks, and optimization opportunities.

Why Cisco

Cisco describes a global technology environment focused on connecting and protecting organizations in the AI era.

The company emphasizes innovation, collaboration, and opportunities to grow while delivering security, visibility, and infrastructure solutions.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Cisco