Software Engineering Tech Lead (SRE + AI)
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Lead the architectural vision and implementation of an AI-powered Production Intelligence platform.
Combine Site Reliability Engineering with agentic AI, reusable skills, MCP and LLM tooling.
Improve how engineering and leadership teams monitor, diagnose and automatically remediate global SaaS infrastructure.
Key Skills for This Role
Full Job Posting
Meet the Team
Cisco’s Collaboration Technology Group builds services that connect people across devices, locations and time zones.
The team operates platform services for collaboration products at global scale across multiple datacentres.
Your Impact
Lead the architectural vision and implementation of an AI-powered Production Intelligence platform.
Combine Site Reliability Engineering with agentic AI, reusable skills, MCP and LLM tooling.
Improve how engineering and leadership teams monitor, diagnose and automatically remediate global SaaS infrastructure.
What You’ll Do
- Define architecture for AI-assisted observability, automated incident response and self-healing cloud infrastructure.
- Build production-grade AI agents, MCP integrations and deterministic evaluation pipelines.
- Architect telemetry ingestion and correlation across logs, metrics, traces, changes and runbooks.
- Develop anomaly detection and human-in-the-loop remediation with safety guardrails.
- Define SLIs and SLOs, manage error budgets and lead post-incident reviews.
- Partner with application and infrastructure teams on reliability and scalability.
- Mentor engineers and establish best practices across global development and operations teams.
- Manage priorities, deadlines and communication across teams.
Minimum Qualifications
- Bachelor’s degree plus 8 years, Master’s degree plus 6 years, or PhD plus 3 years of related technical experience.
- Proven Technical Lead, Lead SRE or Lead Software Engineer experience delivering distributed, highly available SaaS platforms at scale.
- Strong proficiency in Python, Go, Java or C++ and experience designing microservices, APIs and production automation.
- Deep Kubernetes, Docker and container orchestration experience in large-scale multi-cluster environments.
- SRE experience covering SLI/SLO design, observability, incident management and automated root-cause analysis.
Preferred Qualifications
- Hands-on experience with LLM pipelines, AI agents, MCP, RAG architectures and evaluation frameworks.
- Experience with OpenTelemetry, Prometheus, Grafana, Splunk, ThousandEyes or distributed tracing.
- Expertise in AWS, GCP or Azure, Terraform, infrastructure as code, GitOps or CI/CD.
- Experience with responsible AI guardrails, deterministic fallbacks and policy-driven remediation.
- Experience with Kafka, Redis, PostgreSQL, Elasticsearch or vector databases.
Why Cisco
Cisco develops technology that connects and protects organizations across physical and digital environments.
The company emphasizes collaboration, innovation, empathy and large-scale impact.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at Cisco
Consulting Engineer
, IND
Cisco is seeking a senior Consulting Engineer to lead complex Data Center network problem resolution and premium technical support for high-demand customers. The role requires deep Cisco Data Center expertise, advanced n
Virtual Sales Account Executive - Splunk Defense Joint Forces (Remote)
, USA
Customer Project Manager | 9+ Years
, IND
Consulting Engineer
, IND
Paid Media Marketing Manager
, USA
Virtual Sales Account Executive - Splunk Defense Joint Forces (Remote)
, USA
Solutions Engineer - Texas - ThousandEyes (Remote)
, USA
Strategy & Planning Manager, AI Product Operations
, USA
Software Engineer
Bengaluru, IND
Data Engineer
Mumbai, IND