{bc}
indeed

Software Engineering Tech Lead (SRE + AI)

Cisco
London, GBR
Full-time
Onsite
Discovered 1 weeks ago
Site Reliability EngineeringProduction intelligenceAgentic AILLM pipelinesAI agentsModel Context Protocol
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site Reliability EngineeringProduction intelligenceAgentic AI
Smart Apply

Full Job Posting

Meet the Team

Cisco’s Collaboration Technology Group builds services that connect people across devices, locations and time zones.

The team operates platform services for collaboration products at global scale across multiple datacentres.

Your Impact

Lead the architectural vision and implementation of an AI-powered Production Intelligence platform.

Combine Site Reliability Engineering with agentic AI, reusable skills, MCP and LLM tooling.

Improve how engineering and leadership teams monitor, diagnose and automatically remediate global SaaS infrastructure.

What You’ll Do

  • Define architecture for AI-assisted observability, automated incident response and self-healing cloud infrastructure.
  • Build production-grade AI agents, MCP integrations and deterministic evaluation pipelines.
  • Architect telemetry ingestion and correlation across logs, metrics, traces, changes and runbooks.
  • Develop anomaly detection and human-in-the-loop remediation with safety guardrails.
  • Define SLIs and SLOs, manage error budgets and lead post-incident reviews.
  • Partner with application and infrastructure teams on reliability and scalability.
  • Mentor engineers and establish best practices across global development and operations teams.
  • Manage priorities, deadlines and communication across teams.

Minimum Qualifications

  • Bachelor’s degree plus 8 years, Master’s degree plus 6 years, or PhD plus 3 years of related technical experience.
  • Proven Technical Lead, Lead SRE or Lead Software Engineer experience delivering distributed, highly available SaaS platforms at scale.
  • Strong proficiency in Python, Go, Java or C++ and experience designing microservices, APIs and production automation.
  • Deep Kubernetes, Docker and container orchestration experience in large-scale multi-cluster environments.
  • SRE experience covering SLI/SLO design, observability, incident management and automated root-cause analysis.

Preferred Qualifications

  • Hands-on experience with LLM pipelines, AI agents, MCP, RAG architectures and evaluation frameworks.
  • Experience with OpenTelemetry, Prometheus, Grafana, Splunk, ThousandEyes or distributed tracing.
  • Expertise in AWS, GCP or Azure, Terraform, infrastructure as code, GitOps or CI/CD.
  • Experience with responsible AI guardrails, deterministic fallbacks and policy-driven remediation.
  • Experience with Kafka, Redis, PostgreSQL, Elasticsearch or vector databases.

Why Cisco

Cisco develops technology that connects and protects organizations across physical and digital environments.

The company emphasizes collaboration, innovation, empathy and large-scale impact.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Cisco