Base Career helps you apply smarter for this job.
Key skills for this role
Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering • Job requisition ID : 108935 • Location : Bengaluru • Entity : Deloitte Touche Tohmatsu India LLP
Engineering helps Reimagine and re-engineer mission-critical operations and processes; Leverage engineering-led design, deep industry knowledge, and AI and data-driven insights to transform the technology platforms at the heart of business.
Working alongside team, we empower and drive mission-critical solutions whether we need to modernize existing systems or implement new technology products and platforms. Through innovation, we improve financial performance, accelerate new digital businesses and fuel growth. Learn more about Engineering, AI and Data
We are looking for a highly skilled Site Reliability Engineer (SRE) to manage and scale mission-critical, production-grade distributed systems running on Google Cloud Platform (GCP). The ideal candidate will focus on reliability, automation, observability, and operational excellence while minimizing toil and improving system availability. Maintaining and improving 4 Nines of uptime to 5 Nines with engineering efforts.
This role requires deep technical expertise in cloud-native technologies, Kubernetes, infrastructure automation, Linux administration and TCP/IP fundamentals, programming in one language and strong troubleshooting capabilities for distributed systems. The candidate needs to participate in the overall lifecycle management of mission critical banking services with a 24x7 operations mode in an rotational on-call basis. The job requires the candidate to have strong troubleshooting skills in a distributed environment spread across multiple cloud environments. The bare minimum ask would be to maintain high level of agility, learnability and adaptability in different scenarios. An engineer with a zeal to learn fast and having a bias for action would be the best fit for the role.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
, CAN
R8dius recherche un conseiller principal pour piloter des stratégies de gestion du changement liées au déploiement de solutions technologiques. La personne élaborera des plans d’adoption, de communication et de formation
Toronto, CAN
Deloitte recherche un directeur ou une directrice pratique en ingénierie et développement de produits pour concevoir des applications par pile complète et diriger des missions de transformation technologique. La personne
Toronto, CAN
Deloitte is seeking a hands-on Manager, Product Engineering & Development to build full-stack software, shape client solutions, lead pursuits, and remain accountable through delivery. The role requires recent full-stack
, IND
Deloitte is seeking a Senior Consultant to lead treasury reporting for mutual funds, private equity, hedge funds, and related vehicles. The role covers cash forecasting, reconciliations, financing activity, controls, aud
, CAN
Deloitte is seeking a Senior Manager to lead strategy, delivery, and operational execution for its Global Contact Center technology portfolio. The role requires 15+ years of technology delivery experience, contact center
, USA
, USA
Jersey City, USA
, CAN
Toronto, CAN
Toronto, CAN
, IND
, CAN
Reliability & Operations
Own end-to-end production systems reliability, availability, scalability, cost and performance.
Drive measurable improvements in MTTR, MTTA, and incident response practices using automation and runbook additions and process enhancements.
Participate in 24x7 on-call rotations and handle high-severity incidents and document the learnings on ongoing basis.
Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics for mission critical services and partner with engineering teams with full accountability for upholding the SLOs.
Partner with the various engineering, operations and cloud management teams to deliver highly reliable service in a timely manner.
Cloud & Infrastructure
Design, deploy, and manage infrastructure on Google Cloud Platform (GCP).
Work extensively on: GKE (Kubernetes Engine), Compute, networking, IAM, Load Balancers, TLS Certs, BigQuery, Pub/Sub, cloud logging enhancement, metrics and logs analysis
Implement and manage infrastructure using Terraform (Infrastructure as Code).
Kubernetes & Containers
• Deploy and manage containerized workloads using Kubernetes (GKE).
• Troubleshoot issues related to: Pods, nodes, networking, storage, services on an ongoing basis
• Manage deployments using Helm, YAML, and rollout strategies (Canary/Blue-Green).
Automation & CI/CD
Build and maintain CI/CD pipelines using: Jenkins (pipeline-based, Groovy / Shell / Python scripting) Strong experience in using GitHub as a PowerUser
Develop automation using Python and Shell scripting.
Reduce operational toil through automation initiatives.
Observability & Monitoring
Implement and manage monitoring systems using: Dynatrace, Grafana, logs and metrics explorer
Work with logs, metrics, and traces for deep observability to identify trends and arrest problems proactively.
Define alerting strategies based on system behaviour and SLOs and create runbooks.
Work alongside operations teams to identify, fix the production incidents and own the problem resolution.
Work with engineering teams to isolate infra and application issues and set up right tooling for debugging production incidents.
System & Application Troubleshooting
Perform deep troubleshooting for: Distributed systems, Microservices-based architectures on containerised workloads, Java and Golang applications
Strong debugging of: Application issues, Infrastructure issues, Network-related problems. Plan and execute continuous improvement
Identify and eliminate repetitive manual tasks.
Drive reliability engineering practices and culture. (DRY – Don’t Repeat Yourself)
Collaborate with development teams to improve system design and resilience.
Professional services firm providing audit, consulting, and advisory services.
Visit company websiteFull-time
Mid · 4+ years experience
Onsite
Apply faster on company sites with our extension.