{bc}
ripplehire

Tech & Digital-Lead Site Reliability Engineer

HDFC Bank
Bengaluru, IND
Full-time
Senior · 11+ years experience
Onsite
Discovered 6 days ago
DockerKubernetesAWSTerraformAnsibleKafka
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

DockerKubernetesAWS
Smart Apply

Full Job Posting

Job Details

Location: Bangalore

Experience: 0 - 1 Years

Business Unit: Tech & Digital

Job Type: Permanent

Openings: 1

Role Description

Job Title:

Lead Site Reliability Engineer

Job Details:

Team: Ent Factory-Channels, Mobility, Payments

Reports to: SRE Manager

Location: Mumbai, Chennai, Gurgaon & Bangalore

Role Type: Non-Supervisory

No of direct reportees : 0

Travel Required: No

Job Band Range: D1/D2

JD Created date: 28th Jan 2023

Job Purpose:

Analyzing , troubleshooting, and designing vital services, platforms, and infrastructure on GCP with a focus on reliability, scalability, resilience, security, and performance.

Lead and Men tor a team of SRE engineers.

Job Responsibilities:

Execute reliability initiatives for the team and organization.

Mentor and lead a team of SRE engineers.

Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentat ion, and code with other engineering teams.

Solid understanding of observability tools and ability to express reliability metrics via observability.

Define KPI in the form of RPO/RTO/SLI/SLO/Error Budget.

Apply automation and software to any manually perfo rmed tasks or system parts.

Troubleshoot complicated, cross-platform issues handling OS, Networking, Database in a cloud-based SaaS environment and manage live production incidents.

Monitor application performance, take steps to improve overall application performance and stability, and follow through with implementation.

Conduct system analysis, configuration management, and develop improvements for system software performance, availability, and reliability.

Design, write, ship, and motivate the creation o f software and systems to increase observability, product reliability, and organizational efficiency.

Maintain and monitor deployment, orchestration of servers, docker containers, databases, and general backend infrastructure.

Develop Run Books/Standard Operating Procedure for recurring Production issues and work on permanent solutions.

Perform Incident Analysis regularly to prevent and find long-term solutions for Incidents.

Educational Qualifications:

B Tech in Computer Scie nce or related discipline preferred.

Key Skills:

Experience in monitoring and analyzing infrastructure performance using standard performance monitoring tools.

Demonstrable experience in Containerization (Docker) and orchestration (Kubernetes).

Experienc e with Infrastructure As Code (Terraform, Cloud Formation, Ansible).

Knowledge and proven hands-on experience in large-scale databases and distributed technologies, such as Kafka and Confluent Platform Kafka.

Basic programming and scripting skills.

Solid u nderstanding of at least 2 observability technologies.

Experience Required:

Total Yrs of experience: 11-13

Major Stakeholders:

Internal: Product Manager from Digital Factory, Business Analyst from BTG team, Incident Management team, Development Team.

Skills

  • <p><span data-teams="true">Docker, Kubernates, AWS, Monitoring, terraform, ansible, networking, reliability engineering</span></p>

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at HDFC Bank