Base Career helps you apply smarter for this job.
Key skills for this role
Principal infrastructure Engineer – Technical Major Incident Manager
The principal infrastructure engineer is responsible for leading and managing high-priority (P1/P2/P3) infrastructure incidents impacting critical business services. The role serves as the central point of coordination during major incidents, driving technical bridge calls, facilitating collaboration across resolver teams, managing escalations, and ensuring the timely restoration of services.
The Major Incident Manager is accountable for coordinating cross-functional teams, including Server, Virtualization, Unix, Network, Database, Storage, Backup, Security and other Support teams, to identify root causes, restore services, and minimize business impact. The individual must be capable of driving incident calls, maintaining executive-level communications, managing stakeholder expectations, and ensuring all incident activities are executed within established service management processes.
Position Title: Principal Infrastructure Engineer
Career Level: P4
Job Category: Assistant Vice President
Role Type: Hybrid
Job Location: Bangalore
The Technical Major Incident Management (MIM) team provides 24x7 global coverage across India and the United States, leading the response and recovery of critical technology incidents. The team partners with Infrastructure, Engineering, Cybersecurity, SRE and other Support teams to ensure service availability, operational resilience and continuous improvement across enterprise platforms
Key Deliverables (Duties and Responsibilities)
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Bengaluru, IND
, USA
Boston, USA
Bengaluru, IND
Bengaluru, IND
, USA
Charlotte, USA
Atlanta, USA
, USA
Lead and manage P1/P2/P3 major incident bridge calls, coordinating cross-functional technical teams to ensure timely service restoration
Drive incident resolution efforts, escalations, and stakeholder communications throughout the incident lifecycle
Facilitate Post-Incident Reviews (PIRs) and Root Cause Analysis (RCA) activities, ensuring corrective actions are tracked to closure
Publish major incident reports, trend analysis, and service improvement recommendations
Drive continuous service improvement initiatives focused on reducing incident volumes, minimizing business impact, and improving operational stability
Lead and manage P1/P2/P3 major incident bridge calls, ensuring timely service restoration and minimal business impact
Coordinate cross-functional infrastructure teams across Windows, Unix, Network, Database, Storage, Backup domains, Container Platform during critical incidents
Drive technical operations execution by ensuring operational activities, service recovery actions, escalation procedures, and incident response processes are executed effectively and within SLA targets
Facilitate stakeholder communications, executive updates, and incident governance throughout the incident lifecycle
Lead Post-Incident Reviews (PIRs), Root Cause Analysis (RCA), and corrective action tracking to prevent recurring incidents
Deliver incident reporting, operational trend analysis, and continuous service improvement initiatives to enhance service stability and operational efficiency
Prepare and publish major incident reports, dashboards, and operational metrics for leadership review
Ensure compliance with Major Incident Management, Problem Management, and ITIL governance processes
Track and report key performance indicators (KPIs) such as MTTR, incident trends, SLA compliance, and recurring incidents
Facilitate Post-Incident Reviews (PIRs) and ensure Root Cause Analysis (RCA) reports and corrective actions are completed within agreed timelines
Provide regular updates to stakeholders and senior management on incident performance, risks, and service improvement initiatives
12+ years of experience in Technology Operations, Infrastructure Operations, Production Support, Major Incident Management, Site Reliability Engineering (SRE), Network Operations, or Enterprise Technology Services
Proven experience managing and leading P1/P2/P3 major incidents in large-scale enterprise environments
Strong experience coordinating cross-functional teams across Windows, Unix/Linux, Virtualization, Network, Database, Storage, Backup, Cloud and Security domains
Experience driving technical bridge calls, incident escalations, stakeholder communications, and service restoration activities
Demonstrated experience in Problem Management, Root Cause Analysis (RCA), Post-Incident Reviews (PIRs), and Continuous Service Improvement initiatives
Strong understanding of ITIL Incident, Problem, Change, and Service Management processes.
Experience in executive reporting, operational governance, KPI tracking, and service performance management
Proven ability to lead technical operations during critical incidents while maintaining focus on business impact, customer experience, and service availability
Hands-on experience in at least three of the following technology domains:
Windows Server Administration
Virtualization Platforms (VMware vSphere, ESXi, vCenter, Hyper-V)
Unix Systems (Linux & AIX)
Network Services & Platforms: Routing & Switching, DDI (Infoblox, EfficientIP), Load Balancers (F5, A10), Cisco ISE/NAC, Wireless (Aruba, Cisco), Gluware, and Ansible Automation Platform
Database (Oracle, MSSQL, PostgresSQL)
Storage & Backup Technologies (SAN, NAS, NetApp, EMC, Veeam, Commvault)
Container platforms
Experience with automation and operational tooling is preferred
Working knowledge of Ansible, Python, Shell Scripting, and PowerShell
Ability to identify and implement automation opportunities that improve incident response, operational efficiency, and service reliability
Artificial Intelligence & Productivity Tools
Experience using Generative AI (GenAI) tools to enhance incident analysis, reporting, knowledge management, and operational effectiveness
ITIL Foundation (V3/V4 or higher)
Site Reliability Engineering (SRE) Certification
DevOps Certifications
Six Sigma Green Belt or higher
CCNA\ VMware Certified\ RHEL Certified
Reports to: Director - Technology
Partners: Senior leaders and cross-functional teams across geographies (India and US team)
We are committed to providing an inclusive and accessible hiring process. If you require accommodation at any stage (e.g. application, interviews, onboarding) please let us know, and we will work with you to ensure a seamless experience
FC Global Services India LLP (First Citizens India) is an Equal Employment Opportunity Employer. We are committed to fostering an inclusive and accessible environment and prohibit all forms of discrimination on the basis of gender, religion, caste, disability, sexual orientation, economic status or any other characteristics protected by the law. We strive to foster a safe and respectful environment in which all individuals are treated with respect and dignity. Our EEO policy ensures fairness throughout the employee life cycle.
Financial institution providing personal, business, and commercial banking services.
Visit company websiteJobs and hiring trendsFull-time
Senior · 12+ years experience
Hybrid
Apply faster on company sites with our extension.