Leads, coach, and develop a high-performing Production Management Engineering team in Bangalore (hiring, performance management, career development, succession planning). Establishes clear team operating rhythms (shift/on-call expectations where applicable), decision rights, escalation paths, and accountability. Build a culture of operational excellence: documentation discipline, continuous improvement, and rigorous follow-through on actions.
Drives effective stakeholder management across engineering, product, security, infrastructure, and business teams; manage competing priorities and align execution to organizational goals. Owns service availability and reliability targets (e.g., 99.999% where applicable) and define service tiers/criticality for the platform.
Establishes and drives a reliability engineering roadmap: automation/self-healing, reduced toil, improved observability, and improved recoverability. Set production standards for monitoring, alerting, logging, runbooks, and production readiness gates for onboarding new apps/services.
Oversees end-to-end management of production and UAT environments, ensuring availability, resiliency, performance, and security across application and infrastructure layers. Establish and monitor KPIs for production systems (availability, latency, MTTR, incident rates, change success rate, capacity/utilization, and control/compliance evidence health), using SLI/SLO frameworks to proactively prevent customer impact..
Acts as escalation leader for Sev1/Sev2 incidents and high-risk production events; ensure rapid mobilization across global teams. Drive high-quality root cause analysis and corrective actions with measurable outcomes and verified closure.
Leads governance forums (service health reviews, incident/problem reviews, change quality, operational risk) and provide executive-level reporting. Ensures processes and solutions adhere to regulatory requirements and internal policies, maintaining robust compliance and security standards (including IAM lifecycle controls).
Partners with application and platform engineering leadership to improve architecture resiliency, reduce blast radius, and remove single points of failure. Evaluate cross-LOB impact of platform changes and data/process initiatives; ensure operational readiness and safe delivery. Sponsor modernization of tooling and operational analytics to improve MTTR and prevent recurrence.
Champions automation and DevOps practices (CI/CD enablement, infrastructure-as-code, standardized release/run pipelines) to streamline operations and reduce manual intervention and toil. Stay abreast of industry trends and emerging technologies in SRE/production operations, IAM, and cloud infrastructure, recommending enhancements to production management practices. Own DR/SR/HA strategy, testing cadence, evidence capture, and outcomes; ensure documentation is current and executable. Proactively manage capacity, performance, and dependency risks through forecasting, testing, and engineered action plans.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
We have an exciting and rewarding opportunity for you to take your software engineering career to the next level. As a Software Engineer III at JPMorganChase within the Chief Technology Office (cto)Enterprise Observabili
Associate Banker, 40 Hours, Kansas City, MO - Shoal Creek Branch
Kansas City, USA
Mid
At Chase, we are passionate about creating memorable experiences for our clients and employees, making them feel welcomed, valued, and understood. We build lasting relationships by doing the right thing, exceeding expect
At Chase, we are passionate about creating memorable experiences for our clients and employees, making them feel welcomed, valued, and understood. We build lasting relationships by doing the right thing, exceeding expect
Compliance Risk Mgmt Lead – Employee Compliance O&C – Vice President
Jersey City, USA
Executive
Bring your expertise to JPMorganChase’s Conduct Risk Program. As part of Risk Management and Compliance, you are at the center of keeping JPMorgan Chase strong and resilient. You help the firm grow its business in a resp
Part Time (30 Hours) Associate Banker, Middleton Burleys Corner Branch, Middleton, MA
Middleton, USA
Mid
At Chase, we are passionate about creating memorable experiences for our clients and employees, making them feel welcomed, valued, and understood. We build lasting relationships by doing the right thing, exceeding expect
Public Sector Group Business Management - Associate
, USA
Mid
Join a high-impact Business Management team supporting the Public Sector Group within Global Corporate Banking and play a pivotal role in optimizing outcomes for a complex, multi-faceted client segment. You’ll collaborat
Part Time (30 Hours) Associate Banker, Gratiot Warren Branch, Detroit, MI
Detroit, USA
Mid
At Chase, we are passionate about creating memorable experiences for our clients and employees, making them feel welcomed, valued, and understood. We build lasting relationships by doing the right thing, exceeding expect
Associate Banker, 30 Hours, Saint Cloud, MN - Downtown St Cloud Branch
Saint Cloud, USA
Mid
At Chase, we are passionate about creating memorable experiences for our clients and employees, making them feel welcomed, valued, and understood. We build lasting relationships by doing the right thing, exceeding expect
Leads AWS-based resiliency design and execution where applicable (multi-AZ/multi-region patterns, cloud-native scaling and recovery approaches), partnering with security to ensure cloud security best practices and risk management are embedded.
Establishes governance standards for AI-assisted workflows used in incident/problem/change processes (including documentation and trend analysis), ensuring traceability/auditability and alignment to resiliency and security expectations.
Sets reuse-first expectations for enterprise-authorized AI adoption within the work environment across support operations to accelerate incident insights and operational reporting, with human-in-the-loop validation and appropriate handling of sensitive data.
7+ years of experience or equivalent expertise troubleshooting, resolving, and maintaining information technology services, including leadership accountability for global production environments and reliability transformation initiatives, with demonstrated delivery in highly regulated environments and on cloud (AWS) and IAM-critical platforms.
Demonstrated experience leading safe use of enterprise-authorized AI capabilities within the work environment within production support workflows, including validation practices and awareness of data sensitivity.
Ability to define review/approval and escalation expectations for AI-assisted recommendations while maintaining operational, security, and auditability outcomes.
Executive-grade written and verbal communication; calm, decisive leadership during major incidents. Demonstrated ability to lead large-scale, complex production environments in a fast-paced organization; ability to drive strategic initiatives and manage competing priorities effectively.
Strong collaboration skills to work with technical experts, key stakeholders, and team members to resolve complex problems.
Strong understanding of Linux-based platforms, web/app architectures, and distributed system reliability patterns. Expertise in observability practices and tools (logs/metrics/traces, alert design, dashboards; e.g., Splunk/Geneos/Dynatrace/Grafana/ServiceNow), including configuration/integration and metric analysis across application and infrastructure layers.
Proven knowledge of site reliability culture and principles, including use of SLIs/SLOs to drive proactive operations and reliability improvements. Extensive experience with AWS architecture and services (e.g., EC2, IAM, S3, Lambda, CloudFormation) and cloud security best practices; strong understanding of cloud security principles, compliance frameworks, and risk management.
Proven expertise in Identity and Access Management (IAM), including onboarding/support strategy, policy design, implementation, and audit; familiarity with IAM platforms/tools such as ForgeRock and Ping Identity.
Proficiency in at least one programming language such as Python or Java/Spring Boot, plus strong command of Linux scripting for day-to-day operational enablement and automation.
Experience
with automation tooling.
DevOps methodologies, and CI/CD pipelines; familiarity with enterprise scheduling/orchestration tools such as Autosys or Control-M.
Demonstrated experience leading safe use of enterprise-authorized AI capabilities within the work environment within production support workflows, including validation practices and awareness of data sensitivity.
Ability to define review/approval and escalation expectations for AI-assisted recommendations while maintaining operational, security, and auditability outcomes.
Familiarity with container and container orchestration technologies such as Docker, ECS, and Kubernetes.
Familiarity with troubleshooting common networking technologies and issues, plus internet security/network fundamentals.
Familiarity with databases and messaging/integration technologies.
About JPMorgan Chase
Banking323028 employees
Global financial services and investment banking firm.