{bc}
oracle

Open to Jobshare - Sr Lead Software Engineer - Site Reliability Engineer, Python & Infrastructure management - Part time/Jobshare

J.P. Morgan
London, GBR
Part-time
Senior · 10+ years experience
Onsite
Discovered 2 weeks ago
PythonPrometheusDatadogDynatraceSplunkAthena
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

PythonPrometheusDatadog
Smart Apply

Full Job Posting

Job responsibilities

  • Defines the SRE vision, north-star outcomes, and multi-year roadmap for the Production Management team, aligned to both CIB and JPM Global Technology priorities.
  • Establishes the SRE operating model across global regions (ways of working, intake, prioritization, engagement with engineering teams and production support).
  • Partners with business-aligned Production Support leads to embed SRE practices consistently and act as a force multiplier – coaching them on reliability thinking, prioritization, and “engineering out” operational load.
  • Builds and develops a small, high-impact core SRE team (and/or virtual SRE community of practice) that scales reliability improvements across many application flows.
  • Defines and implements standards for: Service cataloging, SLO/SLI frameworks and error budgets, incident response maturity, blameless post-incident reviews, resiliency patterns, capacity, performance, and scalability engineering.
  • Drives service reviews with evidence-based reporting (availability, latency, incident trends, MTTR/MTTD, change failure rate, customer impact).
  • Champions AI adoption and deliver AI-enabled capabilities to reduce operational toil and improve speed/quality of response.
  • Sets direction for observability across logs/metrics/traces, including instrumentation standards, golden signals, and end-user journey monitoring. Improves alert quality and routing: reduce false positives, improve actionable alerts, and tighten feedback loops to engineering teams.
  • Builds strong partnerships with application development teams, platform/infrastructure partners, and governance functions. Communicates clearly and credibly at all levels—from engineers to senior technology and business stakeholders.
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.

Required qualifications, capabilities, and skills

  • 10+ years of experience in technology support, production/application support, DevOps, or infrastructure management.
  • Demonstrated experience leading SRE/reliability engineering or production engineering transformations in a complex enterprise environment.
  • Strong engineering background: ability to design, build, and deliver automation and reliability solutions.
  • Fluency & expertise in Python
  • Deep practical knowledge of: SLOs/SLIs, error budgets, incident management, postmortems, observability design across metrics/logs/traces and distributed systems troubleshooting, resilience engineering, performance/capacity management, and change risk reduction.
  • Proficiency and experience with telemetry (logs/metrics/traces) collection using tools and standards such as Prometheus, Open Telemetry, Datadog, Dynatrace, Splunk.
  • Experience delivering automation at scale (scripting, workflow automation, runbook automation, CI/CD-integrated guardrails).
  • Proven leadership skills: influencing without authority, coaching leaders, and building communities of practice.
  • Strong judgment around risk, security, and controls—especially when applying AI to production workflows.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.

Preferred qualifications, capabilities, and skills

  • Familiarity with Athena / prior experience in Athena
  • Your Pathway to Jobsharing at JPMorgan – IT’S A 2 BRAINER!
  • At JPMorgan we are passionate about supporting different ways of working to support our talent in the flexibility they need.
  • We know that Jobshare is a fantastic way to hire talent for the firm that offers both the flexibility that you need whilst providing the consistency that our business requires. Jobshare is 2 people working part time hours with full time powers. This role is part of a Jobshare o pportunity and therefore a part-time role at 19hours

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at J.P. Morgan