{bc}
indeed

Senior Lead Site Reliability Engineer - Cloud Data & Databricks

JPMorganChase
IND
Full-time
Onsite
Discovered 2 days ago
Site reliability engineeringAWS Databricks platform administrationAWS cloud architecturePython application developmentAutomated unit testingMonitoring and observability tools
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site reliability engineeringAWS Databricks platform administrationAWS cloud architecture
Smart Apply

Full Job Posting

Role Overview

JPMorgan Chase is hiring a Senior Lead Site Reliability Engineer within its AIML Data Platforms and Chief Data and Analytics Team.

The role focuses on advanced technology products, cloud data platforms, data quality and security, AI/ML enablement, and operational excellence.

Job Responsibilities

  • Lead a team of SREs designing, implementing, and maintaining a managed AWS Databricks platform.
  • Provide engineering and operational support to application and engineering teams.
  • Perform platform design, setup, configuration, workspace administration, and resource monitoring.
  • Create multi-availability-zone, multi-region, and multi-cloud resiliency strategies.
  • Evaluate architectural designs and technical credentials with vendors, startups, and internal teams.
  • Improve system observability, alerting, and capacity planning.
  • Optimize infrastructure and deployment processes through automation and operational excellence.
  • Execute software design, development, and technical troubleshooting.
  • Use authorized AI capabilities for reliability design, incident analysis, and requirements traceability.
  • Lead AI-assisted reliability workflows with traceability, auditability, resiliency, and security controls.
  • Develop secure production code and review or debug code written by others.
  • Apply SRE practices, automate recurring issues, and maintain incident response and postmortem procedures.

Required Qualifications and Skills

  • Formal training or certification in software engineering concepts and 10+ years of applied experience.
  • Strong understanding of SRE principles, including SLIs, SLOs, error budgets, and incident management.
  • Experience with monitoring tools, automation frameworks, and CI/CD pipelines.
  • Proficiency in Python application development with automated unit testing.
  • Experience using authorized AI capabilities to improve reliability workflows, with validation and data-sensitivity awareness.
  • Ability to establish safe AI usage practices while maintaining resiliency, security, and auditability.
  • Experience developing with Terraform and understanding Terraform Enterprise.
  • Experience delivering system design, application development, testing, and operational stability.
  • Knowledge of distributed compute frameworks such as Spark, Glue, or MapReduce.
  • Excellent troubleshooting, analytical, and communication skills.

Preferred Qualifications and Skills

  • Experience with data pipelines using Spark.
  • Exposure to AWS and Databricks platform administration.
  • Knowledge of containerization with Docker or Kubernetes and orchestration.
  • Familiarity with distributed systems and large-scale data processing.

About the Company

The employer describes a collaborative environment focused on career growth, innovation, and impactful technology solutions.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at JPMorganChase