Senior Lead Site Reliability Engineer - Cloud Data & Databricks
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
The Senior Lead Site Reliability Engineer will develop and deliver technology products for data and analytics.
The role focuses on cloud data platforms, Data Lake tools, operational excellence, data integrity, security, and AI-enabled reliability practices.
Key Skills for This Role
Full Job Posting
Role Overview
The Senior Lead Site Reliability Engineer will develop and deliver technology products for data and analytics.
The role focuses on cloud data platforms, Data Lake tools, operational excellence, data integrity, security, and AI-enabled reliability practices.
Job Responsibilities
- Lead SREs supporting a managed AWS Databricks platform and application engineering teams.
- Design resilient multi-AZ, multi-region, and multi-cloud strategies for business-critical services.
- Perform platform setup, configuration, workspace administration, resource monitoring, observability, alerting, and capacity planning.
- Evaluate architectural designs and technical credentials with vendors, startups, and internal teams.
- Optimize infrastructure and deployment processes through automation and collaboration with engineering and data teams.
- Develop secure production code, review code, troubleshoot systems, and apply SRE practices to improve reliability, scalability, and performance.
- Use authorized AI capabilities for incident analysis, reliability design, testing, production readiness, and requirements traceability.
- Maintain incident response procedures, root cause analyses, postmortems, auditability, and security controls.
Required Qualifications
- Formal training or certification in software engineering concepts and 10+ years of applied experience are required.
- Strong knowledge of SRE principles, SLIs, SLOs, error budgets, and incident management is required.
- Experience with monitoring tools, automation frameworks, and CI/CD pipelines is required.
- Proficiency in Python application development with automated unit testing is required.
- Experience with authorized AI capabilities and safe AI usage practices is required.
- Experience with Terraform development and Terraform Enterprise is required.
- Experience in system design, application development, testing, and operational stability is required.
- Knowledge of distributed compute frameworks such as Spark, Glue, or MapReduce is required.
- Excellent troubleshooting, analytical, and communication skills are required.
Preferred Qualifications
- Experience building data pipelines with Spark is preferred.
- Exposure to AWS and Databricks platform administration is preferred.
- Knowledge of Docker, Kubernetes, distributed systems, and large-scale data processing is preferred.
Work Environment
- The description states that the role works with cross-functional teams in an agile environment.
- No explicit remote, hybrid, onsite, or field workplace policy is provided.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at JPMorganChase
Control Manager Program - Vice President - Finance Control Management
, IND
Java Lead Software Engineer
Glasgow, GBR
Control Manager Program - Vice President - Finance Control Management
, IND
Payment Lifecycle Manager - Vice President
, IND
Trade Lifecycle Manager - Vice President
, IND
Asset and Wealth Management Audit Manager (Vice President)
, IND
Associate - Business Manager
, IND
Senior Product Delivery Associate
, IND
Lead Software Engineer - Lead Data Architect
, IND