Base Career helps you apply smarter for this job.
Key skills for this role
Lead reliability engineering initiatives across IAM platforms and services.
Drive platform availability, resiliency, service health, and operational readiness improvements.
Establish and implement SRE best practices and operational excellence standards.
Partner with engineering teams to embed reliability, monitoring, and automation requirements throughout the development lifecycle.
Perform production readiness reviews and operational risk assessments.
Design and implement monitoring and observability strategies for IAM platforms.
Develop standards for: Infrastructure Monitoring Application Monitoring User Experience Monitoring Transaction Monitoring Dependency Monitoring Security Event Monitoring Cloud Service Monitoring
Infrastructure Monitoring
Application Monitoring
User Experience Monitoring
Transaction Monitoring
Dependency Monitoring
Security Event Monitoring
Cloud Service Monitoring
Establish logging, metrics, tracing, telemetry, dashboards, and alerting standards.
Build and manage centralized observability solutions leveraging platforms such as Splunk, Dynatrace, Datadog, Grafana, Azure Monitor, and Application Insights.
Develop service health dashboards and executive operational reporting.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Jersey City, USA
Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefro
Jersey City, USA
Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefro
Boston, USA
Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefro
Dallas, USA
Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefro
Tampa, USA
Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefro
Tampa, USA
Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefro
Tampa, USA
Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefro
London, GBR
Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefro
Jersey City, USA
Jersey City, USA
Boston, USA
Dallas, USA
Tampa, USA
Tampa, USA
Tampa, USA
London, GBR
Create operational architecture diagrams, service dependency maps, data flow diagrams, and resiliency models.
Review platform designs and identify scalability, performance, availability, and resiliency risks.
Define and implement standards for: High Availability (HA) Disaster Recovery (DR) Failover Design Capacity Planning Fault Tolerance Service Recovery
High Availability (HA)
Disaster Recovery (DR)
Failover Design
Capacity Planning
Fault Tolerance
Service Recovery
Collaborate with architecture and engineering teams to eliminate single points of failure.
Develop automation frameworks to improve operational efficiency and platform reliability.
Create scripts and tooling using PowerShell, Python, Bash, or similar technologies for monitoring, remediation, health validation, and operational support.
Design and implement automated testing frameworks for IAM platforms, APIs, integrations, and provisioning workflows.
Develop automation for: Platform Health Checks Service Validation Monitoring Configuration Incident Triage Alert Enrichment Automated Recovery Tasks Operational Reporting
Platform Health Checks
Service Validation
Monitoring Configuration
Incident Triage
Alert Enrichment
Automated Recovery Tasks
Operational Reporting
Drive Infrastructure as Code (IaC) and configuration automation practices.
Partner with engineering teams to integrate automated testing into CI/CD pipelines.
Define and monitor SLIs, SLOs, Error Budgets, and Availability Targets.
Drive initiatives to improve uptime, performance, reliability, and recoverability.
Lead DR testing, failover exercises, resiliency validation, and business continuity readiness activities.
Identify operational bottlenecks and implement preventive controls.
Lead major incident response, technical troubleshooting, and root cause analysis.
Drive post-incident reviews and corrective action programs.
Analyze incident trends and recurring service issues.
Implement proactive monitoring and predictive alerting solutions.
Improve Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
Global post-trade market infrastructure for the financial services industry.
Visit company websiteJobs and hiring trendsFull-time
Senior · 10+ years experience
Hybrid
Apply faster on company sites with our extension.