The SRE team develops and maintains platforms and tools which help other Engineering teams in Goldman Sachs to build and operate reliable and resilient systems. These systems span on-premises datacenters and multiple public cloud environments. The platforms we offer include central logging, monitoring, agents and alerting and we provide tools to drive adoption and improvements to capacity planning, operational readiness assessments, production incident postmortems, SLIs / SLOs, and deployment automation including canary releases.
Strategic Reliability & Performance: Drive the strategic direction for availability, scalability, and performance of mission-critical applications and platform services, ensuring alignment with firm-wide objectives.
Architectural Leadership: Lead the design, build, and implementation of highly available, resilient, and scalable infrastructure and application architectures.
Advanced Automation & Tooling: Architect and develop sophisticated platforms, tools, and automation solutions to eliminate toil, optimize operational workflows, and enhance deployment processes across the enterprise.
Complex Incident Management & Post-Mortem Analysis: Lead critical incident response, conduct in-depth root cause analysis for systemic issues, and implement long-term preventative measures to significantly enhance system stability and resilience.
System Design & Capacity Planning: Partner with development teams to embed reliability into application design from inception, provide expert system design consulting, and lead comprehensive capacity planning initiatives for future growth.
Observability & Insights: Define and implement advanced monitoring, high volume logging with multi-user query capabilities, and tracing strategies to provide deep, actionable insights into application performance, infrastructure health, and user experience.
Technical Vision & Mentorship: Provide technical vision, lead complex technical projects, conduct rigorous code reviews, enforce SDLC best practices, and actively mentor and develop senior and staff-level engineers.
Technology Evaluation & Adoption: Stay at the forefront of industry trends and advancements, evaluating and integrating cutting-edge tools and frameworks to significantly improve operational efficiency and reliability.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
On-Call Leadership: Participate in and lead on-call rotations, providing expert guidance and hands-on support for critical system incidents.
Experience: Minimum of 6+ years of hands-on experience in Site Reliability Engineering, with a proven track record in architecting, designing, building, and maintaining highly available, scalable, and fault-tolerant systems at an enterprise level.
Technical Proficiency:
Exceptional programming skills in one or more major languages such as Java, Python, Go with a focus on building robust, scalable software.
Extensive hands-on experience with cloud platforms (e.g., AWS, GCP) and deep expertise in containerization and orchestration technologies (e.g., Docker, Kubernetes).
Mastery of Infrastructure as Code (IaC) tools (e.g., Terraform, CloudFormation) and configuration management tools (e.g., Puppet, Chef, Ansible).
Advanced proficiency in Prompt Engineering and Retrieval-Augmented Generation (RAG) architectures to automate complex SRE workflows, such as the generation of Infrastructure as Code (IaC), dynamic runbooks, and incident response summaries.
Profound understanding of Linux internals, networking, distributed systems, and advanced system performance tuning.
Expertise in designing and implementing comprehensive monitoring, alerting, logging and tracing solutions (e.g., Prometheus, Grafana, ELK stack, Datadog, PagerDuty).
Deep experience with CI/CD tools and practices (e.g., Jenkins, GitLab, Maven).
Strong foundation in databases and distributed systems.
Exceptional problem-solving abilities and analytical skills, with a track record of resolving complex technical challenges.
Preferred Experience:
Experience with Distributed Databases like Elastic Search
Experience with working on GCP Big Query
Experience with messaging Systems Like Kafka
Education: Advanced degree (Bachelor’s or Mas ter's or PhD) in Computer Science or a related technical field involving coding and/or systems engineering, or equivalent practical experience.
Soft Skills: Superior communication, collaboration, and interpersonal skills, with the ability to influence technical direction, lead cross-functional initiatives, and effectively engage with global teams and executive leadership. Proven ability to work independently, manage multiple complex stakeholders, and drive significant organizational change.
About Goldman Sachs
Financial Services10,001+Founded 1869
Goldman Sachs is a global investment banking, securities, and investment management firm. It provides a wide range of financial services including advisory, underwriting, lending, and asset management to corporations, governments, and individuals worldwide.