Base Career helps you apply smarter for this job.
Key skills for this role
We’re hiring a Senior Platform Reliability Engineer to help define and scale reliability as a first-class capability at Grow. In this role you’ll operate horizontally across the organization, shaping how reliability is understood, measured, and built into the developer experience.
You’ll work closely with other members of the platform team as well as our product engineering teams to establish standards around observability, SLOs/SLAs, and incident response—while also helping translate those standards into self-service tooling and “golden paths” that make it easy for teams to adopt them.
This is a high-impact, highly autonomous role where you’ll drive both cultural and technical change, ultimately enabling teams to independently build and operate reliable systems at scale.
We’re hiring a Senior Platform Reliability Engineer to help define and scale reliability as a first-class capability at Grow. In this role you’ll operate horizontally across the organization, shaping how reliability is understood, measured, and built into the developer experience.
You’ll work closely with other members of the platform team as well as our product engineering teams to establish standards around observability, SLOs/SLAs, and incident response—while also helping translate those standards into self-service tooling and “golden paths” that make it easy for teams to adopt them.
This is a high-impact, highly autonomous role where you’ll drive both cultural and technical change, ultimately enabling teams to independently build and operate reliable systems at scale.
You’ll help us establish and scale reliability as a discipline at Grow by:
Defining Reliability Standards Establishing frameworks for SLOs/SLAs, error budgets, and operational readiness; helping teams understand what to measure and why it matters.
Improving Observability & Measurement Identifying gaps in metrics, logging, and tracing; ensuring services are measurable, debuggable, and aligned with reliability goals.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
, USA
About Us: Grow was born to tackle a critical challenge: making mental health care more effective and accessible for everyone. Since launching in 2020, 15 million sessions have started on Grow, and we've barely scratched
San Francisco, USA
, USA
New York City, USA
New York City, USA
New York City, USA
New York City, USA
San Francisco, USA
New York City, USA
, USA
Evolving Incident Response Developing and improving incident response practices, from detection to post-incident learning, and helping teams build sustainable on-call and escalation patterns.
Enabling Self-Service Reliability Partnering with the platform team to build tooling and abstractions (e.g., service scorecards, dashboards, templates, golden paths) that make it easy for teams to adopt and stay compliant with reliability standards.
Driving Adoption Across Teams Working cross-functionally to educate, influence, and guide engineering teams—scaling reliability practices through a combination of clear standards, strong communication, and developer-friendly systems
Experienced in production systems: You have 6+ years of experience operating and improving reliability of production systems at scale.
Strong foundation in cloud and infrastructure: You have hands-on experience with AWS, Kubernetes (e.g., EKS), and infrastructure as code tools like Terraform.
Deep understanding of reliability principles: You’ve defined or worked with SLOs/SLAs, understand error budgets, and have experience improving reliability through measurement and iteration.
Observability expertise: You’ve worked with modern observability tooling (we use DataDog) and understand how to build actionable monitoring systems across metrics, logs, and traces.
Systems thinker: You’re able to zoom out, identify patterns across teams and services, and design solutions that scale beyond a single system.
Impact-oriented: You focus on outcomes over output and care deeply about improving real reliability outcomes—not just adding processes.
Strong communicator and influencer: You can drive change across teams without direct authority, balancing pragmatism with long-term vision.
Self-directed: You thrive in ambiguous environments and are comfortable defining problems, proposing solutions, and executing independently.
Team player : You collaborate well, communicate with empathy, and enjoy mentoring and learning from others.
You’ve helped introduce or scale reliability practices in a growing organization.
You’ve built internal tooling or platforms used by multiple teams.
You have experience designing service-level scorecards or compliance/reporting systems.
You’ve worked with both SaaS (e.g., DataDog) and self-managed observability stacks.
You were previously a product engineer and bring empathy for developer experience.
You have experience with database reliability and performance (we use PostgreSQL)
This is a rare opportunity to define what reliability looks like at a growing, scaling engineering organization—and to do it in a way that actually sticks.
You won’t just be responding to incidents or working within a single team. You’ll be shaping how reliability is measured, enforced, and experienced across the entire company. You’ll work alongside your team mates to turn best practices into intuitive, self-service systems that engineers rely on every day.
Your work will directly improve system reliability, reduce incidents, and enable teams to move faster with confidence, ultimately making reliability a built-in property of how we build software at Grow.
Employment Type: Full Time, Exempt
Base Compensation: The base compensation range for this position is $182,000–$250,000 USD Annually.
This is a hybrid role with the expectation to work onsite from our San Francisco, NYC, or Seattle hub location three days per week (Tuesday, Wednesday, and Thursday) and travel 2–3 times per year (e.g., company and department offsites). The base compensation for this role will vary depending on several factors, including relevant experience, qualifications, and the candidate’s working location.
Mental health technology platform connecting patients, independent providers, and insurers with affordable, insurance-covered therapy and medication care.
Visit company websiteJobs and hiring trendsUSD 182000-250000 yearly / year
Full-time
Senior · 6+ years experience
Hybrid
Apply faster on company sites with our extension.