Senior Site Reliability Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
The SRE practice keeps Encore, LabelTraxx, and supporting platforms reliable, observable, and secure across a multi-account AWS estate.
The Senior Platform Engineer on the SRE track owns reliability for major production domains end to end.
The role defines SLOs, builds observability and automation, leads complex incident response, shapes reliability standards, mentors the SRE team, and partners with cloud engineers.
Key Skills for This Role
Full Job Posting
About Amtech
Amtech provides enterprise software for packaging, printing, and manufacturing industries.
Its integrated systems support order management, production planning, scheduling, inventory, and business analytics.
Amtech is scaling its platform engineering organization as products move to a fully AWS-hosted, multi-tenant SaaS model.
Role Description
The SRE practice keeps Encore, LabelTraxx, and supporting platforms reliable, observable, and secure across a multi-account AWS estate.
The Senior Platform Engineer on the SRE track owns reliability for major production domains end to end.
The role defines SLOs, builds observability and automation, leads complex incident response, shapes reliability standards, mentors the SRE team, and partners with cloud engineers.
Reliability and Performance
- Own SLIs, SLOs, and error budgets for major production domains.
- Design standardized monitoring, alerting, and reliability frameworks using OpenTelemetry, OpenObserve, CloudWatch, and PagerDuty.
- Engineer capacity management, autoscaling, and self-healing.
- Lead root cause analysis for the highest-severity incidents and verify permanent fixes.
Automation and Platform Engineering
- Design reusable Terraform modules, deployment patterns, and progressive delivery in GitHub Actions.
- Operate and optimize ECS Fargate, EKS, Lambda, and RDS PostgreSQL workloads at production scale.
- Quantify and eliminate operational toil through engineering.
- Build reliability into Encore-on-AWS and LabelTraxx as customer counts grow.
Incident Response and On-Call
- Serve as senior incident commander for cross-service, customer-impacting incidents.
- Own escalation policies, alert quality, and rotation sustainability for the on-call program.
- Drive game days and failure testing to validate runbooks and recovery paths.
Security and Compliance
- Engineer IAM boundaries, secrets, and network controls into reliability tooling.
- Ensure operations satisfy SOC 2 and ISO 27001 obligations with automated audit evidence.
AI Competency
- The source begins an AI competency section describing the use of AI in operations and incident response with risk tiering, but the provided description is truncated.
Requirements
- Senior-level ownership of reliability for major production domains is expected.
- Ability to define SLOs, build observability and automation, and lead complex incident response.
- Ability to operate and optimize AWS workloads at production scale.
- Ability to engineer security controls and automated audit evidence for SOC 2 and ISO 27001 obligations.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career