Senior Site Reliability Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Improve the reliability, scalability, and performance of MontyCloud’s cloud management SaaS platform.
Develop and maintain tools and systems that automate and optimize cloud infrastructure and application management.
Guide operational practices toward reliability engineering excellence and continuous improvement.
Key Skills for This Role
Full Job Posting
Role Overview
Improve the reliability, scalability, and performance of MontyCloud’s cloud management SaaS platform.
Develop and maintain tools and systems that automate and optimize cloud infrastructure and application management.
Guide operational practices toward reliability engineering excellence and continuous improvement.
Key Responsibilities
- Design and implement automation for cloud infrastructure and application management and monitoring.
- Collaborate with SREs and cross-functional teams to proactively address reliability issues.
- Monitor infrastructure and application health and performance and implement troubleshooting and optimization strategies.
- Participate in on-call rotations and provide expert incident response and resolution.
- Champion disaster recovery, chaos engineering, and incident response practices.
- Lead post-mortem analyses and integrate lessons learned into future operations.
Must-Have Skills
- Problem-solving skills.
- Experience with AWS cloud infrastructure.
- Python scripting experience.
- Experience with Ansible, Puppet, or Chef for automation and configuration management.
- Experience with Splunk, New Relic, Datadog, AWS CloudWatch, or AWS X-Ray.
- Experience with disaster recovery and incident management.
Experience
- At least 5 years of SRE experience in a SaaS platform environment.
- At least 3 years managing and optimizing SaaS platforms.
- At least 3 years of hands-on AWS experience.
- At least 4 years using automation tools such as Ansible, Puppet, or Chef.
- At least 4 years scripting in Python or similar languages.
- At least 3 years using monitoring and observability tools.
- At least 3 years leading disaster recovery and implementing chaos engineering practices.
- At least 4 years participating in on-call rotations and incident management.
- At least 4 years of end-to-end application development experience.
- At least 3 years leading post-mortem analysis sessions.
Education
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, or equivalent hands-on experience building and operating large-scale cloud or SaaS platforms.
Good-to-Have Skills
- Application development experience in any technology stack is desirable.
- CI/CD experience with tools such as Jenkins or GitLab CI is described as a plus.
- Experience with chaos engineering tools such as Gremlin or Chaos Monkey is described as a plus.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career