{bc}
linkedin

Senior Site Reliability Engineer

MontyCloud
Bengaluru, IND
Full-time
Mid-Senior
Onsite
Discovered 1 weeks ago
AWSPython scriptingAnsible, Puppet, or ChefSplunk, New Relic, Datadog, AWS CloudWatch, or AWS X-RayDisaster recoveryIncident management
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

AWSPython scriptingAnsible, Puppet, or Chef
Smart Apply

Full Job Posting

Role Overview

Improve the reliability, scalability, and performance of MontyCloud’s cloud management SaaS platform.

Develop and maintain tools and systems that automate and optimize cloud infrastructure and application management.

Guide operational practices toward reliability engineering excellence and continuous improvement.

Key Responsibilities

  • Design and implement automation for cloud infrastructure and application management and monitoring.
  • Collaborate with SREs and cross-functional teams to proactively address reliability issues.
  • Monitor infrastructure and application health and performance and implement troubleshooting and optimization strategies.
  • Participate in on-call rotations and provide expert incident response and resolution.
  • Champion disaster recovery, chaos engineering, and incident response practices.
  • Lead post-mortem analyses and integrate lessons learned into future operations.

Must-Have Skills

  • Problem-solving skills.
  • Experience with AWS cloud infrastructure.
  • Python scripting experience.
  • Experience with Ansible, Puppet, or Chef for automation and configuration management.
  • Experience with Splunk, New Relic, Datadog, AWS CloudWatch, or AWS X-Ray.
  • Experience with disaster recovery and incident management.

Experience

  • At least 5 years of SRE experience in a SaaS platform environment.
  • At least 3 years managing and optimizing SaaS platforms.
  • At least 3 years of hands-on AWS experience.
  • At least 4 years using automation tools such as Ansible, Puppet, or Chef.
  • At least 4 years scripting in Python or similar languages.
  • At least 3 years using monitoring and observability tools.
  • At least 3 years leading disaster recovery and implementing chaos engineering practices.
  • At least 4 years participating in on-call rotations and incident management.
  • At least 4 years of end-to-end application development experience.
  • At least 3 years leading post-mortem analysis sessions.

Education

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, or equivalent hands-on experience building and operating large-scale cloud or SaaS platforms.

Good-to-Have Skills

  • Application development experience in any technology stack is desirable.
  • CI/CD experience with tools such as Jenkins or GitLab CI is described as a plus.
  • Experience with chaos engineering tools such as Gremlin or Chaos Monkey is described as a plus.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today