Base Career helps you apply smarter for this job.
Key skills for this role
General Summary:
As a Senior DevOps Engineer – Observability Platform , you will be responsible for building and maintaining scalable, reliable infrastructure and deployment pipelines with a strong emphasis on observability — metrics, logs, and traces — across systems running on Kubernetes and AWS. You will work closely with development teams to improve development velocity while ensuring system reliability, security, and performance. This role is critical in providing a standardized, observability platform that gives both internal engineering teams and external, customer-facing services deep, reliable visibility into system health, performance, and reliability
Minimum Qualifications:
Infrastructure Management : Design, implement, and maintain cloud-based infrastructure using Infrastructure as Code principles
Automation : Develop automation scripts and tools to streamline operations and eliminate manual processes
Containerization : Manage containerization strategies and orchestration using Docker and Kubernetes
Observability Platform: Design, build, and operate a standardized, self-service metrics, logs, and tracing platform (Prometheus, Grafana, Loki, OpenTelemetry) serving both internal teams and external, customer-facing services running on Kubernetes and AWS
Instrumentation & Telemetry: Partner with engineering teams to instrument applications and infrastructure, standardizing telemetry collection with OpenTelemetry
SLOs & Alerting: Define and maintain SLIs/SLOs and error budgets, build actionable dashboards, and tune alerting to maximize signal and reduce noise
Performance Optimization : Use observability data to analyze and optimize system performance, scalability, and cost-efficiency
Documentation : Create and maintain thorough documentation for infrastructure, deployment processes, and operational procedures
Incident Response & Escalation: Provide second-tier engineering escalation during business hours and own the telemetry, SLO, and alerting tooling that powers incident detection and reduces MTTD/MTTR; front-line 24/7 on-call is owned by the dedicated SRE team, not observability engineers. Lead post-mortem analysis for observability-platform incidents
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Bengaluru, IND
Hyderabad, IND
Bengaluru, IND
Bengaluru, IND
Hyderabad, IND
Bengaluru, IND
Hyderabad, IND
Hyderabad, IND
Verified company details for this employer are not available yet.
Full-time
Senior · 5+ years experience
Onsite
Apply faster on company sites with our extension.