Senior Site Reliability Engineer Lead
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
The SRE Engineer operates business applications and platforms with a focus on reliability, availability, observability, and operational support.
The role uses Splunk-based observability, SLI/SLO practices, automation, and continuous service improvement.
Key Skills for This Role
Full Job Posting
Role overview
The SRE Engineer operates business applications and platforms with a focus on reliability, availability, observability, and operational support.
The role uses Splunk-based observability, SLI/SLO practices, automation, and continuous service improvement.
Key responsibilities
- Maintain production applications and platforms for high availability, reliability, performance, and stability.
- Configure and optimize monitoring, logging, alerting, and service health solutions using Splunk Enterprise, Splunk ITSI, and Splunk Observability Cloud.
- Develop dashboards, alerts, reports, service health views, and KPI monitoring solutions.
- Perform incident management, troubleshooting, root cause analysis, blameless postmortems, and service restoration.
- Implement SLI, SLO, error budget tracking, reliability reporting, and continuous improvement initiatives.
- Reduce alert noise and false positives while improving monitoring coverage.
- Support event correlation, anomaly detection, alert enrichment, and operational analytics.
- Automate operational tasks, monitoring, reporting, and remediation activities.
- Collaborate with application, cloud, infrastructure, and operations teams.
- Maintain runbooks, standard operating procedures, knowledge articles, and operational documentation.
- Participate in reliability reviews, Agile ceremonies, and service improvement activities.
- Support releases, platform upgrades, maintenance, and operational readiness reviews.
Must-have skills
- 6+ years of experience in production operations, SRE, observability, application support, or operations engineering.
- Experience with Splunk ITSI, Splunk Observability Cloud, and Splunk Enterprise logging.
- Experience with SLI, SLO, error budgets, service health monitoring, and reliability reporting.
- Experience with incident, problem, and change management, troubleshooting, RCA, and blameless postmortems.
- Experience supporting applications on virtual machine and container platforms.
- Experience with Azure, AWS, GCP, or other cloud platforms.
- Experience with Ansible, Python, RPA, or other automation technologies.
- Experience with ITSM processes and tools such as ServiceNow.
- Strong analytical, troubleshooting, communication, and collaboration skills.
Good-to-have skills
- Knowledge of DevOps, CI/CD, GitOps, and release automation.
- Experience with Java, .NET, SAP, Salesforce, SaaS/COTS, or other enterprise applications.
- Experience with resilience testing and chaos engineering.
- Exposure to OpenTelemetry, agentic AI, AI-driven operations, or AI-assisted observability.
- Splunk, SRE, cloud, or observability-related certifications are advantageous.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at HCLTech
Technical Architect (158792)
, IND
Automation Test Lead Embedded (152157)
, IND
SME - Program & Project Management (165346)
, IND
SME - IBM AIX, Power HA (166020)
, IND
Senior Technical Lead (164957)
, USA
Senior Project Lead - Scrum Master (165971)
, IND
Senior .NET Developer (165974)
, IND
DataStage Technical Lead (165576)
, IND