{bc}
indeed

Senior Site Reliability Engineer Lead

HCLTech
Karnataka, IND
Onsite
Discovered 2 weeks ago
Site Reliability EngineeringProduction operationsSplunk EnterpriseSplunk ITSISplunk Observability CloudSLI/SLO and error budgets
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site Reliability EngineeringProduction operationsSplunk Enterprise
Smart Apply

Full Job Posting

Role overview

The SRE Engineer operates business applications and platforms with a focus on reliability, availability, observability, and operational support.

The role uses Splunk-based observability, SLI/SLO practices, automation, and continuous service improvement.

Key responsibilities

  • Maintain production applications and platforms for high availability, reliability, performance, and stability.
  • Configure and optimize monitoring, logging, alerting, and service health solutions using Splunk Enterprise, Splunk ITSI, and Splunk Observability Cloud.
  • Develop dashboards, alerts, reports, service health views, and KPI monitoring solutions.
  • Perform incident management, troubleshooting, root cause analysis, blameless postmortems, and service restoration.
  • Implement SLI, SLO, error budget tracking, reliability reporting, and continuous improvement initiatives.
  • Reduce alert noise and false positives while improving monitoring coverage.
  • Support event correlation, anomaly detection, alert enrichment, and operational analytics.
  • Automate operational tasks, monitoring, reporting, and remediation activities.
  • Collaborate with application, cloud, infrastructure, and operations teams.
  • Maintain runbooks, standard operating procedures, knowledge articles, and operational documentation.
  • Participate in reliability reviews, Agile ceremonies, and service improvement activities.
  • Support releases, platform upgrades, maintenance, and operational readiness reviews.

Must-have skills

  • 6+ years of experience in production operations, SRE, observability, application support, or operations engineering.
  • Experience with Splunk ITSI, Splunk Observability Cloud, and Splunk Enterprise logging.
  • Experience with SLI, SLO, error budgets, service health monitoring, and reliability reporting.
  • Experience with incident, problem, and change management, troubleshooting, RCA, and blameless postmortems.
  • Experience supporting applications on virtual machine and container platforms.
  • Experience with Azure, AWS, GCP, or other cloud platforms.
  • Experience with Ansible, Python, RPA, or other automation technologies.
  • Experience with ITSM processes and tools such as ServiceNow.
  • Strong analytical, troubleshooting, communication, and collaboration skills.

Good-to-have skills

  • Knowledge of DevOps, CI/CD, GitOps, and release automation.
  • Experience with Java, .NET, SAP, Salesforce, SaaS/COTS, or other enterprise applications.
  • Experience with resilience testing and chaos engineering.
  • Exposure to OpenTelemetry, agentic AI, AI-driven operations, or AI-assisted observability.
  • Splunk, SRE, cloud, or observability-related certifications are advantageous.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at HCLTech