{bc}
workday

Site Reliability Engineer Shift Manager

NCR Voyix
Chennai, IND
Full-time
Entry · 2+ years experience
Hybrid
Discovered 1 weeks ago
AppDynamicsDynatraceDatadogSplunkMicrosoft AzureGoogle Cloud Platform (GCP)
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

AppDynamicsDynatraceDatadog
Smart Apply

Full Job Posting

Position Summary

The Global Command Center (GCC) Site Reliability Engineer (I) is responsible for ensuring the availability, reliability, performance, and operational stability of business-critical applications, infrastructure, and customer-facing services. This role serves as a key member of the Global Command Center, providing proactive monitoring, incident response, problem management, automation, and operational excellence across enterprise environments.

The GCC works closely with application teams, infrastructure teams, cloud engineering, network operations, vendors, and business stakeholders to prevent outages, reduce operational risk, and drive continuous service improvement. The role participates in major incident management, root cause analysis, operational readiness reviews, and reliability engineering initiatives.

Reliability & Service Availability

Monitor enterprise applications, infrastructure, cloud platforms, and customer-facing services.

Ensure platform availability, performance, and service health within established SLAs and SLOs.

Identify service degradation trends and proactively address reliability risks.

Drive operational improvements to reduce incidents and improve system resiliency.

Establish and track reliability metrics, KPIs, and service health indicators.

Develop preventative measures to reduce future service disruptions.

Observability & Monitoring

Configure and maintain monitoring platforms including: AppDynamics Dynatrace Datadog Splunk Azure Monitor Google Cloud Operations ServiceNow Event Management

AppDynamics

Dynatrace

Datadog

Splunk

Azure Monitor

Google Cloud Operations

ServiceNow Event Management

Develop dashboards, alerts, and health monitoring solutions.

Reduce alert fatigue through alert tuning and optimization.

Automation & Engineering

Automate operational processes and repetitive tasks.

Develop scripts and tooling using: PowerShell Python Bash APIs

PowerShell

Python

Bash

APIs

Improve operational efficiency through self-healing and automated remediation capabilities.

Support infrastructure-as-code and reliability engineering initiatives.

Operational Readiness

Participate in Operational Readiness Reviews (ORR).

Validate monitoring, alerting, runbooks, and support procedures prior to production go-live.

Ensure escalation paths and support models are documented and operational.

Review application deployments for supportability and operational risks.

Cloud & Infrastructure Support

Support hybrid environments across: Azure Google Cloud Platform (GCP) AWS VMware

Azure

Google Cloud Platform (GCP)

AWS

VMware

Analyze application, network, database, and infrastructure performance issues.

Work with engineering teams to optimize platform stability and scalability.

Governance & Reporting

Produce incident reports, service health updates, operational reviews, and executive summaries.

Maintain operational documentation, runbooks, and knowledge articles.

Track service performance metrics and reliability improvements.

Support audit and compliance initiatives as required.

Required Qualifications

  • 2+ years of experience in: Site Reliability Engineering Production Support Systems Engineering DevOps Network Operations Center (NOC) Command Center Operations
  • Site Reliability Engineering
  • Production Support
  • Systems Engineering
  • DevOps
  • Network Operations Center (NOC)
  • Command Center Operations
  • Experience supporting mission-critical production environments.
  • Strong understanding of: Incident Management Problem Management Change Management Service Level Management
  • Incident Management
  • Problem Management
  • Change Management
  • Service Level Management
  • Experience with ServiceNow or similar ITSM platforms.

Operating Systems

Windows Server

Linux/Unix

Cloud Platforms

Microsoft Azure

Amazon Web Services (AWS)

Monitoring & Observability

New Relic

Preferred Qualifications

  • Experience working in a Global Command Center environment.
  • AWS, Azure, or GCP certifications.
  • Experience supporting retail, hospitality, payments, or enterprise SaaS platforms.
  • Experience with CI/CD pipelines and DevOps practices.
  • Knowledge of SRE concepts including: SLI/SLO/SLA management Error budgets Chaos testing Resiliency engineering
  • SLI/SLO/SLA management
  • Error budgets
  • Chaos testing
  • Resiliency engineering

Key Competencies

Critical Incident Leadership

Technical Troubleshooting

Problem Solving

Customer Focus

Operational Excellence

Communication Skills

Executive Presence

Collaboration

Continuous Improvement

Decision Making Under Pressure

Success Metrics

The GCC Site Reliability Engineer will be measured on:

Service availability and uptime

Incident response times

Mean Time to Detect (MTTD)

Mean Time to Restore (MTTR)

Reduction in recurring incidents

Monitoring effectiveness

Automation adoption

Operational readiness compliance

Customer impact reduction

Service reliability improvements

Work Environment

24x7 operational support organization.

Participation in on-call and major incident rotations.

Collaboration with global teams across multiple regions.

Hybrid cloud and enterprise production environments.

Fast-paced, mission-critical operational setting.

Offers of employment are conditional upon passage of screening criteria applicable to the job

EEO Statement

Integrated into our shared values is NCR Voyix’s commitment to equal employment opportunity. All qualified applicants will receive consideration for employment without regard to sex, age, race, color, creed, religion, national origin, disability, sexual orientation, gender identity, veteran status, military service, genetic information, or any other characteristic or conduct protected by law. NCR Voyix is committed to being a globally inclusive company where all people are treated fairly, recognized for their individuality, promoted based on performance and encouraged to strive to reach their full potential. We believe in understanding and respecting differences among all people. Every individual at NCR Voyix has an ongoing responsibility to respect and support a globally diverse environment.

Statement to Third Party Agencies To ALL recruitment agencies: NCR Voyix only accepts resumes from agencies on the preferred supplier list. Please do not forward resumes to our applicant tracking system, NCR Voyix employees, or any NCR Voyix facility. NCR Voyix is not responsible for any fees or charges associated with unsolicited resumes

“When applying for a job, please make sure to only open emails that you will receive during your application process that come from a @ncrvoyix.com email domain.”

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at NCR Voyix