Base Career helps you apply smarter for this job.
Key skills for this role
Monitor and manage batch execution: Oversee the execution of critical daily, weekly, and monthly batch processing cycles scheduled via Autosys.
Troubleshoot batch failures: Rapidly diagnose and resolve Autosys job failures, analyzing log files, identifying dependency issues, and performing necessary job overrides, force-starts, or hold/release actions to minimize business impact.
Optimize job flows: Collaborate with development and engineering teams to define, configure, and optimize Autosys job definitions using JIL (Job Information Language).
Identify and eliminate manual bottlenecks: Actively analyze daily support activities to identify repetitive, manual tasks ("toil") and design automated solutions to eliminate them.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Chennai, IND
Mumbai, IND
Mumbai, IND
Gurugram, IND
Pune, IND
Gurugram, IND
Pune, IND
Mumbai, IND
Develop automation scripts: Write, test, and deploy robust scripts (using Python, Bash, or PowerShell) to automate routine operations, such as daily health checks, application restarts, log archiving, and data reconciliation.
Drive process improvements: Evaluate existing support workflows, runbooks, and escalation paths, implementing enhancements to streamline operations and reduce Mean Time to Repair (MTTR).
Own and drive the end-to-end resolution of L2/L3 production incidents, ensuring strict adherence to corporate Service Level Agreements (SLAs) and Service Level Objectives (SLOs).
Lead technical triage during Major Incidents (MIM) and high-severity outages. Coordinate effectively with cross-functional global teams (Infrastructure, Database, Networks, Development, and Business Operations) to restore services rapidly.
Act as the primary technical escalation point during incidents, translating complex technical issues into clear, concise, and business-friendly updates for senior leadership and stakeholders.
Ensure accurate and timely logging, categorization, and tracking of incidents within ServiceNow.
Lead proactive Problem Management initiatives by analyzing incident trends, identifying systemic patterns, and pinpointing recurring failure points.
Conduct deep-dive technical investigations—including log analysis, database queries, and infrastructure health checks—to perform comprehensive Root Cause Analysis (RCA).
Author high-quality Post-Incident Reviews (PIRs) and RCA documents, detailing the timeline, root cause, impact, and preventative actions.
Collaborate closely with Development and Engineering teams to prioritize, track, and implement permanent bug fixes, structural workarounds, and long-term remediations.
Troubleshoot database-related application issues by writing and executing complex SQL queries (including multi-table joins, subqueries, and aggregations) on databases such as Oracle, MS SQL Server, or Sybase.
Analyze database performance, identify slow-running queries, and collaborate with DBAs to resolve locks, blocks, and indexing issues affecting production.
Actively monitor application health and infrastructure performance using ITRS Geneos (Active Console).
Configure, customize, and maintain ITRS Geneos samplers, rules, alerts, and Netprobes to ensure comprehensive coverage of critical system components.
Working experience supporting Big Data ecosystems and Enterprise Application/Analytics Platforms (EAP). Hands-on familiarity with Hadoop, HDFS, Hive, Spark, YARN, and Kafka. Ability to troubleshoot distributed job failures and monitor cluster health.
Strong hands-on experience with Autosys (or similar enterprise job schedulers). Proficient in monitoring batch cycles, troubleshooting job failures, managing dependencies, performing run-time overrides, and writing/modifying JIL (Job Information Language) configurations.
Advanced hands-on experience navigating Unix/Linux file systems, managing processes, analyzing system performance (CPU, memory, disk I/O), and writing/debugging Shell scripts (Bash/Ksh).
Proficient in writing complex SQL queries to extract, analyze, and troubleshoot data issues across relational databases (Oracle, Sybase, or MS SQL Server). Understanding of database locks, indexing, and basic performance tuning.
Hands-on experience using ITRS Geneos for real-time monitoring. Ability to navigate the Active Console, interpret alerts, and configure basic samplers, rules, and alerts.
Proven experience in automating manual support tasks and optimizing operational processes. Strong ability to identify automation opportunities and implement scripting solutions to reduce operational toil.
Strong expertise in ITIL processes, specifically Incident and Problem Management. Proven track record of leading major incident triage, managing SLAs, conducting Root Cause Analysis (RCA), and managing the ticket lifecycle in ServiceNow.
------------------------------------------------------
Global financial services organization enabling growth and economic progress.
Visit company websiteJobs and hiring trendsFull-time
Senior · 5+ years experience
Hybrid
Apply faster on company sites with our extension.