Base Career helps you apply smarter for this job.
Key skills for this role
• Manage the Tier 1 team: directly manage a team of 14 AIOps Support Engineers performing manual triage of alarms and alerts across a diverse, 50-to-320-application portfolio including hiring, coaching, scheduling, and performance management.
• Own triage quality and speed: set and monitor standards for how quickly and accurately the team detects, classifies, and routes incidents, and drive continuous improvement in mean-time-to-triage.
• Drive data stewardship: partner with application teams to standardize alarm and alert data across heterogeneous log aggregation tools (Dynatrace, New Relic, ManageEngine, Glass box, and others) into a clean, consistent telemetry backbone built on Open Telemetry.
• Manage the reactive-to-proactive shift: reduce reliance on reactive, manual triage over time by improving alert quality, correlation, and early-warning signals laying the groundwork for future automated and Agentic triage.
• Navigate a diverse, moving application landscape: support applications spanning different technology stacks and different architecture dispositions (Invest, Tolerate, Retire, Migrate), reprioritizing team focus as the portfolio shifts.
• Coordinate onboarding of new apps: run a repeatable process for bringing new applications into Tier 1 coverage as the program scales from 50 to 320 applications, including support group and application owner mapping.
• Manage stakeholders: act as the primary point of contact for support groups, application owners, and AIOps program leadership on Tier 1 status, incidents, and data-quality issues.
• Manage shift/roster coverage: ensure the team of 14 provides consistent triage coverage across required hours as the application count grows.
• Report on outcomes: track and report team KPIs triage time, alert-to-incident accuracy, false-positive rates, coverage growth to program leadership.
What We're Looking For
• Experience: 7+ years in application/production support (L1/L1.5/L2) or site reliability, with 2+ years directly managing a technical support team.
• Hybrid environment expertise: proven experience supporting applications across both on-premises and cloud environments, with exposure to modern microservices architectures.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Bengaluru, IND
, IND
Bengaluru, IND
Bengaluru, IND
, IND
, IND
Bengaluru, IND
• Observability tooling: hands-on experience with monitoring and observability platforms such as Dynatrace, New Relic, AWS CloudWatch, ManageEngine, or Glass box; working knowledge of Open Telemetry and distributed tracing concepts.
• ITIL discipline: strong grounding in incident, problem, and change management practices, with ServiceNow or Jira ticket management experience.
• Technical range: comfortable with Linux and Windows troubleshooting, basic networking (TCP/IP, DNS, HTTP/HTTPS, SSL, load balancers), SQL/database query analysis, and API/integration troubleshooting.
• People management: demonstrated ability to hire, coach, and retain a team of 10+ technical support staff through a period of significant scale-up (5x application coverage growth).
• Analytical mindset: able to turn noisy, inconsistent alert data into clear, actionable insight, and to build repeatable frameworks rather than one-off fixes.
• Comfort with ambiguity: willing to support a moving target a portfolio spanning Invest, Tolerate, Retire, and Migrate applications and to adapt priorities as the program evolves.
• Incident Management Lifecycle, and working knowledge of Problem, Change Request, and Service Request concepts (ITIL)
• CMDB concepts and their use in incident and asset traceability
• Hands-on experience with log aggregation technologies (Dynatrace, New Relic, ManageEngine, Glass box, or similar)
• Working knowledge of JSON and XML, and basic file/task automation
• Understanding of IT infrastructure and basic networking: VMs, firewalls, load balancers, containers, OpenShift (OCP), Kubernetes
• Unix Shell scripting; Windows batch file creation
• Basic cloud concepts: compute, storage, and security fundamentals
• Security fundamentals: TLS, SSL, tokens, and secret management
• Familiarity with API gateways and API testing toolkits (Postman, SOAP UI, or similar)
• Outage response management and experience leading cross-functional coordination during major incidents
• Ability to drive Root Cause Analyses (RCAs) and build reusable knowledge articles/runbooks
• SLA/SLO management and reporting, including availability calculation
• Working knowledge of data concepts: data latency, data fragmentation, data lineage, and data marts
• Familiarity with AI concepts such as prompt engineering, knowledge graphs, and Retrieval-Augmented Generation (RAG)
• Experience with cloud-native observability on AWS, Azure, or GCP
• Exposure to Agentic AI or automation-driven triage tooling
Success Looks Like
• A 14-person Tier 1 team that reliably triages alerts across 50+ applications with clear, standardized data.
• A measurable, ongoing reduction in reactive manual triage as proactive detection improves.
• A clean, well-governed Open Telemetry-based data backbone that the future Agentic automation layer can build on.
A repeatable onboarding process ready to scale coverage from 50 to 320 applications.
Indian private technology centre and Bell Canada capability hub building telecommunications software for consumer and enterprise customers.
Visit company websiteJobs and hiring trendsFull-time
Senior · 7+ years experience
Hybrid
Apply faster on company sites with our extension.