Base Career helps you apply smarter for this job.
Key skills for this role
Escalation Handling: Acknowledge and own incidents escalated by the L1 support team or monitoring alerts.
Troubleshooting & Diagnosis: Perform in-depth technical analysis and root cause investigation of application, database, middleware, and infrastructure failures.
Workaround Implementation: Apply approved temporary workarounds or permanent hotfixes to restore services quickly and minimize business impact.
Collaboration: Partner with L3 developers, database administrators (DBAs), network engineers, and system administrators to resolve complex, multi-tiered issues.
Incident Lifecycle Management: Track and document all incident progression within the ITSM platform (e.g., ServiceNow) from creation through to resolution.
System Monitoring: Actively monitor application health, batch jobs, integration feeds, and system performance using enterprise observability suites.
Proactive Interventions: Identify recurring patterns, error trends, or capacity bottlenecks and initiate preventative actions.
Health Checks: Perform daily health checks, start-of-day (SOD) and end-of-day (EOD) verifications, and batch processing runs.
Deployment Validation: Support the verification of application deployments, system patches, and infrastructure upgrades during maintenance windows.
Configuration Controls: Maintain and update application configuration files, environment variables, and parameter settings under strict change management processes.
Environment Management: Assist in maintaining non-production (UAT/Staging) and production environments to ensure consistency.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Mississauga, CAN
Sydney, AUS
London, GBR
London, GBR
Mississauga, CAN
Toronto, CAN
Mississauga, CAN
Sydney, AUS
London, GBR
London, GBR
London, GBR
London, GBR
Status Updates: Provide clear, timely, and precise communications to business stakeholders, product owners, and technology leadership during critical (Sev-1/Sev-2) incidents.
Bridges & War Rooms: Participate in or lead technical incident bridges to coordinate restoration efforts.
Post-Incident Reviews: Contribute technical inputs to Post-Incident Reviews (PIRs) and Root Cause Analysis (RCA) documentation.
Runbooks & SOPs: Document and update Standard Operating Procedures (SOPs), application architecture maps, and troubleshooting runbooks.
Knowledge Sharing: Train L1 teams on common issue resolution paths to shift-left workloads and improve first-contact resolution rates.
------------------------------------------------------
Global financial services organization enabling growth and economic progress.
Visit company websiteJobs and hiring trendsFull-time
Senior · 5+ years experience
Hybrid
Apply faster on company sites with our extension.