{bc}
linkedin

Site Reliability Engineer(Production Reliability, Azure Operations, Databricks Support and Incident Response)

K&K Global Talent Solutions INC.
Toronto, CAN
Full-time
Onsite
Discovered 1 weeks ago
Site reliability engineeringMicrosoft Azure cloud infrastructureDatabricksAzure Storage, ADLS Gen2, and Blob StorageAzure networking, VNets, NSGs, private endpoints, DNS, and routingEntra ID, RBAC, managed identities, and Key Vault
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Site reliability engineeringMicrosoft Azure cloud infrastructureDatabricks
Smart Apply

Full Job Posting

Role Overview

This is an Intermediate Site Reliability Engineer role supporting and continuously improving enterprise Azure and Databricks platforms.

The work focuses on production reliability, monitoring, incident response, availability, operational readiness, and platform support.

The engineer will collaborate with platform engineering, security, network, application, and data teams.

Location and Schedule

  • The role is based in Toronto, Ontario and is onsite three days a week.
  • The position is full-time and permanent.

Key Responsibilities

  • Monitor production Azure and Databricks environments for availability, performance, and operational readiness.
  • Respond to incidents and service requests and participate in on-call rotations.
  • Troubleshoot Azure, Databricks, networking, storage, identity, access, and application issues.
  • Support Databricks workspaces, compute, policies, jobs, workflows, access, monitoring, and cost controls.
  • Operate Unity Catalog and integrations with Azure data and security services.
  • Support Azure networking, connectivity, and storage services.
  • Manage alerts, dashboards, incident bridges, escalations, and emergency changes.
  • Perform root cause analysis and remediation for recurring incidents.
  • Assist with patching, upgrades, maintenance, validation, and disaster recovery exercises.
  • Maintain runbooks, documentation, and support procedures.
  • Manage work through JIRA and ServiceNow.
  • Improve reliability and operational support with engineering teams.

Must-Have Skills

  • At least 3 years of production Azure cloud infrastructure support experience.
  • At least 1 year of hands-on Databricks support experience.
  • Experience with Azure Storage, ADLS Gen2, Blob Storage, Azure networking, identity, and security.
  • Experience with production incidents, on-call support, escalation, change management, root cause analysis, and problem management.
  • At least 1 year of Windows Server administration experience.
  • At least 1 year of Linux administration experience.
  • Basic understanding of network troubleshooting and connectivity concepts.
  • Experience with JIRA, ServiceNow, runbooks, and enterprise support procedures.

Preferred Skills

  • Databricks jobs, clusters, workflows, workspaces, and user access support experience is preferred.
  • Unity Catalog concepts, permissions, and governance knowledge is preferred.
  • Azure SQL and Azure Data Factory operational support experience is preferred.
  • Business continuity and disaster recovery knowledge, including RTO and RPO, is preferred.
  • Experience with AI or GenAI platforms, enterprise data platforms, capacity review, cost monitoring, and performance troubleshooting is preferred.
  • Strong communication, knowledge-sharing, and documentation skills are preferred.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at K&K Global Talent Solutions INC.