Base Career helps you apply smarter for this job.
Key skills for this role
Design, implement, test, and operate components of monitoring, alerting, telemetry, diagnostics, and operational intelligence platforms for large-scale Azure services.
Develop automation that improves incident detection, enrichment, triage, routing, mitigation, recovery, and operational communications.
Build dashboards and analytics that provide actionable visibility into service health, reliability, and customer impact.
Contribute to scalable logging, metrics, tracing, diagnostics, and audit infrastructure.
Develop deployment pipeline and release engineering capabilities, including automated validation, deployment health monitoring, operational readiness checks, safe rollout automation, and release verification.
Investigate production issues using service telemetry and diagnostics, and partner with engineers to implement durable fixes.
Participate in DevOps and livesite operations, contributing to service quality, availability, performance, and operational excellence.
Use AI-assisted engineering and automation tools to improve development, testing, incident investigation, and operational efficiency.
Collaborate with engineering, operations, and support teams to improve service supportability and the customer experience.
Follow engineering best practices for security, privacy, quality, reliability, accessibility, and inclusive product development.
Bachelor's Degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience. 3+ years of software engineering experience.
with one or more modern programming languages such as C#, C++, Java, or Python.
building, testing, shipping, and supporting production software, cloud services, distributed systems, infrastructure platforms, or operational tooling.
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Redmond, USA
Warrenton, USA
Redmond, USA
London, GBR
, USA
Hillsboro, USA
London, GBR
Bengaluru, IND
Redmond, USA
with telemetry, monitoring, alerting, logging, metrics, tracing, diagnostics, automation, or incident-management systems.
Understanding of distributed systems fundamentals, cloud platforms, or service operations.
Strong problem-solving, debugging, troubleshooting, analytical, and communication skills.
Ability to independently own scoped features from design through implementation, validation, deployment, and production support.
Effectiveness at collaborating with diverse groups of people, openness to feedback, bias for action, and comfort working through ambiguity.
with operational intelligence, site reliability engineering (SRE), production operations, or service engineering.
building deployment pipelines, release orchestration systems, rollout automation, deployment health monitoring, or service validation frameworks.
with service APIs, data-processing systems, event-driven architectures, or reusable infrastructure components.
applying AI-assisted engineering tools, agents, or intelligent automation to coding, testing, incident investigation, or operational workflows.
with search technologies, information retrieval, vector databases, retrieval-augmented generation (RAG), or AI-powered applications.
Ability to reason about distributed services and troubleshoot across application, compute, network, storage, caching, messaging, and load-balancing layers.
engaging with customers or support teams to understand scenarios and resolve service issues.
Commitment to security, privacy, accessibility, responsible AI, and customer-focused engineering.
Microsoft is a global technology company that develops software, hardware, and cloud services, known for products like Windows, Office, Azure, and Xbox.
Visit company websiteJobs and hiring trendsFull-time
Mid · 3+ years experience
Apply faster on company sites with our extension.