{bc}
company_site

Senior Cloud Hardware Storage Engineer

Microsoft
Aliso Viejo, USA
Full-time
Senior · 5+ years experience
Discovered Today
azurecpluspluscsharppythonrustsql
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

azurecpluspluscsharp
Smart Apply

Full Job Posting

Build the platform behind Azure's fault self-healing and failure-prediction system with telemetry pipelines, prediction services, decision logic, and automated repair workflows. Ship AI agents to production: prompt and tool design, retrieval over diagnostics data, evaluation harnesses, guardrails, and the CI/CD path that deploys new agent skills safely. Close the loop safely: automated remediation with staged rollout, blast-radius limits, and verification. Own the developer experience: SDKs, APIs, and dashboards that let engineers across the org author and deploy new detection and repair logic themselves. Build and monitor measurements: prediction precision/recall, false-repair rate, action success rate, and regression gates that block a bad model from shipping. 8+ years building and operating production software, distributed services, data platforms, or large-scale automation. Python and C# (or C++/Rust), plus cloud-scale data pipelines at high volume. AI/ML in production, not just experimentation: agent or model serving, evaluation, versioning and rollback, drift and regression monitoring. LLM application patterns: agent/tool-calling, RAG, structured output and judgment about when an LLM is the wrong answer. A track record of automation that takes real actions on real infrastructure, with the safety engineering that requires. Bonus: anomaly detection on time-series data; Kusto/SQL; server hardware, firmware, or datacenter operations. Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 3+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 5+ years technical engineering experience OR equivalent experience These requirements include but are not limited to the following specialized security screenings: * Bachelor's Degree in Computer Engineering, Computer Science, Electrical Engineering, or related field AND 8+ years of firmware or embedded systems engineering experience OR Master's Degree in Computer Engineering, Computer Science, Electrical Engineering, or related field AND 6+ years of firmware or embedded systems engineering experience OR equivalent experience 6+ years developing SSD or storage device firmware, including 4+ years working directly with NVMe and PCIe protocols Demonstrated depth in storage device resiliency and fault analysis — failure mode characterization, error handling and recovery paths, and root-cause investigation of field failures Experience supporting live-site operations for storage at fleet scale, including on-call ownership and production incident resolution Track record of owning end-to-end technical design across the full reliability lifecycle: detection, prediction, mitigation, and repair Proven experience building automation-heavy systems that operate safely at hyperscale, with the guardrails, staged rollout, and blast-radius controls that requires

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Microsoft