Microsoft Logo

Microsoft

Senior Cloud Hardware Storage Engineer

Reposted 3 Days Ago
Be an Early Applicant
In-Office
Redmond, WA, USA
120K-261K Annually
Senior level
In-Office
Redmond, WA, USA
120K-261K Annually
Senior level
Build and operate Azure’s fault self-healing and failure-prediction platform at hyperscale. Responsibilities include telemetry pipelines, prediction services, AI agents, automated repair workflows, safe staged rollouts, developer SDKs and APIs, dashboards, and reliability metrics. The role requires production AI/ML, distributed systems, cloud-scale automation, and storage or hardware resiliency expertise, including failure analysis, live-site operations, and NVMe/PCIe or firmware experience.
The summary above was generated by AI
Overview

Microsoft Silicon and Cloud Hardware Infrastructure Engineering (SCHIE) is the team behind Microsoft’s expanding Cloud Infrastructure and responsible for powering Microsoft’s “Intelligent Cloud” mission. CHIE delivers the core infrastructure and foundational technologies for Microsoft's over 200 online businesses including Bing, MSN, Office 365, Xbox Live, Skype, OneDrive and the Microsoft Azure platform globally with our server and data center infrastructure, security and compliance, operations, globalization, and manageability solutions. Our focus is on smart growth, high efficiency, and delivering a trusted experience to customers and partners worldwide and we are looking for passionate, high-energy engineers to help achieve that mission.  
  
As Microsoft's cloud business continues to grow the ability to deploy new offerings and HW infrastructure on time, in high volume with high quality and lowest cost is of paramount importance. To achieve this goal, the Silicon Cloud Hardware Infrastructure Engineering (SCHIE) team is instrumental in defining and delivering measures of success for hardware design, qualification, fleet support, scale, and sustainability related to Microsoft cloud hardware.  

Azure Memory and Storage Center of Excellence (AMS CoE) is part of the SCHIE organization focusing on Memory and Storage devices going into the Cloud hardware servers. AMS provide memory and storage solutions to Azure, drive memory and storage suppliers to deliver high quality products, meeting our requirements.   

Every hour a server sits unhealthy is an hour of lost customer capacity. We build the software that catches hardware failures before they take a node down and repairs the ones that do, without a human in the loop. Production AI agents reasoning over fleet telemetry, the data platform beneath them, and the tooling engineers use to ship them safely. Across millions of nodes. 


We are looking for a Senior Cloud Hardware Engineer to scale Azure’s Fault Self‑Healing and Failure Prediction systems.


Responsibilities

What you'll do 

  • Build the platform behind Azure's fault self-healing and failure-prediction system with telemetry pipelines, prediction services, decision logic, and automated repair workflows. 

  • Ship AI agents to production: prompt and tool design, retrieval over diagnostics data, evaluation harnesses, guardrails, and the CI/CD path that deploys new agent skills safely. 

  • Close the loop safely: automated remediation with staged rollout, blast-radius limits, and verification. 

  • Own the developer experience: SDKs, APIs, and dashboards that let engineers across the org author and deploy new detection and repair logic themselves. 

  • Build and monitor measurements: prediction precision/recall, false-repair rate, action success rate, and regression gates that block a bad model from shipping. 

What you'll bring 

  • 8+ years building and operating production software, distributed services, data platforms, or large-scale automation.  

  • Python and C# (or C++/Rust), plus cloud-scale data pipelines at high volume. 

  • AI/ML in production, not just experimentation: agent or model serving, evaluation, versioning and rollback, drift and regression monitoring. 

  • LLM application patterns: agent/tool-calling, RAG, structured output and judgment about when an LLM is the wrong answer. 

  • A track record of automation that takes real actions on real infrastructure, with the safety engineering that requires. 

  • Bonus: anomaly detection on time-series data; Kusto/SQL; server hardware, firmware, or datacenter operations. 

Why this role

  • Your impact shows up in fleet availability metrics and customer uptime, quarter over quarter. The AI agents you ship take autonomous action on production infrastructure at a scale very few teams can offer. 

Qualifications
  • Required/minimum qualifications
    • Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 3+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 5+ years technical engineering experience OR equivalent experience

 

  • Other Requirements:
    • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings:
    • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.  

 

  • Preferred Qualifications:
    • Bachelor's Degree in Computer Engineering, Computer Science, Electrical Engineering, or related field AND 8+ years of firmware or embedded systems engineering experience OR Master's Degree in Computer Engineering, Computer Science, Electrical Engineering, or related field AND 6+ years of firmware or embedded systems engineering experience OR equivalent experience
    • 6+ years developing SSD or storage device firmware, including 4+ years working directly with NVMe and PCIe protocols 
    • Demonstrated depth in storage device resiliency and fault analysis — failure mode characterization, error handling and recovery paths, and root-cause investigation of field failures 
    • Experience supporting live-site operations for storage at fleet scale, including on-call ownership and production incident resolution 
    • Track record of owning end-to-end technical design across the full reliability lifecycle: detection, prediction, mitigation, and repair
    • Proven experience building automation-heavy systems that operate safely at hyperscale, with the guardrails, staged rollout, and blast-radius controls that requires 

Hardware Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

HQ

Microsoft Redmond, Washington, USA Office

1 Microsoft Way, Redmond, WA, United States, 98052

Microsoft Bellevue, Washington, USA Office

Bellevue, United States

Microsoft Issaquah, Washington, USA Office

Issaquah, United States

Microsoft Kirkland, Washington, USA Office

Kirkland, United States

Microsoft Mercer Island, Washington, USA Office

Mercer Island, United States

Microsoft Sammamish, Washington, USA Office

Sammamish, United States

Microsoft Seattle, Washington, USA Office

Seattle, United States

Similar Jobs

4 Days Ago
In-Office
Redmond, WA, USA
120K-261K Annually
Senior level
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Designs and builds hyperscale storage fleet resiliency systems for Azure, including live monitoring, fault detection, prediction, isolation, repair automation, and firmware deployment for SSDs, HDDs, and storage accelerators. The role analyzes fleet data, improves reliability and quality, collaborates with hardware, firmware, software teams and suppliers, and supports Azure service stakeholders.
Top Skills: AutomationAzureFirmwareHddNvmePcieSsdTelemetry
2 Hours Ago
In-Office or Remote
99K-135K Annually
Senior level
99K-135K Annually
Senior level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Support export compliance for unmanned aerial systems: draft, submit, and manage ITAR/EAR licenses and commodity jurisdiction requests; classify technical data; manage license lifecycle and documentation; coordinate with stakeholders and provide licensing guidance as an empowered official.
Top Skills: ExcelMicrosoft TeamsOcr EaseSharepoint
2 Hours Ago
In-Office
177K-239K Annually
Senior level
177K-239K Annually
Senior level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Leads interdisciplinary systems engineering for space payloads and advanced technologies. Responsibilities include establishing and tracing requirements, validating and verifying system needs, defining system architectures and interfaces, managing cross-domain interactions, and developing system models to support lifecycle decisions. The role requires an active U.S. Top Secret clearance, U.S. citizenship, and may involve up to 10% domestic or international travel.
Top Skills: 3DxCameo/MsosaCatiaEnterprise ArchitectureModel-Based EngineeringSysmlSystems EngineeringUml

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account