Microsoft Logo

Microsoft

HPC Operations Engineering Manager

Reposted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in United States
166K-331K Annually
Senior level
Remote
Hiring Remotely in United States
166K-331K Annually
Senior level
Lead a team of SREs to operate large-scale hybrid CPU/GPU AI training and inference infrastructure. Design observability, automation, deployments, incident management, security/compliance, and collaborate with ML teams to accelerate research-to-production workflows.
The summary above was generated by AI
Overview

 

Microsoft AI is seeking  an experienced High Performance Computing Operations Engineering Manager to join our infrastructure team on the MAI SuperIntelligence Team. In this role, you’ll lead a team of Site Reliability Engineers who blend software engineering and systems engineering to keep our large-scale distributed AI infrastructure reliable and efficient. You’ll work closely with ML researchers, data engineers, and product developers to design and operate the platforms that power training, fine-tuning, and serving generative AI models. 

 

Microsoft Superintelligence Team  

Microsoft Superintelligence Team’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. This role is part of Microsoft AI's Superintelligence Team. The MAIST is a startup-like team inside Microsoft AI, created to push the boundaries of AI toward Humanist Superintelligence — ultra-capable systems that remain controllable, safety-aligned, and anchored to human values.   

Our mission is to create AI that amplifies human potential while ensuring humanity remains firmly in control. We aim to deliver breakthroughs that benefit society — advancing science, education, and global well-being. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in — come and join us as we work on our next generation of models!  


By applying to this Mountain View, CA position, you are required to be local to the San Francisco area and in office 4 days a week.    


Responsibilities

Responsibilities 

  • Team leadership: Lead a team of experienced SREs to ensure uptime, resiliency and fault tolerance of AI model training and inference systems.  

  • Observability: Design and help maintain monitoring, alerting, and logging systems to provide real-time visibility into model serving pipelines and infra. 

  • Automation & Tooling: Lead building of automation for deployments, incident response, scaling, and failover in hybrid cloud/on-prem CPU+GPU environments. 

  • Incident Management: Lead on-call rotations, troubleshoot production issues, conduct blameless postmortems, and drive continuous improvements. 

  • Security & Compliance: Ensure data privacy, compliance, and secure operations across model training and serving environments. 

  • Collaboration: Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows. 


Qualifications

Required Qualifications

  • Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with Site Reliability Engineering, DevOps, or Infrastructure Engineering Leadership roles AND 8+ years experience with Kubernetes, Docker, and container orchestration, AND 6+ years experience with programming/scripting skills not limited to Python, Go, or Bash  

  • OR equivalent experience 

Preferred Qualifications:

  • Master's Degree in Computer Science or related technical field AND 12+ years technical engineering experience AND 10+ years experience with Kubernetes, Docker, and container orchestration, AND 10+ years' experience with public cloud platforms like Azure/AWS/GCP and infrastructure-as-code
    • OR equivalent experience
  • 6+ years people management experience. 

  • 8+ years experience in monitoring & observability tools (Grafana, Datadog, OpenTelemetry, etc.). 

  • Knowledge of CI/CD pipelines for Inference and ML model deployment. 

  • Solid knowledge of distributed systems, networking, and storage. 

  • Experience running large-scale GPU clusters for ML/AI workloads (preferred). 

  • Familiarity with ML training/inference pipelines. 

  • Experience with high-performance computing (HPC) and workload schedulers ( Kubernetes operators). 

  • Background in capacity planning & cost optimization for GPU-heavy environments 

 


Software Engineering IC6 - The typical base pay range for this role across the U.S. is USD $165,600 - $296,400 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $220,800 - $331,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

HQ

Microsoft Redmond, Washington, USA Office

1 Microsoft Way, Redmond, WA, United States, 98052

Microsoft Bellevue, Washington, USA Office

Bellevue, United States

Microsoft Issaquah, Washington, USA Office

Issaquah, United States

Microsoft Kirkland, Washington, USA Office

Kirkland, United States

Microsoft Mercer Island, Washington, USA Office

Mercer Island, United States

Microsoft Sammamish, Washington, USA Office

Sammamish, United States

Microsoft Seattle, Washington, USA Office

Seattle, United States

Similar Jobs

4 Minutes Ago
Remote or Hybrid
Senior level
Senior level
Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Lead and manage tax compliance and consulting engagements for technology clients, mentor tax teams, develop tax strategies, prepare and review returns, conduct research on complex issues, and maintain strong client relationships while ensuring timely, proactive service.
4 Minutes Ago
Remote or Hybrid
77K-104K Annually
Senior level
77K-104K Annually
Senior level
Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Manage day-to-day financial reporting and general ledger accuracy for nonprofit clients, lead month-end close, prepare monthly reporting packages and KPIs, oversee account classifications and balance sheet reviews, maintain client procedure manuals, coordinate project tasks, supervise and estimate accountants' work, and support development of best practices and cross-team collaboration. Requires nonprofit grant reporting expertise.
Top Skills: Accounting SoftwareMicrosoft Office Suite
4 Minutes Ago
Remote or Hybrid
Senior level
Senior level
Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Lead tax compliance and consulting engagements for technology clients; manage and mentor a tax team; prepare and review federal and multi-state tax returns; perform tax research; develop tax strategies; maintain strong client relationships and ensure proactive, timely service.

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account