Microsoft Logo

Microsoft

Senior Software Engineer

Posted 12 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
120K-261K Annually
Senior level
Remote
Hiring Remotely in United States
120K-261K Annually
Senior level
Design and build monitoring, telemetry, and data pipelines to operate flagship supercomputers at scale. Diagnose and troubleshoot GPU hardware, networking, datacenter and software stack issues, implement systemic mitigations, improve observability, and drive reliability and performance improvements through incident reviews and tooling.
The summary above was generated by AI
Overview

Microsoft Azure High Performance Computing & AI Engineering (HPC & AI Eng) team is responsible for managing the core platform & fleet of AI High Performance Computing products that customers use to run their most performant and demanding workloads. The AI Customer Experience (AICE) engineering team within the HPC & AI Eng. team is on the frontlines managing the flagship supercomputers used by top tier AI customers that enable breakthroughs such as ChatGPT and are highlighted in Top500, MLPerf and Graph500 rankings.

We run lean, obsess about customer experience and use evidence-based approach to decision making. We have live-site first, metrics-driven culture that prevents us from accumulating debt and necessity to put out fires on daily basis. You will be in a position that carries a ton of responsibility and provides opportunities to directly impact customers satisfaction.

As a Senior Software, you will design and develop capabilities needed to monitor and efficiently operate across the infrastructure and fleet of supercomputers at scale. You will diagnose and troubleshoot the largest scale supercomputing systems across the infrastructure stack (GPU hardware, networking, datacenter and core software). To enable first to know of critical incidents impacting customer capacity, you will create end to end data pipelines that process and synthesize large volume of telemetry, log files and other data sources to create actionable alerts.


Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.


Responsibilities
  • Contribute to improving key metrics such as Job Mean Time to Interrupt, Nodes in Service, Mean Time to Resolve on flagship supercomputers.
  • Manages operations of supercomputers by responding quickly to mitigate issues.
  • Implements systemic solutions and mitigations to more complex issues impacting performance or functionality of supercomputers
  • Reviews and writes incident postmortem and presents insights that drive changes to reduce or eliminate incidents.
  • Independently improves troubleshooting guides (TSGs), wikis, tests, and telemetry, adding comprehensive observability and monitoring capabilities.
  • Proactively seeks new knowledge and adapts to new trends, technical solutions, and patterns that will improve the availability, reliability, efficiency, observability, and performance of supercomputers while also driving consistency in monitoring and operations at scale. 

Qualifications

Required Qualifications:

  • Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, OR Java, JavaScript, or Python
    • OR equivalent experience.

Other Requirements:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings: 
    • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.

Preferred Qualifications:

  • Bachelor's Degree in Computer Science
    • OR related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, OR Python
    • OR Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience.
  • Experience diagnosing and troubleshooting GPU based systems such as H100, A100 or networking technologies such as InfiniBand or Ethernet.
  • Experience with large scale data pipelines using tools such as Prometheus, Grafana, etc.

Software Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

HQ

Microsoft Redmond, Washington, USA Office

1 Microsoft Way, Redmond, WA, United States, 98052

Microsoft Bellevue, Washington, USA Office

Bellevue, United States

Microsoft Issaquah, Washington, USA Office

Issaquah, United States

Microsoft Kirkland, Washington, USA Office

Kirkland, United States

Microsoft Mercer Island, Washington, USA Office

Mercer Island, United States

Microsoft Sammamish, Washington, USA Office

Sammamish, United States

Microsoft Seattle, Washington, USA Office

Seattle, United States

Similar Jobs

6 Hours Ago
Remote
USA
180K-240K Annually
Senior level
180K-240K Annually
Senior level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
Design and build scalable, automated evaluation pipelines and infrastructure to validate speech, audio, and multimodal models. Translate research benchmarks into enforceable pass/fail gates, run batch and streaming tests, operate canaries and continuous monitoring, integrate quality gates into CI/CD, and partner with research, training, inference, and infra teams to prevent regressions and ensure production model quality.
Top Skills: Agent FrameworksAnomaly DetectionCi/CdContainersDockerGoGpu ClustersGrafanaKubernetesLlmsMultimodal ModelsPythonRagReact NativeRust
Yesterday
Easy Apply
Remote
United States
Easy Apply
128K-155K Annually
Senior level
128K-155K Annually
Senior level
Healthtech • Software
Lead full‑stack engineering efforts to design, build, and operate production healthcare systems. Own epic-level features, mentor engineers, design for reliability and scalability, integrate with partners, and ensure launch readiness across backend (Java/Groovy), frontend (React/TypeScript), data (MongoDB), and event-driven pipelines (Kafka).
Top Skills: AWSCypressFhirGitGrailsGroovyHl7JavaJavaScriptJestJunitKafkaMongoDBPlaywrightReactTypescript
Yesterday
Easy Apply
Remote
United States
Easy Apply
173K-255K Annually
Senior level
173K-255K Annually
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Join a small AI-native engineering team building and operating Affirm's bank systems. Own end-to-end bank infrastructure, durable workflows, vendor integrations, data platform, CI/CD, monitoring, and examiner-ready documentation. Use AI-assisted tooling to accelerate development, collaborate with Compliance/Risk/Finance, and mentor growing engineering talent.
Top Skills: AWSCi/CdClaudeCodexCursorIamSnowflakeTerraform

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account