NVIDIA Logo

NVIDIA

Distinguished Engineer, AI Resiliency Lead

Sorry, this job was removed Sorry, this job was removed at 08:13 p.m. (PST) on Friday, May 30, 2025
Be an Early Applicant
In-Office
2 Locations
In-Office
2 Locations

Similar Jobs

6 Hours Ago
5 Locations
210K-283K
Senior level
210K-283K
Senior level
Cloud • Information Technology • Machine Learning
The Senior Director will lead risk management, program execution, and executive enablement in supply chain strategy for AI infrastructure.
Top Skills: Enterprise Program ManagementRisk ManagementSupply Chain Management
6 Hours Ago
Remote
Hybrid
9 Locations
174K-254K
Senior level
174K-254K
Senior level
Fintech • HR Tech
The Senior Staff Content Designer will drive growth through content, creating scalable messaging systems across customer journeys while collaborating with multiple teams and setting long-term content strategy.
Top Skills: Ai-Generated Content
6 Hours Ago
Remote
Hybrid
68 Locations
89K-315K
Senior level
89K-315K
Senior level
Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
As a Senior Manager, you will lead projects, manage analytics teams, and transform data into strategic insights for clients at PwC.
Top Skills: AlteryxLookerPower BIPythonRSQLTableau

We are seeking a Distinguished Engineer to lead AI Resiliency at NVIDIA!

Join NVIDIA and help push the boundaries of AI. In this role, you will architect, design, and develop world-class software resiliency features for training ground breaking AI models on the largest AI superclusters in the world. Leading a team of cross-functional experts, you will drive and shape our end-to-end AI software stack, ensuring seamless training of frontier models on industry-leading frameworks like PyTorch and JAX/XLA, with near-zero downtime. Your optimizations will span from algorithmic innovations to robust software architecture, with a significant impact on NVIDIA’s most critical customers. This highly visible role demands exceptional technical expertise and leadership across organizations, with direct exposure to NVIDIA's senior leadership.

What You'll Be Doing:

  • Define a scalable software architecture to enable single-job resilient training on hundreds of thousands of GPUs with minimal downtime.

  • Design and deliver modular, resilient software features to support large-scale AI training for our top customers.

  • Innovate and evolve resilient architecture designs to achieve stringent uptime requirements (downtime < 1%), through solutions like in-memory check-pointing, in-process restart, and anomaly/SDC detection.

  • Collaborate closely with internal partners, spearheading successful project execution and communicating regular progress updates to senior leadership.

What We Need to See:

  • A Master’s or Ph.D. in Computer Science, Electrical or Computer Engineering from a top-tier university, or equivalent experience.

  • 15+ years of experience in software architecture or related fields, with a deep understanding of AI-optimized systems.

  • Excellent and proven ability to collaborate and communicate effectively across multiple engineering teams.

  • At least 5 years of hands-on experience in software development on high-complexity projects involving HPC or AI.

Ways to Stand Out from the Crowd:

  • Proven experience with large-scale AI supercomputing applications, particularly in the training phase.

  • 5+ years of experience with using and contributing to modern AI frameworks like PyTorch and JAX/XLA, specifically for large-scale training workloads.

  • A strong passion for designing system architectures tailored for AI, covering CPU, GPU, memory, storage, and networking.

  • Hands-on involvement in the entire lifecycle—from design to deployment—of large-scale High-Performance Computing (HPC) systems.

  • Experience in implementing HPC software development best practices in large-scale systems.

NVIDIA continues to expand its presence in the Datacenter space, and our team plays a pivotal role in enhancing the value of our rapidly growing datacenter deployments. We also drive a data-driven approach to hardware design and system software development. You will collaborate with a wide array of teams across NVIDIA, including deep learning research, CUDA kernel and framework development, and silicon architecture. NVIDIA is widely recognized as one of the technology industry’s most desirable employers, with some of the most dedicated and innovative minds working with us. If you’re creative, driven, and autonomous, we want to hear from you!

The base salary range is 308,000 USD - 471,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

HQ

NVIDIA Seattle, Washington, USA Office

4545 Roosevelt Way NE 6th Floor, Seattle, Washington, United States, 98105

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute
By clicking Apply you agree to share your profile information with the hiring company.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account