NVIDIA Logo

NVIDIA

Principal Software Engineer – Infrastructure

Posted 19 Days Ago
Be an Early Applicant
In-Office
Redmond, WA, USA
248K-391K Annually
Expert/Leader
In-Office
Redmond, WA, USA
248K-391K Annually
Expert/Leader
Leads the architecture and development of enterprise infrastructure automation and configuration management platforms across compute, storage, networking, data centers, cloud, and hybrid environments. Defines long-term technical strategy, builds secure and scalable integrations, applies AI to configuration intelligence and remediation, and remains hands-on with software development. Mentors senior engineers, establishes engineering standards, drives enterprise adoption, and partners with leadership on technical roadmaps and business outcomes.
The summary above was generated by AI

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

We are looking for a Principal Software Engineer to join our Configuration Management team and define the future of enterprise infrastructure automation, configuration management and operational platforms. This is a deeply technical, hands-on leadership role, you will architect and build foundational software systems that manage infrastructure consistently across compute, storage, networking, data centers, cloud, and hybrid environments. You will establish the multi-year technical direction, identify organization-wide challenges, and drive large-scale transformations from initial strategy through architecture, implementation, production adoption, and measurable outcomes.

What You’ll Be Doing:

  • Define the multi-year technical vision and architecture for IT infrastructure automation, configuration management, orchestration, and self-service platforms.

  • Set the technical direction for infrastructure automation and configuration management across technologies such as Ansible Automation Platform, AWX, Salt or equivalent platforms.

  • Establish the strategy for applying AI to configuration management, including configuration intelligence, drift and compliance analysis, change-risk identification, root-cause assistance, intelligent recommendations, and guarded automated remediation.

  • Remain deeply hands-on by developing prototypes and production software, reviewing critical code and designs, and resolving the most challenging technical and scalability problems.

  • Build secure and scalable integrations across infrastructure platforms, cloud services, CMDB, secrets management, observability, and enterprise data systems.

  • Set and drive configuration management strategy across networking, storage, and compute domains. Bring deep cross-domain technical expertise to identify gaps and challenge current methods. Establish architectures and engineering standards. Lead organizations in implementation and adoption within large-scale environments.

  • Ensure platforms meet enterprise requirements for availability, scalability, performance, security, disaster recovery, observability, and operational support.

  • Act as a force multiplier by mentoring senior and staff engineers, raising engineering standards, facilitating architectural decisions, and developing technical leaders across teams.

  • Partner with engineering and executive leadership to translate critical business challenges into technical strategy, prioritized roadmaps, and measurable business outcomes.

What We Need to See:

  • 15+ years of progressive software engineering experience, with a sustained record of delivering complex, business-critical platforms and distributed systems.

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a similar domain, or a Master’s degree or equivalent experience.

  • Deep knowledge of enterprise configuration management or infrastructure automation using technologies such as Ansible Automation Platform, AWX, Salt or equivalent platforms.

  • Deep hands-on expertise in Go, Python, Java, or a comparable systems programming language, including APIs, concurrency, distributed systems, testing, debugging, and performance engineering.

  • Strong hands-on experience in Linux, Kubernetes, containers, cloud and hybrid infrastructure, CI/CD and Infrastructure as Code.

  • Strong experience developing and implementing enterprise data and automation pipelines using databricks or comparable large-scale data platforms.

  • Proven experience defining, owning, and evolving the architecture of large-scale infrastructure platforms operating across multiple teams, data centers, or cloud environments.

  • Demonstrated ability to identify organization-wide business and technical challenges, establish a clear strategy, and drive implementation across teams.

  • Outstanding communication and technical leadership skills, with a proven ability to build consensus, influence senior leaders, mentor experienced engineers, and lead through ambiguity.

Ways to Stand Out from the Crowd:

  • Experience in architecting, deploying and managing Ansible Automation Platform at enterprise scale.

  • Experience developing platforms that manage large global infrastructure fleets across data centers, public clouds, compute, storage, and networking environments.

  • A history of leading the development and enterprise-wide adoption of an AI-enabled configuration management, infrastructure automation, or autonomous operations platform.

  • Experience using AI to solve infrastructure challenges such as configuration drift, compliance, change-risk analysis, incident diagnosis, capacity management, predictive operations, or automated remediation.

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/ 

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 248,000 USD - 391,000 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 29, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

HQ

NVIDIA Seattle, Washington, USA Office

4545 Roosevelt Way NE 6th Floor, Seattle, Washington, United States, 98105

NVIDIA Bellevue, Washington, USA Office

Bellevue, United States

NVIDIA Redmond, Washington, USA Office

Redmond, United States

Similar Jobs

3 Days Ago
Hybrid
Seattle, WA, USA
231K-399K Annually
Expert/Leader
231K-399K Annually
Expert/Leader
AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Leads the strategy, design, implementation, and operation of Expedia’s cloud platform and infrastructure. Responsibilities include Kubernetes and AWS platform engineering, infrastructure as code, CI/CD, observability, security, reliability, cost optimization, capacity planning, incident response, and cloud migration. The role contributes production code, drives cross-team adoption and technical standards, collaborates with stakeholders, and mentors engineers.
Top Skills: AlbApi GatewaysAuto ScalingAWSCi/CdCloudFormationDatadogDockerEc2EcsEksElbGithub ActionsGitlab CiGoGrafanaHelmIamInfrastructure As CodeIstioJaegerJavaJenkinsKubernetesLinkerdObservabilityOpentelemetryPrometheusPythonRdsS3SpinnakerSreTerraformVpcVpc Lattice
5 Days Ago
Hybrid
Seattle, WA, USA
263K-355K Annually
Senior level
263K-355K Annually
Senior level
Artificial Intelligence • Internet of Things • Semiconductor
Designs, builds, and operates large-scale AI compute infrastructure for training, fine-tuning, evaluation, and inference. Responsibilities include managing Kubernetes clusters, enabling CPU and GPU systems, optimizing scheduling and capacity, troubleshooting distributed infrastructure, and automating provisioning, upgrades, monitoring, and maintenance. The role partners closely with AI researchers and engineers to improve infrastructure reliability, scalability, performance, and developer productivity.
Top Skills: Argo CdAws EksContainersCudaDcgmEfaGoGpuGrafanaHelmKubernetesLinuxNcclNetworkingNvlinkNvswitchPrometheusPythonPyTorchRaySglangStorageTensorrt-LlmTerraformVllm
25 Days Ago
Hybrid
2 Locations
235K-414K Annually
Expert/Leader
235K-414K Annually
Expert/Leader
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Lead technical strategy, architecture, and implementation for ML inference platform services. Design and scale distributed, high-throughput inference systems, collaborate across teams, drive availability, scalability, operational excellence, cost management, and provide company-wide technical direction and mentorship.
Top Skills: Distributed SystemsGpuKubernetesLlm InferenceMl Inference PlatformPyTorchRpcTensorFlow

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account