Designworks Talent LLC Logo

Designworks Talent LLC

Vice President - AI Infrastructure Engineering

Posted 2 Days Ago
Be an Early Applicant
Hybrid
Bellevue, WA, USA
Expert/Leader
Hybrid
Bellevue, WA, USA
Expert/Leader
Lead the engineering organization responsible for bringing data center hardware into production-ready AI and HPC infrastructure. Own automated provisioning, Linux deployment, configuration, validation, GPU clusters, Kubernetes environments, monitoring, and workload readiness. Establish standards across servers, GPUs, networking, storage, firmware, and software automation while partnering with hardware, network, SRE, data center operations, and software teams. Build and develop a high-performing engineering organization capable of scaling infrastructure across thousands of servers or GPUs.
The summary above was generated by AI
Vice President AI Infrastructure Engineering

Bellevue, WA Area | Hybrid | Senior Leadership

A rapidly growing, well-funded technology company is seeking a Vice President AI Infrastructure Engineering Leader to lead the engineering organization responsible for transforming newly deployed data center hardware into reliable, production-ready compute infrastructure.
This is a high-impact leadership opportunity for someone who combines deep technical expertise in AI/HPC infrastructure with strong engineering leadership. The ideal candidate understands both the physical infrastructure layer and the software automation required to operate large-scale GPU and compute environments efficiently.
You’ll operate at the intersection of servers, GPUs, Linux, networking, Kubernetes, distributed systems, automation, and infrastructure software, helping establish the architecture, standards, tooling, and engineering practices required to deploy and operate infrastructure at scale.


What You'll Do
  • Lead the team of 15+ responsible for data center infrastructure bring-up and production readiness.

  • Own the platform lifecycle from installed hardware through automated provisioning, configuration, validation, and workload readiness.

  • Build and scale automation for bare-metal provisioning, Linux deployment, configuration management, and infrastructure validation.

  • Lead deployment and configuration of GPU clusters, Kubernetes environments, and distributed compute infrastructure.

  • Establish engineering standards for servers, GPUs, networking, storage, firmware, and system configuration.

  • Drive infrastructure automation using Terraform, Ansible, Bash, and similar technologies.

  • Oversee integration with technologies such as Redfish, IPMI, BMCs, PXE, MAAS, Ironic, Foreman, or comparable platforms.

  • Partner closely with Network, Hardware/GPU, Data Center Operations, SRE, and Software Engineering teams to deliver production-ready infrastructure.

  • Establish automated testing, health checks, monitoring, and validation processes to identify infrastructure issues before workloads reach production.

  • Improve deployment speed, reliability, automation, scalability, and operational efficiency.

  • Build and develop a highly capable engineering organization while establishing processes that can scale with the business.

What We're Looking For
  • 12+ years of experience across infrastructure, software, systems, platform engineering, or related technical disciplines.

  • 5+ years of engineering leadership experience, including managing and developing highly technical teams.

  • Proven experience building and operating large-scale data center, cloud, HPC, or AI infrastructure.

  • Strong technical understanding of Linux, distributed systems, networking, and infrastructure automation.

  • Hands-on understanding of Kubernetes, containers, and infrastructure-as-code.

  • Demonstrated ability to lead complex infrastructure deployments and bring new environments into production.

  • Ability to operate comfortably across both hardware and software organizations.

  • Strong communication and cross-functional leadership skills.

  • A hands-on, high-ownership leadership style with the ability to operate effectively in a fast-moving, build-from-the-ground-up environment.

Preferred Experience

Experience in one or more of the following areas is highly valued:

  • GPU infrastructure, NVIDIA platforms, AI or HPC environments.

  • Bare-metal provisioning technologies such as MAAS, Ironic, xCAT, Foreman, or similar.

  • Hardware management technologies including Redfish, IPMI, BMC, PXE, and firmware management.

  • NVIDIA technologies such as CUDA, NVML, DCGM, NVIDIA drivers, or GPU Operator.

  • High-performance networking including InfiniBand, RoCE, RDMA, or high-speed Ethernet.

  • Cluster orchestration and scheduling technologies such as Kubernetes, Slurm, or similar.

  • Automated infrastructure validation and hardware health testing.

  • Experience scaling infrastructure across thousands of servers or GPUs.


The Opportunity

This is an opportunity to join an organization at an early and highly consequential stage of its growth. You’ll have significant influence over architecture, automation, engineering standards, tooling, and team development, rather than simply inheriting an established infrastructure environment.


The company operates with a startup mentality of fast, lean, highly collaborative, and high ownership while having the resources to build infrastructure for significant scale.

The role is particularly well suited to a leader who enjoys building something new, moving quickly, solving complex infrastructure challenges, and creating software-driven systems that replace manual processes with scalable automation.


Location
  • Bellevue, WA area

  • Hybrid work model with three days per week in the office

  • Candidates currently outside the area may be considered if they are willing to relocate

  • U.S. work authorization required; visa sponsorship is not currently available

 

Similar Jobs

26 Minutes Ago
Remote or Hybrid
United States
161K-241K Annually
Senior level
161K-241K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Develop low-level, high-performance instrumentation within CICS and z/OS using HLASM. Design production features for mission-critical mainframe environments, diagnose complex system behavior through dumps and traces, analyze performance bottlenecks, and collaborate with an experienced engineering team. The role requires deep z/OS and CICS internals expertise, systems programming knowledge, and proficiency with JCL, SDSF, and USS.
Top Skills: AICCicsCobolCtgDb2HlasmJavaJclMqPl/ISdsfUssWebsphereZ/OsZ/Os Connect
44 Minutes Ago
In-Office
Bellevue, WA, USA
28-52 Hourly
Internship
28-52 Hourly
Internship
Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Support Ericsson’s CTO organization on emerging technology enablement, customer-focused innovation, and technology strategy initiatives. Research industry trends, analyze information with AI-enabled tools, prepare recommendations and presentations, and contribute to network transformation, automation, standardization, regulatory, and technology planning activities. Collaborate with technical and business stakeholders while gaining exposure to telecommunications, cloud technologies, AI, customer engagement, and strategic technology decisions.
Top Skills: 5G6GAIAutomationCloud ComputingCloud-Native PlatformsNetwork TechnologiesPhysical AiProgrammingScriptingWireless Networks
45 Minutes Ago
Hybrid
Seattle, WA, USA
155K-248K Annually
Senior level
155K-248K Annually
Senior level
AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Lead creative strategy and production for paid social, display, programmatic, paid search, video, creator, and app store assets. Develop channel-specific strategies, platform partnerships, innovation programs, and testing frameworks. Collaborate with marketing, growth, brand, design, copy, agencies, and technology partners to optimize creative performance against conversion, CPA, and iROAS goals while maintaining brand consistency.
Top Skills: A/B TestingAIApp StoreDv360Google DisplayGoogle PlayMetaMultivariate TestingPinterestTiktokYoutube

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account