Softchoice Logo

Softchoice

Domain Architect - AI Compute (Nvidia Ecosystem)

Posted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in United States
110K-150K Annually
Expert/Leader
Remote
Hiring Remotely in United States
110K-150K Annually
Expert/Leader
Leads architecture, provisioning, automation, performance validation, and day-two operations for large-scale NVIDIA GPU clusters and AI compute environments. Designs repeatable bare-metal and Kubernetes-based infrastructure using schedulers, monitoring, networking, and infrastructure-as-code. Serves as the technical authority for DGX, HGX, MGX, NVL72, and Cisco AI POD deployments, while supporting pre-sales discovery, bills of materials, solution estimates, client workshops, and executive-level communication.
The summary above was generated by AI
Job Summary & Responsibilities

Minimum Qualifications 

  • 10+ years in HPC, AI, or data center infrastructure engineering, including 7+ years in a customer-facing architecture role. 
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. 
  • Deep working knowledge of NVIDIA GPU platforms across the Hopper and Blackwell architectures, including Grace Hopper and Grace Blackwell superchip systems, and the associated software stack (CUDA, cuDNN, NCCL). 
  • Strong Linux systems engineering skills, including distributions tuned for HPC and AI workloads (Ubuntu, RHEL), kernel tuning, driver management, and system hardening. 
  • Automation and infrastructure-as-code experience, with proficiency in Ansible and Python, and familiarity with declarative tooling such as Terraform. 
  • Demonstrated ability to translate technical architecture into commercial outcomes for senior stakeholders. 

 

Preferred Qualifications 

  • Industry background: experience within a systems integrator (SI) or managed service provider (MSP) environment. 
  • Cluster management: hands-on experience with NVIDIA Base Command Manager and/or NVIDIA Mission Control. 
  • Network integration: understanding of high-speed interconnects (InfiniBand NDR/HDR, RoCEv2) and how they interact with host PCIe and NVLink topologies. 
  • Container platforms: Kubernetes and Red Hat OpenShift installation and administration. 
  • Cloud architecture: AI-oriented compute design on AWS, Azure, Google Cloud, or OCI. 
  • AI and orchestration tooling: exposure to Rafay, NVIDIA Run:ai, NVIDIA Omniverse, Red Hat OpenShift AI, and Kubeflow. 

Certain states and localities require employers to post a reasonable estimate of salary range. A reasonable estimate of the current base pay range for this position is $110,000.00 to $150,000.00 annually. Actual salary will be based on a variety of factors, including shift, location, experience, skill set, performance, licensure and certification, and business needs. The range for this position in other geographic locations may differ. Certain positions may also be eligible for variable incentive compensation, such as bonuses or commissions, that is not included in the base pay.


The well-being of WWT employees is essential. When it comes to our benefits package, WWT has one of the best. We offer the following benefits to all full-time employees:

  • Health and Wellbeing: Health (Medical & Prescription), Dental, and Vision Care, Onsite Health Centers (MO & IL), Employee Assistance Program, Wellness program
  • Financial Benefits: Competitive Pay, Profit Sharing, 401k Plan with Company Matching, Life and Disability Insurance, Flexible Spending Accounts, Tuition Reimbursement
  • Paid Time Off: PTO & Holidays, Parental Leave, Medical Leave, Military Leave, Bereavement, Day of Caring
  • Additional Perks: Family Planning Benefits, Nursing Mothers Benefits, Voluntary Legal, Voluntary Supplemental Accident/Illness/Hospital, Voluntary ID Theft, Pet Insurance, Employee Discount Program

Note: This is not an all-encompassing list and should not be used as a complete description of the plan’s benefits. For more information, see our US benefits website at wwt.com/us-benefits.


We strive to create an environment where all employees are empowered to succeed based on their skills, performance, and dedication. Our goal is to cultivate a culture of belonging that encourages innovation, collaboration, and respect for all team members, ensuring that WWT remains a great place to work for all!


If you require accessibility accommodation(s) or adjustment during any stage of the hiring process, please let your WWT Recruiter know. The recruiter will work with you to understand your needs and help ensure an accessible experience throughout the interview process.


World Wide Technology is an Equal Opportunity Employer.


If you have any questions or concerns about this posting, please email [email protected]


#LI-AF1

#LI-Remote

Preferred Qualifications

World Wide Technology (WWT) strives to make a new world happen. WWT's work benefits clients and partners as much as it does its people and community across the globe.


Founded in 1990, WWT brings together strategy, deep technical expertise and world-class partnerships to help public and private sector organizations design, build and scale intelligent AI, digital, cybersecurity, cloud and infrastructure solutions. Through its Advanced Technology Center (ATC)—a collaborative ecosystem featuring state-of-the-art hardware and software—WWT enables clients and partners to conceptualize, test and validate innovative technology and then deploy solutions at scale using its global integration and distributions capabilities.


With more than 14,000 team members and over 60 locations globally, WWT's culture—grounded in core values and leadership philosophies—has been recognized by Fortune® and Great Place to Work for its commitment to innovation, trust and creating a great place to work for all. WWT provides products and services to large enterprise, global service provider and public sector clients in up to 130 countries across six continents. Softchoice, a World Wide Technology company, supports commercial and SMB markets in the U.S. and Canada.


Want to work with highly motivated individuals on high-performance teams? Join WWT today!


What is the Solutions Consulting & Engineering Team and why join?


Solutions Consulting & Engineering is an organization that is customer-focused and solutions-led. We deliver end-to-end and emerging solutions to drive customer satisfaction and increase profitability and growth. Our world-class management consulting, delivery excellence, and engineering brilliance enable our success. We embody the OneWWT mindset by bringing the right talent at the right time from anywhere within WWT to solve our customer’s problems. Our goal is to bring together business acumen with full-stack technical know-how to develop innovative solutions for our clients’ most complex challenges. 


About the Role 

As Domain Architect - AI Compute, you will be the primary technical authority for the physical and logical lifecycle of high-performance GPU compute fleets across a diverse portfolio of client environments, bridging the gap between architectural design and hands-on execution. You are a builder as much as an advisor: as comfortable configuring a cluster from the CLI as you are explaining that configuration to a C-level audience. 

 

As a global systems integrator, we don't simply operate static cloud environments. We design and deliver purpose-built, high-scale AI factories for some of the world's leading enterprises. In this role you will define the reference standard for compute infrastructure, moving beyond single-server administration to architect repeatable, scalable, and automated compute fabrics. You will act as technical lead on NVIDIA Cloud Partner (NCP) and private enterprise AI cloud deployments, owning the compute layer of the compute, network, and storage stack. 

 

Your time will be split roughly 60/40 between delivering complex AI infrastructure (60%) and providing pre-sales subject matter expertise (40%). You will lead the physical provisioning of NVIDIA DGX SuperPOD, NVIDIA DGX BasePOD, and Cisco AI POD environments, ensuring clients inherit platforms that are genuinely ready for day-2 operations, while helping the sales team scope and cost future deployments. 

 

Key Responsibilities 

Delivery and implementation  

Bare-metal build and provisioning 

  • Lead the physical provisioning of GPU platforms and clusters, including NVIDIA GB200/GB300 NVL72 rack-scale systems, NVIDIA DGX SuperPOD and DGX BasePOD reference architectures, HGX- and MGX-based systems, and Cisco AI PODs. 
  • Use NVIDIA Base Command Manager (BCM) for cluster provisioning, diskless boot, image management, and firmware lifecycle management. 
  • Harden and baseline the host operating system across the fleet. 
  • Establish cluster monitoring and lifecycle management with NVIDIA Mission Control. 
  • Build and execute zero-touch provisioning (ZTP) workflows that turn bare-metal hardware into production-ready nodes. 

 

Scheduler and workload configuration 

  • Define and enforce fair-share policies, fractional GPU allocation using Multi-Instance GPU (MIG), and preemption logic for multi-tenant environments. 
  • Ensure orchestration layers respect hardware topology, including NUMA affinity and PCIe topology, to protect performance. 
  • Implement and tune advanced schedulers: Slurm on bare metal, and NVIDIA Run:ai, Kueue, or Volcano on Kubernetes. 

 

Orchestration and day-2 operations 

  • Deploy and configure multi-cluster management and observability platforms such as Rafay. 
  • Implement high-fidelity telemetry using NVIDIA Data Center GPU Manager (DCGM) to monitor GPU health, thermal throttling, and Xid error rates. 
  • Lead the transition to day-2 operations, ensuring the environment is fully integrated with the client's identity providers and storage backends before handoff. 

 

Performance engineering 

  • Conduct acceptance and validation testing using NCCL tests and HPL to verify cluster performance against expected baselines. 
  • Perform kernel and OS-level tuning (hugepages, sysctl, and driver parameters) to optimize for high-bandwidth InfiniBand and RoCEv2 fabrics. 

 

Pre-sales SME and consulting 

Technical scoping and estimation 

  • Support the sales team by validating customer technical requirements and producing accurate level-of-effort (LOE) estimates for statements of work. 
  • Define the standard operating environment (SOE) used in proposals to ensure repeatability across engagements. 

 

Bill of materials validation and architecture 

  • Own the technical accuracy of the compute bill of materials. 
  • Verify that all components, including memory, NVMe storage, NICs, and optical transceivers, are aligned with NVIDIA-Certified Systems requirements and the qualified component lists in the relevant DGX, HGX, MGX, and NVL72 reference architectures. 
  • Match GPU platform selection to the client's actual workload profile and growth expectations. 

 

Client workshops 

  • Lead technical discovery workshops to determine workload requirements, distinguishing between the very different infrastructure profiles of traditional machine learning, LLM training, LLM inference, and Omniverse digital twin and visualization workloads. 

Similar Jobs

Yesterday
Remote
United States
110K-150K Annually
Expert/Leader
110K-150K Annually
Expert/Leader
Big Data • Cloud • Hardware • Software • App development
Architect and deliver high-scale NVIDIA GPU compute environments for enterprise AI workloads. Lead bare-metal provisioning, cluster management, operating-system hardening, scheduler and workload configuration, monitoring, performance validation, and day-two operations. Serve as a customer-facing technical authority, supporting presales discovery, architecture, bills of materials, estimation, and client workshops. Design repeatable AI infrastructure across DGX, HGX, MGX, NVL72, and Cisco AI POD platforms.
Top Skills: AnsibleAWSBlackwellCisco Ai PodCudaCudnnGb200Gb300 Nvl72GCPGrace BlackwellGrace HopperHgxHopperHplInfinibandKubeflowKubernetesKueueLinuxMgxAzureMulti-Instance GpuNcclNumaNvidia Base Command ManagerNvidia Data Center Gpu ManagerNvidia Dgx BasepodNvidia Dgx SuperpodNvidia GpusNvidia Mission ControlNvidia OmniverseNvidia Run:AiNvlinkOciPciePythonRafayRed Hat OpenshiftRed Hat Openshift AiRhelRocev2SlurmTerraformUbuntuVolcano
14 Minutes Ago
Remote or Hybrid
US
135K-210K Annually
Senior level
135K-210K Annually
Senior level
Artificial Intelligence • Cloud • Payments • Software • Business Intelligence • Generative AI • Automation
Lead product vision, strategy, adoption, roadmap prioritization, user research, metrics, and cross-functional execution for an AI-powered B2B insurance submissions platform. Partner across design, engineering, marketing, customer experience, and operations; manage integrations with adjacent products; resolve dependencies; communicate strategy and outcomes to senior leadership; and mentor other product managers.
Top Skills: AgileAIApplied EpicConfluenceJIRA
An Hour Ago
Remote or Hybrid
264K-500K Annually
Senior level
264K-500K Annually
Senior level
Cloud • Software
Own strategic enterprise accounts across New England for ThousandEyes, driving new-logo acquisition and customer expansion. Build executive relationships, independently generate pipeline, leverage Cisco’s sales ecosystem, and lead complex enterprise sales cycles through technical evaluation, negotiation, and close. Manage six- and seven-figure transactions, develop strategic account plans, expand customers into $1M+ ARR relationships, and maintain quota attainment, pipeline coverage, and forecast accuracy.
Top Skills: Cisco ThousandeyesCloud NetworkingCybersecurityEnterprise Technology SolutionsObservabilitySaaS

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account