Top Tech Jobs & Startup Jobs in Seattle, WA

3 Days AgoSaved
In-Office or Remote
4 Locations
108K-173K Annually
Senior level
108K-173K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Provide Tier 1 support for complex cloud platforms, troubleshoot distributed software and customer issues, investigate root causes, improve operational workflows, create runbooks and documentation, build support tooling, and coordinate with engineering, SRE, and other internal teams. The role supports production systems through an on-call rotation and requires expertise across cloud infrastructure, networking, storage, Kubernetes, Linux, and DevOps tooling, with GPU, MLOps, HPC, or SLURM experience preferred.
Top Skills: AWSAzureBlob StorageBlock StorageDatabasesDevOpsDistributed Training SystemsFile StorageGoogle Cloud Platform (Gcp)Gpu WorkloadsHigh-Performance Computing (Hpc)InfrastructureKubernetesLinuxMachine Learning InfrastructureMlopsNetworkingOracle Cloud Infrastructure (Oci)SlurmStorage
5 Days AgoSaved
Remote or Hybrid
2 Locations
152K-288K Annually
Senior level
152K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Develop CUDA-Q core infrastructure for inter-device communication and efficient execution across hybrid quantum-classical systems. Build extensible compiler toolchains integrating quantum architecture components, solve problems spanning compilers, high-performance computing, and quantum computing, and collaborate on software design. Improve development processes and infrastructure while delivering performant, robust production software.
Top Skills: CudaCuda-QDistributed SystemsFpgaGpuHigh-Performance ComputingLlvmMlirQuantum Computing
12 Days AgoSaved
In-Office or Remote
2 Locations
184K-357K Annually
Senior level
184K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, deploy, maintain, tune, and secure enterprise Linux environments supporting high-performance EDA compute farms. Lead infrastructure-as-code automation, configuration management, continuous deployment, and tooling development using Python, Bash, or Go. Manage LDAP/SSSD naming services, NFS and Autofs storage, virtualization, container infrastructure, registries, networking, and infrastructure telemetry. Support highly available, scalable systems and potentially integrate cloud platforms, CI/CD pipelines, EDA toolchains, and HPC clusters.
Top Skills: AnsibleAutofsAWSBashCadenceDockerGCPGithub ActionsGitlab CiGoInfrastructure As CodeJenkinsKubernetesLdapLinuxLsfMicrovmsNfsPodmanPythonRpm-Based Linux DistributionsSdnSlurmSssdSynopsysTcpdumpTerraformWireshark
12 Days AgoSaved
In-Office or Remote
2 Locations
184K-357K Annually
Senior level
184K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design, deploy, operate, and automate reliable large-scale EDA infrastructure services on NVIDIA hardware. Build tools that reduce manual work, support distributed private and public cloud systems, define service-level objectives, improve observability, and participate in incident response and on-call rotations. The role also provides systems-design guidance to peer teams and may involve Kubernetes, OpenStack, Docker, Slurm, multi-cloud infrastructure, networking, storage, and GPU-based systems.
Top Skills: Ai AgentsBare Metal As A ServiceCloud InfrastructureContainersContinuous IntegrationDistributed SystemsDockerGoInfrastructure AutomationKubernetesLinuxMcp ServersNetworkingNvidia GpusOpenstackPythonSlurmStorageWorkflow Engines
13 Days AgoSaved
Remote
5 Locations
248K-391K Annually
Senior level
248K-391K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Leads a distributed team securing NVIDIA DGX Cloud’s large-scale GPU infrastructure. Responsibilities include hiring and developing engineers, setting strategy for security control planes and foundational services, delivering measurable risk reduction, maintaining engineering standards, communicating security risks to executives, and collaborating with platform, SRE, networking, and security teams. The role requires strong engineering management, infrastructure or security expertise, cloud-native fluency, and influence across partner organizations.
Top Skills: Cloud-Native ArchitectureDistributed SystemsGpu InfrastructureHigh-Performance ComputingIdentity And Access ManagementKubernetesPolicy EnforcementVulnerability Management
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
14 Days AgoSaved
In-Office or Remote
2 Locations
184K-357K Annually
Senior level
184K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Build and operate AI infrastructure systems, including telemetry pipelines, observability dashboards, incident-management automation, hardware and software catalogs, service ownership data models, and self-service discovery tools. Lead cross-functional infrastructure initiatives and develop AI-assisted operational tooling. The role requires production software and infrastructure expertise, programming experience, and the ability to improve reliability, visibility, incident response, and operational efficiency across NVIDIA GPU cloud environments.
Top Skills: Ai Agent FrameworksConfiguration Management Databases (Cmdb)GoJavaMachine Learning ModelsObservability PlatformsPythonSaas Incident-Management ToolsTypescript
14 Days AgoSaved
In-Office or Remote
2 Locations
152K-288K Annually
Senior level
152K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design and operate large-scale infrastructure supporting GPU- and CPU-based EDA workloads. Build automation for provisioning, configuration, deployment, lifecycle management, monitoring, health remediation, cluster enrollment, and recovery across Linux compute fleets. Integrate services with schedulers, infrastructure systems, and observability platforms. Participate in incident response, root-cause analysis, capacity planning, and reliability improvements while collaborating with EDA, networking, storage, and hardware teams.
Top Skills: Bright Cluster ManagerCluster SchedulingCpu InfrastructureDistributed SystemsEdaFirmwareGoGpu InfrastructureKubernetesLinuxLsfNetworkingObservabilityPythonSlurmStorage
17 Days AgoSaved
In-Office or Remote
4 Locations
184K-288K Annually
Senior level
184K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Designs and deploys large-scale AI and HPC GPU cloud infrastructure for NVIDIA Cloud Partners. Serves as the primary technical advisor across solution design, development, integration, production, debugging, and customer adoption. Builds proof-of-concepts, integrates AI software stacks, libraries, frameworks, and models, and delivers technical presentations and workshops. Partners with engineering, product, sales, and business teams to secure opportunities, guide product strategy, and support complex customer engagements throughout the lifecycle.
Top Skills: AWSBase Command ManagerDcgmDeep Learning FrameworksDistributed Cloud ArchitecturesFoundation ModelsGoogle Cloud PlatformGpu Cloud InfrastructureHybrid Cloud InfrastructureKubernetesLarge Language ModelsLoad BalancingAzureMlopsNcclNemo RetrieverNvidia Cuda-XNvidia DynamoNvidia GpusNvidia Mission ControlNvidia Nemo FrameworkNvidia NemotronNvidia Triton Inference ServerPythonPyTorchSlurmTensorFlowTensorrtTensorrt-LlmUfm
17 Days AgoSaved
In-Office or Remote
4 Locations
224K-357K Annually
Senior level
224K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Develop and deploy deep learning models and petabyte-scale multimodal data curation pipelines for foundation-model training. Responsibilities include document extraction, OCR, layout and table analysis, deduplication, distributed processing, dataset and metric development, experimentation, production scaling, and deployment through NVIDIA Inference Microservices. The role also involves publishing research, creating technical documentation, communicating findings, and mentoring team members.
Top Skills: SparkDaskGpu ClustersNvidia Inference Microservices (Nims)OcrPythonPyTorchRay
19 Days AgoSaved
Remote or Hybrid
4 Locations
272K-431K Annually
Senior level
272K-431K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Lead NVIDIA’s RL post-training frameworks strategy and engineering ecosystem across distributed training, inference, rollout, evaluation, orchestration, and NVIDIA platforms. Build and manage globally distributed teams, prioritize upstream and internal investments, establish benchmarks and execution metrics, and drive reliable open-source integrations. Partner with research, product, hardware, CUDA, networking, and external communities to scale reinforcement learning workloads across GPUs and heterogeneous systems.
Top Skills: CudaCudnnDistributed SystemsDpoGrpoHigh-Performance ComputingKubernetesMegatron-CoreMilesMonarchNcclNemoNemo-AlignerNixlNsightNvidia GpusOpen-Source SoftwareOpenrlhfPpoRayReinforcement LearningReward ModelingRlhfSglangSkyrlSlimeSlurmTensorrt-LlmTorchtitanTransformer EngineVerl
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account