Bright Vision Technologies Logo

Bright Vision Technologies

ML Infrastructure Engineer

Posted 2 Days Ago
Be an Early Applicant
In-Office
Bellevue, WA, USA
100K-150K Annually
Senior level
In-Office
Bellevue, WA, USA
100K-150K Annually
Senior level
Designs and operates large-scale AI infrastructure for GPU-based training and inference. Responsibilities include distributed training platforms, scheduling, storage and networking performance, observability, fault tolerance, security, automation, cost optimization, and developer tooling. Partners with ML research teams on capacity planning and maintains operational documentation. Requires expertise in GPU clusters, distributed systems, Linux, networking, cloud ML services, and software engineering, with proficiency in Python and Go or C++.
The summary above was generated by AI
ML Infrastructure Engineer - Remote 
 
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. 
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. 
 
Job Title: ML Infrastructure Engineer
Location: 100% Remote (U.S.) 
Position Type: Full-time, Direct W2 
Salary Range: $100,000–$150,000 Annually 
Experience Required: 6+ years 
 
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position. 
 
Job Summary 
We are seeking an AI Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role focuses on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control. The ideal candidate has built or operated production AI infrastructure at scale, understands the interaction between hardware, kernel, scheduler, and ML framework, and brings strong software engineering discipline to platform work. 
Key Responsibilities 
  • Design and operate GPU and accelerator infrastructure for training and inference, spanning on-prem clusters, cloud-managed services, and hybrid configurations. 
  • Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams. 
  • Integrate frameworks such as PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering. 
  • Operate high-performance storage systems and data pipelines that keep accelerators fed with training data at near-line-rate. 
  • Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication. 
  • Build observability for AI workloads including utilization, throughput, training stability, and failure-mode analytics. 
  • Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs at scale. 
  • Drive cost optimization across compute, storage, and networking through scheduling, spot capacity, and right-sizing. 
  • Develop developer tooling and paved-road workflows that let researchers launch experiments safely and efficiently. 
  • Partner with research and applied ML teams to plan capacity for upcoming training runs. 
  • Implement security controls, isolation, and access management for multi-tenant AI infrastructure. 
  • Drive automation across cluster provisioning, lifecycle management, and configuration enforcement. 
  • Maintain runbooks, capacity dashboards, and operational documentation for the AI platform. 
  • Stay current with AI infrastructure research, accelerator hardware, and emerging open-source AI tooling. 
Required Qualifications 
  • Bachelor’s or Master’s degree in Computer Science or a related field. 
  • Six or more years of experience in infrastructure, platform, or HPC engineering. 
  • Hands-on experience operating GPU clusters or large-scale ML training infrastructure. 
  • Strong proficiency in Python and at least one systems language such as Go or C++. 
  • Deep understanding of distributed training, accelerator architectures, and collective communication. 
  • Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads. 
  • Strong understanding of Linux internals, networking, and high-performance storage. 
  • Experience with at least one major cloud provider’s ML infrastructure offerings. 
  • Strong software engineering practices including testing, CI/CD, and code review. 
  • Excellent communication and cross-functional collaboration skills. 
Preferred Qualifications 
  • Experience operating InfiniBand or RDMA networking at scale. 
  • Contributions to open-source ML infrastructure projects. 
  • Familiarity with custom orchestrators or research-grade training stacks. 
  • Exposure to frontier model training operations. 
  • Experience with FinOps for AI workloads. 
How to Apply 
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3544. Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.
 

Similar Jobs

11 Days Ago
Hybrid
Bellevue, WA, USA
133K-235K Annually
Junior
133K-235K Annually
Junior
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Design, build, and optimize scalable ML infrastructure for training, evaluation, and high-performance inference. Develop feature generation/serving pipelines, data management and labeling systems, and collaborate with ML engineers to deploy production models while ensuring reliability, scalability, and production-quality code.
Top Skills: AdkCaffe2FlinkFrontier LlmGoJavaLangfusePythonPyTorchRayScikit-LearnSparkSpark MlTensorFlow
10 Days Ago
Hybrid
Bellevue, WA, USA
200K-261K Annually
Senior level
200K-261K Annually
Senior level
AdTech • Artificial Intelligence • Gaming • Machine Learning • Software • Virtual Reality • Metaverse
Design, build, and operate large-scale offline data infrastructure and pipelines for ML training and experimentation. Improve reproducibility, observability, performance, and cost-efficiency across batch and streaming workloads while integrating orchestration systems and partnering with ML engineers.
Top Skills: AirflowData LakeData WarehouseFlinkFlytePythonRaySparkStreaming Platforms
11 Days Ago
Remote or Hybrid
Seattle, WA, USA
187K-243K Annually
Senior level
187K-243K Annually
Senior level
AdTech • Artificial Intelligence • Gaming • Machine Learning • Software • Virtual Reality • Metaverse
Design, build, and operate large-scale online ML inference infrastructure to serve production models with low latency and high reliability. Optimize inference performance, enable distributed training workflows, integrate ML pipelines with orchestration systems, improve observability, and support safe model rollout, canary testing, and automated rollback. Lead architectural improvements and collaborate with ML engineers and product teams to scale, monitor, and streamline model deployment and iteration.
Top Skills: AirflowAutoscalingCachingDynamic BatchingFlyteGkeGpu AccelerationGpu Kernel OptimizationKubernetesModel CompilationNvidia Triton Inference ServerObservabilityPythonPyTorchQuantizationRayRay DataRay ServeRay TrainTensorflow ServingTorchserve

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account