Bright Vision Technologies Logo

Bright Vision Technologies

ML Performance Engineer

Posted 2 Days Ago
Be an Early Applicant
In-Office
Redmond, WA, USA
100K-150K Annually
Senior level
In-Office
Redmond, WA, USA
100K-150K Annually
Senior level
Optimize large-scale AI training and inference systems for throughput, latency, and cost. Responsibilities include GPU and kernel optimization, distributed training, model compression, LLM serving, compiler optimization, data pipeline tuning, benchmarking, hardware evaluation, and performance regression analysis. The engineer will collaborate across ML and platform teams, document optimization practices, evaluate emerging technologies, and mentor junior engineers.
The summary above was generated by AI
ML Performance Engineer - Remote 
 
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. 
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. 
 
Job Title: ML Performance Engineer
Location: 100% Remote (U.S.) 
Position Type: Full-time, Direct W2 
Salary Range: $100,000–$150,000 Annually 
Experience Required: 6+ years 
 
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position. 
 
Job Summary 
We are seeking an AI Performance Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems. The role spans the full stack from low-level kernel optimization to distributed system tuning, requiring deep understanding of GPU architecture, model parallelism, memory management, and compiler-level optimization. The ideal candidate has demonstrated impact on production AI workloads, with strong instrumentation and measurement discipline that enables rigorous, data-driven optimization decisions. In this role you will work closely with cross-functional partners — product, design, engineering, operations, and business stakeholders — to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production. 
Key Responsibilities 
  • Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost. 
  • Identify and eliminate bottlenecks across data loading, model compute, communication, and memory. 
  • Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference. 
  • Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding. 
  • Tune attention implementations using FlashAttention, paged attention, and related techniques. 
  • Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving. 
  • Drive compiler-level optimizations using Triton, XLA, TorchInductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains. 
  • Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training. 
  • Build and maintain rigorous benchmark suites and regression frameworks across workloads. 
  • Collaborate with ML and platform engineering teams to embed best practices in standard pipelines. 
  • Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies. 
  • Evaluate new hardware and software offerings, and advise on adoption. 
  • Document performance tuning playbooks and share findings broadly across engineering teams. 
  • Stay current with AI systems research and translate advances into production improvements. 
Required Qualifications 
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field. 
  • Six or more years of experience in performance engineering, ML systems, or HPC. 
  • Strong proficiency in Python and C++. 
  • Hands-on experience optimizing deep learning workloads on modern GPUs. 
  • Deep understanding of distributed training and inference techniques. 
  • Experience with profiling tools across CPU, GPU, and distributed systems. 
  • Familiarity with model compression techniques and their accuracy implications. 
  • Strong grasp of memory hierarchies, communication primitives, and parallelism strategies. 
  • Excellent measurement, debugging, and analytical reasoning skills. 
  • Strong communication and collaboration skills. 
Preferred Qualifications 
  • Experience optimizing LLM inference at production scale. 
  • Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects. 
  • Familiarity with custom kernel authoring in Triton or CUTLASS. 
  • Experience with FinOps for AI workloads. 
  • Publications or talks on AI systems performance. 
How to Apply 
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3544. Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.
 

Similar Jobs

22 Days Ago
In-Office or Remote
Seattle, WA, USA
242K-389K Annually
Senior level
242K-389K Annually
Senior level
Artificial Intelligence • Machine Learning • Robotics • Software • Transportation • Design • Manufacturing
Lead ML Performance Optimization for autonomous driving, collaborating across teams, implementing optimization techniques, and mentoring engineers for career growth.
Top Skills: C++PythonPyTorchRay ServeTensorrt
32 Minutes Ago
Hybrid
Seattle, WA, USA
136K-234K Annually
Senior level
136K-234K Annually
Senior level
AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Design, build, and operate low-latency, high-throughput ad delivery systems for bidding, ranking, auctions, and budget enforcement. Integrate ML signals, collaborate with product and data teams, improve observability and reliability, lead design discussions, and mentor junior engineers.
Top Skills: A/B TestingApache FlinkSparkCloud EnvironmentsGrpcJavaKotlinMachine LearningMicroservices
An Hour Ago
Hybrid
Seattle, WA, USA
50K-94K Annually
Mid level
50K-94K Annually
Mid level
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Provides underwriting and processing support for Subcontract Default Insurance, including new business, renewals, cancellations, endorsements, broker setup, licensing, reconciliations, regulatory requirements, and reporting. Responds to internal and external customer inquiries, maintains files, supports technical issue resolution and user testing, and assists with billing discrepancies. Mentors colleagues, serves as a workflow subject-matter expert, and performs office support and special projects.
Top Skills: Bond Processing SystemsCorporate SystemsSdi Systems

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account