Dragonfly Logo

Dragonfly

Senior Inference Optimization Engineer - Dragonfly Portfolio

Posted One Month Ago
Remote or Hybrid
Hiring Remotely in Greece
Senior level
Remote or Hybrid
Hiring Remotely in Greece
Senior level
Optimize large-model inference at scale: improve throughput, reduce latency and cost per token, build benchmarking harnesses, tune parallelism and quantization strategies, implement load-balancing in routing, and evaluate custom kernels and emerging inference hardware.
The summary above was generated by AI
Dragonfly is a crypto-native Venture Capital and research firm with $3.6B+ in assets under management and 160+ portfolio companies. Our Talent team connects people with roles across our portfolio, opening the door to opportunities through our Talent Network.

This is an application to join our talent network. This is not a listing for an internal role at Dragonfly.

We're actively sourcing for a Senior Inference Optimization Engineer for one of our portfolio companies building privacy-first consumer AI infrastructure. You'll be on the bleeding edge of LLM inference performance, pushing throughput, driving down latency, and optimizing cost per token at significant scale.

Location: Remote, USA (open to excellent candidates outside the USA)

What We’re Looking For:
  • 5+ years in performance optimization or HPC with deep GPU architecture and parallel programming knowledge
  • Hands-on experience with at least one production LLM inference engine (vLLM, SGLang) running at high volume
  • Demonstrated experience with LLM inference optimization: continuous batching, PagedAttention, KV cache management, speculative decoding, quantization, CUDA graphs, torch.compile
  • Experience with distributed inference strategies: tensor parallelism, pipeline parallelism, MoE parallelism in multi-GPU and multi-node environments
  • GPU profiling fluency: Nsight Systems, Nsight Compute, PyTorch Profiler
  • Proficiency in Python, Rust, or Go. C++/CUDA a strong plus
  • Bonus: custom Triton kernels, diffusion/image model inference optimization, open-source inference framework contributions

About the role:
  • Stand up and optimize GPU infrastructure including B300 nodes in owned data centers
  • Drive down TTFT and TPOT, push throughput, and improve cost per token for LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines to identify optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
  • Optimize multivariate inference load-balancing algorithms within the inference routing system
  • Evaluate emerging inference optimization techniques including custom CUDA/Triton kernels, novel attention variants, new quantization schemes, and compilation stack improvements
  • Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in the stack

Even if you don't match every point above but are an engineer passionate about AI and/or crypto, we encourage you to apply. There may be other opportunities that fit your skill set.

Process: 
  • We'll review your application and assess fit for this role.
  • If there's a match, we'll facilitate a warm introduction to the team.
  • If the timing isn't right, we'll keep you in mind for future opportunities across the portfolio.

Compensation
  • Early Career: $110,000–$150,000
  • Mid: $150,000–$200,000
  • Senior: $180,000–$250,000
  • Staff/Leadership: $230,000–$330,000

The compensation range reflects variation by seniority, location, and hiring company. Final compensation is confirmed with the specific hiring company.

Submit your information below, and we’ll reach out if there’s a potential fit.

Similar Jobs

7 Minutes Ago
Remote or Hybrid
231K-385K Annually
Senior level
231K-385K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Leads Pfizer’s strategic technology initiatives across eight Digital Core Platforms, aligning enterprise architecture, AI programs, reuse, governance, and delivery. Oversees Senior Technical Architects, sets portfolio priorities, resolves cross-program trade-offs, advises executive stakeholders, and develops senior technical talent. Establishes responsible AI-enabled delivery practices and ensures reliability, security, documentation, operational excellence, and measurable outcomes across a complex global technology portfolio.
Top Skills: Artificial IntelligenceAutomationCloud PlatformsCybersecurityData IntegrationEnterprise ArchitectureModern Engineering PracticesReliability Engineering
7 Minutes Ago
Remote or Hybrid
208K-385K Annually
Senior level
208K-385K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Leads business technology strategy across Pfizer divisions, aligning divisional programs and projects with enterprise digital platforms, architecture standards, security, and governance. Serves as the primary business-Digital liaison, chairs or co-leads architecture reviews, resolves cross-division technology trade-offs, and oversees Director-level technical specialists. The role establishes technology direction, promotes platform reuse, assures delivery and operational excellence, advises senior leaders, and develops technical talent in a global, regulated enterprise.
Top Skills: Artificial IntelligenceCloud PlatformsCybersecurityData IntegrationDigital PlatformsEnterprise ArchitectureTechnology Governance
5 Hours Ago
Remote or Hybrid
4K-4K Annually
Senior level
4K-4K Annually
Senior level
Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Manages strategic, transformational, and cross-functional projects from planning through go-live. Responsibilities include coordinating teams and vendors, managing budgets, timelines, quality, testing, resources, risks, reporting, governance, and operational readiness. The role also mentors project managers, leads stakeholder meetings, and drives consistent project delivery across multiple regions.
Top Skills: MS Office

What you need to know about the Seattle Tech Scene

Home to tech titans like Microsoft and Amazon, Seattle punches far above its weight in innovation. But its surrounding mountains, sprinkled with world-famous hiking trails and climbing routes, make the city a destination for outdoorsy types as well. Established as a logging town before shifting to shipbuilding and logistics, the Emerald City is now known for its contributions to aerospace, software, biotech and cloud computing. And its status as a thriving tech ecosystem is attracting out-of-town companies looking to establish new tech and engineering hubs.

Key Facts About Seattle Tech

  • Number of Tech Workers: 287,000; 13% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Amazon, Microsoft, Meta, Google
  • Key Industries: Artificial intelligence, cloud computing, software, biotechnology, game development
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Madrona, Fuse, Tola, Maveron
  • Research Centers and Universities: University of Washington, Seattle University, Seattle Pacific University, Allen Institute for Brain Science, Bill & Melinda Gates Foundation, Seattle Children’s Research Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account