Dragonfly logo

Senior Inference Optimization Engineer

Dragonfly
Remote
Remote· about 2 hours ago

Straight from Dragonfly’s careers page. Apply on the company site — no recruiter, no middleman.

Senior Inference Optimization Engineer - Dragonfly Portfolio

Location: United States (Remote)

Department: Portfolio

Location Type: HYBRID

Employment Type: FULL_TIME

Dragonfly is a crypto-native Venture Capital and research firm with $3.6B+ in assets under management and 160+ portfolio companies. Our Talent team connects people with roles across our portfolio, opening the door to opportunities through our Talent Network.

This is an application to join our talent network. This is not a listing for an internal role at Dragonfly.

Were actively sourcing for a Senior Inference Optimization Engineer for one of our portfolio companies building privacy-first consumer AI infrastructure. Youll be on the bleeding edge of LLM inference performance, pushing throughput, driving down latency, and optimizing cost per token at significant scale.

Location: Remote, USA (open to excellent candidates outside the USA)

What We’re Looking For:
  • 5+ years in performance optimization or HPC with deep GPU architecture and parallel programming knowledge
  • Hands-on experience with at least one production LLM inference engine (vLLM, SGLang) running at high volume
  • Demonstrated experience with LLM inference optimization: continuous batching, PagedAttention, KV cache management, speculative decoding, quantization, CUDA graphs, torch.compile
  • Experience with distributed inference strategies: tensor parallelism, pipeline parallelism, MoE parallelism in multi-GPU and multi-node environments
  • GPU profiling fluency: Nsight Systems, Nsight Compute, PyTorch Profiler
  • Proficiency in Python, Rust, or Go. C++/CUDA a strong plus
  • Bonus: custom Triton kernels, diffusion/image model inference optimization, open-source inference framework contributions

About the role:
  • Stand up and optimize GPU infrastructure including B300 nodes in owned data centers
  • Drive down TTFT and TPOT, push throughput, and improve cost per token for LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines to identify optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
  • Optimize multivariate inference load-balancing algorithms within the inference routing system
  • Evaluate emerging inference optimization techniques including custom CUDA/Triton kernels, novel attention variants, new quantization schemes, and compilation stack improvements
  • Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in the stack

Even if you dont match every point above but are an engineer passionate about AI and/or crypto, we encourage you to apply. There may be other opportunities that fit your skill set.

Process: 
  • Well review your application and assess fit for this role.
  • If theres a match, well facilitate a warm introduction to the team.
  • If the timing isnt right, well keep you in mind for future opportunities across the portfolio.

Submit your information below, and we’ll reach out if there’s a potential fit.

Similar remote jobs

More like this →
FreedomPay logo

FreedomPay

Snowflake Data Engineer

Remote
✓ From careers page· 11 minutes ago
Cloudera logo

Cloudera

Forward Deployed AI Engineer

Remote
Dubai, AE
✓ From careers page· 41 minutes ago
Airbnb logo

Airbnb

Staff Machine Learning Engineer, Traffic Intelligence (Remote)

Remote
$212k–$265k
✓ From careers page· about 1 hour ago
Airbnb logo

Airbnb

Software Engineer, Secure Development Engineering

Remote
$162k–$186k
✓ From careers page· about 1 hour ago

Discover More than 100,000 Hidden Remote Jobs Before Everyone Else

Unlock All Remote Jobs Today

Simple pricing. Big savings on Quarterly and Yearly.

Monthly Access

$19/month
  • Instant access to fresh remote jobs from 500+ companies
  • New opportunities added hourly, often 3-7 days before anywhere else
  • Advanced filtering by role type, stack, pay, and location
  • Priority customer support
Start 7-day trial — $2.95
Most Popular

Yearly Access

$59/year
  • Everything in Monthly
  • Save $169 (~74%) vs paying monthly
  • Average job search takes ~6 months - get covered for the whole journey
  • Less than the cost of one lunch per month for competitive advantage
  • Equivalent to just ~$4.92/month
Start 7-day trial — $2.95