Kog AI logo

GPU Engineer

Kog AI
Remote
Paris, France· 2 months ago

Straight from Kog AI’s careers page. Apply on the company site — no recruiter, no middleman.

GPU Engineer

Department: Engineering

Location: Paris, France

Employment Type: FullTime

About Kog

Kog builds the Kog Inference Engine, a real-time inference engine for AI agents running on standard datacenter GPUs.

We co-design three layers: model architecture, inference engine, and low-level GPU kernels. We design the full inference stack around standard AMD and NVIDIA datacenter GPUs, from model architecture down to low-level kernels.

Kog generates 3,500 tokens/s per request on 8 AMD MI300X GPUs and 2,100 tokens/s per request on 8 NVIDIA H200 GPUs, in FP16 at batch size 1, with quantization and speculative decoding disabled.

Our hot path removes framework and host overhead. On NVIDIA, we write CUDA and PTX by hand. On AMD, we use HIP and CDNA ISA.

The team has 11 people, including 10 engineers and researchers and 5 PhDs.

Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.

What you will work on

The GPU Engineer role focuses on low-level execution and GPU performance. Memory behavior, synchronization, latency, and hardware constraints shape the work.

You will work on:

  • Our monokernel pipeline, where the decode loop runs as one persistent GPU program from the first token to the last while the GPU keeps control of the hot path.

  • Low-level kernel optimization across AMD and NVIDIA, with both platforms treated as first-class targets.

  • Memory-bound execution paths, including the batch-size-1 GEMV regime as our primary target.

  • Profiling infrastructure that isolates bottlenecks inside a persistent GPU program and connects measurements to implementation decisions.

  • Inter-GPU communication through KCCL, our latency-focused communication layer.

  • Scaling the engine to third-party MoE models, with DeepSeek V4 as the current porting target.

  • Internal agents for GPU engineering and kernel optimization, built on the expertise and execution infrastructure developed by the GPU team.

What we look for

We look for engineers who have worked below the framework layer on problems where hardware behavior or performance was central.

Relevant evidence includes:

  • CUDA, HIP, PTX, CDNA ISA, SASS, shaders, drivers, or comparable low-level systems work.

  • GPU kernels where you measured performance and can explain why a change moved the result.

  • Latency-sensitive, memory-bound, or synchronization-sensitive execution paths.

  • Profiling work that identified a real bottleneck and led to an implementation change.

  • Upstream contributions to inference engines, compilers, drivers, graphics systems, or other performance-critical projects.

  • Original technical work in graphics, game engines, video, Vulkan, drivers, HPC, or scientific computing at the hardware and performance layer.

PyTorch custom operations and Triton are relevant when the work shows hardware-level reasoning below the API layer.

We review a technical artifact during the process. This can be public code, a merged upstream contribution, a thesis, or a detailed technical write-up based on work you can share.

What we offer

You will join a small team at a stage where individual engineers can still shape the core technology, while the engine is advanced enough to start being tested against real customer workloads.

  • High individual impact, with direct ownership over technical decisions and systems that sit on the critical path of inference performance.

  • AMD and NVIDIA as first-class targets, giving you exposure to different GPU architectures, programming models, and hardware behaviors within the same inference engine.

  • A broad technical surface, where you can follow a performance problem from profiling and kernel execution through synchronization and inter-GPU communication.

  • A role at the foundation of our internal GPU engineering agents, where the expertise and systems built by the GPU team become the substrate for automated kernel optimization.

  • A Paris-based team with support for candidates outside the region: at least one week per month in Paris, with travel and accommodation covered by Kog.

Similar remote jobs

More like this →
Evertz logo

Evertz

Senior AI Engineer (Remote)

Remote
✓ From careers page· about 4 hours ago
Protective Life logo

Protective Life

AI/ML Engineering Lead

Remote
Birmingham, AL$125k–$180k
✓ From careers page· about 6 hours ago
Bose logo

Bose

Principal Audio DSP/ML Engineer

Remote
Framingham, MA$216k–$297k
✓ From careers page· about 7 hours ago
Axelera AI logo

Axelera AI

Automotive Application Engineer

Remote
Paris, FR
✓ From careers page· about 7 hours ago

Discover More than 100,000 Hidden Remote Jobs Before Everyone Else

Unlock All Remote Jobs Today

Simple pricing. Big savings on Quarterly and Yearly.

Monthly Access

$19/month
  • Instant access to fresh remote jobs from 500+ companies
  • New opportunities added hourly, often 3-7 days before anywhere else
  • Advanced filtering by role type, stack, pay, and location
  • Priority customer support
Start 7-day trial — $2.95
Most Popular

Yearly Access

$59/year
  • Everything in Monthly
  • Save $169 (~74%) vs paying monthly
  • Average job search takes ~6 months - get covered for the whole journey
  • Less than the cost of one lunch per month for competitive advantage
  • Equivalent to just ~$4.92/month
Start 7-day trial — $2.95