← All open roles

AI Inference Kernel Optimization Engineer

Engineering5+ yearsRemote

Run BiOS serves enterprise AI at production scale through an optimized inference engine. As an AI Inference Kernel Optimization Engineer, you improve the compute kernels, memory paths and serving runtimes that determine throughput, latency, reliability and token cost.

What you'll do

  • 01Write and tune custom compute kernels (attention, matrix multiplication, normalization) in CUDA, Triton and CUTLASS where the serving stack leaves performance on the table.
  • 02Optimize memory movement: cache reuse, KV-cache management and operator fusion, so kernels are bound by math, not by bandwidth.
  • 03Implement low-precision inference (INT4, FP8, INT8 quantization paths) matched to the hardware we actually run.
  • 04Profile at the hardware level with Nsight Systems, Compute Sanitizer and PTX/SASS inspection and turn what you find into shipped improvements.
  • 05Work inside the serving runtimes we deploy (vLLM, TensorRT-LLM) and upstream what belongs upstream.
  • 06Partner with research and platform engineering so new model architectures land on kernels that can actually serve them.

What you bring

  • At least 5 years in systems software or ML systems engineering, with real time on GPU performance work.
  • Deep proficiency in C++ and Python, and mastery of at least one low-level parallel programming framework: CUDA, Triton or CUTLASS.
  • Working fluency with hardware-level profiling tools and the ability to read assembly-level execution (PTX/SASS) when the profiler is not enough.
  • Familiarity with the modern serving stack (vLLM, TensorRT-LLM or a comparable runtime) and with how transformer inference actually spends its time and memory.
  • Evidence, not adjectives: kernels you wrote or tuned, with the before-and-after numbers.

Nice to have

  • Experience with deep learning compiler stacks: MLIR, LLVM or torch.compile internals.
  • Contributions to vLLM, TensorRT-LLM, Triton or another inference-relevant open-source project we can read.
  • KV-cache and paged-attention internals.
  • Hardware-software co-design exposure: working with compiler or silicon teams on what next-generation hardware should support.
  • An advanced degree in CS, CE or EE is welcome. Equivalent depth built in production is equally valued.

Apply for this role

Five minutes: your resume, your links and a paragraph about why this one. We read every application.

← All open roles