AI Inference Kernel Optimization Engineer
Engineering5+ yearsRemote · San Francisco, CA · Bangalore, India
Run BiOS is a capacity business behind an API: every fraction of utilization we win in the serving stack is room to lower the price of a token. As an AI Inference Kernel Optimization Engineer you work at the layer where that is decided — the compute kernels, memory paths, and serving runtimes that execute open models on our GPUs. Your work shows up in two places customers feel: latency, and the price list.
What you'll do
- 01Write and tune custom compute kernels — attention, matrix multiplication, normalization — in CUDA, Triton, and CUTLASS where the serving stack leaves performance on the table.
- 02Optimize memory movement: cache reuse, KV-cache management, and operator fusion, so kernels are bound by math, not by bandwidth.
- 03Implement low-precision inference — INT4, FP8, INT8 quantization paths — matched to the hardware we actually run.
- 04Profile at the hardware level with Nsight Systems, Compute Sanitizer, and PTX/SASS inspection, and turn what you find into shipped improvements.
- 05Work inside the serving runtimes we deploy — vLLM, TensorRT-LLM — and upstream what belongs upstream.
- 06Partner with research and platform engineering so new model architectures land on kernels that can actually serve them.
What you bring
- At least 5 years in systems software or ML systems engineering, with real time on GPU performance work.
- Deep proficiency in C++ and Python, and mastery of at least one low-level parallel programming framework: CUDA, Triton, or CUTLASS.
- Working fluency with hardware-level profiling tools and the ability to read assembly-level execution (PTX/SASS) when the profiler is not enough.
- Familiarity with the modern serving stack — vLLM, TensorRT-LLM, or a comparable runtime — and with how transformer inference actually spends its time and memory.
- Evidence, not adjectives: kernels you wrote or tuned, with the before-and-after numbers.
Nice to have
- Experience with deep learning compiler stacks: MLIR, LLVM, or torch.compile internals.
- Contributions to vLLM, TensorRT-LLM, Triton, or another inference-relevant open-source project we can read.
- KV-cache and paged-attention internals.
- Hardware-software co-design exposure — working with compiler or silicon teams on what next-generation hardware should support.
- An advanced degree in CS, CE, or EE is welcome; equivalent depth built in production is equally valued.