Research Engineer
Engineering5+ yearsRemote · San Francisco, CA · Bangalore, India
Run BiOS runs serverless inference across the major open model families and fine-tunes custom models on dedicated GPUs. As a Research Engineer you improve the models themselves — from bios-adaptive, our automatic router across the model pool, to the fine-tuning methods and alignment recipes customers run. You turn research into behavior that serves production traffic.
What you'll do
- 01Develop and evaluate routing strategies for bios-adaptive across the open-model pool, and measure what they do to quality, latency, and cost.
- 02Design fine-tuning and alignment recipes — SFT, LoRA/QLoRA, DPO-family objectives, continued pre-training — that behave predictably on dedicated-GPU infrastructure.
- 03Build evaluation harnesses that catch regressions before customers do.
- 04Read the literature, reproduce what matters, discard what does not, and ship what wins.
- 05Work with product engineering to make research outcomes visible and controllable in the product.
- 06Write down what you learn — internal notes, docs, and occasionally public posts.
What you bring
- At least 5 years in machine learning engineering or research, with meaningful time on large language models.
- Hands-on experience fine-tuning open-weight models: PyTorch, distributed training, parameter-efficient methods.
- Evaluation rigor: when you say a model got better, you can show how you know.
- Engineering fundamentals strong enough that someone else can reproduce your experiments.
Nice to have
- Publications or substantial open-source work in post-training, routing, or evaluation.
- Preference optimization beyond the basics — DPO variants, reward modeling, online RL.
- Inference-time optimization: quantization, speculative decoding, serving-stack internals.
- Mixture-of-experts familiarity.