← All open roles

Applied Machine Learning Engineer

Engineering5+ yearsRemote · San Francisco, CA · Bangalore, India

Customers come to Run BiOS to fine-tune open models on dedicated GPUs, billed per second, and to serve what they train through one API. As an Applied Machine Learning Engineer you make their models work: you take fine-tuning jobs from dataset to deployed endpoint, build the tooling that makes that path boring, and turn individual customer problems into platform capabilities.

What you'll do

  • 01Own fine-tuning engagements end to end: dataset inspection, method selection, run configuration, evaluation, and handoff.
  • 02Build internal tooling for dataset validation, training orchestration, and model evaluation across a catalog of 250k+ open models.
  • 03Debug the weird ones: loss spikes, adapter pathologies, tokenization edge cases, VRAM ceilings.
  • 04Feed recurring customer patterns back to product and research as concrete, prioritized proposals.
  • 05Help customers serve their fine-tuned weights through our inference stack.

What you bring

  • At least 5 years applying machine learning in production environments.
  • Real experience fine-tuning LLMs or vision-language models — configuring runs, reading curves, not only calling APIs.
  • Strong Python and PyTorch, and comfortable reasoning about GPU memory when things get tight.
  • Data instincts: you look at the dataset before you touch the model.
  • Clear communication — you can explain a failed run to a customer without hand-waving.

Nice to have

  • Breadth across model families — DeepSeek, Qwen, Kimi, GLM, Gemma, and friends.
  • Depth in the Hugging Face ecosystem: transformers, PEFT, TRL, datasets.
  • Multi-GPU training frameworks and checkpoint management.
  • Serving experience — vLLM or similar inference stacks.

Apply for this role

Five minutes: your resume, your links, and a paragraph about why this one. We read every application.

← All open roles