Build your online resume. Claim your username
LiveKit logo

Reinforcement Learning Research Engineer at LiveKit at LiveKit

Remote 🌍 Work from Anywhere Full time Senior Posted  Apply before Oct 17, 2026

Job Description

About LiveKit

LiveKit is at the forefront of building the essential infrastructure for the next generation of voice-driven computing. Our robust platform provides developers with all necessary tools for building, testing, deploying, scaling, and observing AI agents in production environments. Established in 2021, LiveKit proudly powers voice AI applications for industry leaders such as OpenAI, xAI, Salesforce, Coursera, and Spotify, along with thousands of other innovative companies, collectively handling billions of calls annually.

About This Role

We are seeking an exceptional engineer to specialize in post-training development at LiveKit. Our advanced agents operate across voice and increasingly text-based channels like SMS and chat. This role focuses on solving complex, long-horizon problems: ensuring agents remain effective across multiple sessions, managing context that evolves over time, and reliably utilizing tools within live conversations.

What You Will Do

  • Construct the environments and verifiers essential for our models' training processes.
  • Take ownership of the synthetic data pipeline, from initial generation through rigorous quality assurance gates.
  • Execute training experiments comprehensively, from start to finish, and clearly articulate the impact of model changes.
  • Develop and define the evaluation criteria that every release must successfully meet.
  • Select and adapt open-weight base models to suit our specific task requirements.
  • Ensure consistent trained behavior across both voice and text-based agents.
  • Deploy models into production and continuously enhance their performance based on real-world usage data.

Who You Are

  • A highly skilled Python engineer.
  • Have successfully guided a machine learning model from raw data input all the way to production deployment.
  • Possess a product-centric view of data, emphasizing coverage, diversity, and preventing leakage.
  • Anticipate that models will exploit weak rewards and proactively design robust solutions against such vulnerabilities.
  • Proficient with GPUs and possess a realistic understanding of their capabilities and limitations.
  • Understand when model training is necessary and when it is not.
  • Comfortable collaborating effectively in a remote work environment.

Nice to Have

  • Previous experience with post-training methodologies, including fine-tuning, reward design, or reinforcement learning (e.g., GRPO).
  • Familiarity with RL and fine-tuning frameworks such as TRL, verl, or OpenRLHF, or experience developing custom training loops.
  • Knowledge of fast rollout techniques like vLLM or SGLang, and multi-GPU training with FSDP.
  • Experience training tool-using or multi-turn agents.
  • Proficiency with execution sandboxes, verifiers, evaluation harnesses, or creating tools relied upon by other engineers.
  • Familiarity with open-weight model families such as Qwen or Llama, and techniques like LoRA.

Our Commitment to You

LiveKit is dedicated to providing an outstanding employee experience, including:

  • The unique chance to significantly influence the brand and direction of a rapidly expanding developer platform.
  • Opportunities for deep collaboration within a small, highly senior team that places a strong emphasis on craftsmanship and creative problem-solving.
  • A competitive compensation and equity package.
  • Comprehensive health, dental, and vision benefits.
  • A flexible vacation policy, promoting work life balance.

LiveKit is an equal opportunity employer committed to diversity and inclusion. We do not discriminate on the basis of any characteristic protected by applicable law. Should you require a reasonable accommodation during the application or interview process, please reach out to [email protected].

Ready to Apply?

Take the next step in your career journey.

Apply Now

You will be redirected to the company's application page

Link verified about 10 hours ago

💜 Please mention that you found the job on True Work From Home, this helps us grow. Thanks!