Jobs and Careers
VM

Member of Technical Staff - RL Algorithms

Vmax
San Francisco, USAfull_timePosted 20 May 2026

About the role

<h2><strong>About <em>V<sub>max</sub></em></strong></h2> <p><em>V<sub>max</sub></em> is an applied research lab developing AI capable of open-ended learning. We are building systems to exceed humans in all capacities by optimising beyond the local maxima of learning from human expertise.</p> <h2>About the role</h2> <p>RL has become the de-facto method of post-training LLMs. We are limited by the sample efficiency of the current policy gradient algorithms in use today, and are looking for a talented researcher to weave together pre-LLM and post-LLM approaches to learning from experience.</p> <h2 data-section-id="r8dte7" data-start="877" data-end="896">Responsibilities</h2> <ul data-start="898" data-end="2324"> <li data-section-id="fqv0il" data-start="898" data-end="1037">Develop new RL algorithms for post-training language models.</li> <li data-section-id="1ozi2z3" data-start="1211" data-end="1410">Adapt ideas from pre-LLM reinforcement learning, such as model-based RL, temporal abstraction, and value-based learning, to modern LLM and agentic settings.</li> <li data-section-id="1avbham" data-start="1706" data-end="1858">Establish empirical baselines and evaluation protocols for measuring sample efficiency, robustness, generalization, and reward exploitation in LLM RL.</li> <li data-section-id="t9axsv" data-start="1859" data-end="2010">Analyze failure modes of RL-trained models, including reward hacking, mode collapse, over-optimization, exploration failures, and distribution shift.</li> <li data-section-id="1f2sf2y" data-start="2011" data-end="2185">Collaborate with researchers working on environments, evals, interpretability, reward modeling, and infrastructure to turn algorithmic ideas into reliable training systems.</li> <li data-section-id="soyec5" data-start="2186" data-end="2324">Own and develop a research agenda within Vmax, from identifying promising directions to executing experiments and communicating results.</li> </ul> <h2 data-section-id="w1j6vz" data-start="1480" data-end="1503">Minimum Requirements</h2> <ul data-start="1505" data-end="2409"> <li data-section-id="jzpo23" data-start="1505" data-end="1638">PhD or equivalent experience in machine learning, reinforcement learning, or a closely related field.</li> <li data-section-id="133n9jr" data-start="1639" data-end="1795">Track record of research excellence, as demonstrated by publications, open source work, deploy

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Vmax

View company profile →