ServiceNow-AI researchers led by Rafael Pardinas and Ehsan Kamalloo migrated PipelineRL from vLLM V0 to vLLM V1, achieving parity with their reference run after fixing four backend issues. The team corrected processed rollout logprobs, V1-specific runtime defaults, the inflight weight-update path, and the fp32 lm_head final projection. They prioritized fixing backend correctness before modifying the RL objective, ensuring train-inference logprob consistency.
No score is assigned. Sources and their independence are shown in the citation chain below.