Yikai Wang and two co-authors submitted a paper titled "Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback" to arXiv on April 30, 2026, with two subsequent revisions through July 9, 2026. The paper is classified under cs.LG and addresses distributionally robust approaches to RLHF using Wasserstein metrics.
No score is assigned. Sources and their independence are shown in the citation chain below.