Distributionally Robust Policy Gradient for Offline Contextual Bandits

Yang, Zhouhao; Guo, Yihong; Xu, Pan; Liu, Anqi; Anandkumar, Animashree

Distributionally Robust Policy Gradient for Offline Contextual Bandits

Zhouhao Yang, Yihong Guo, Pan Xu, Anqi Liu, Animashree Anandkumar

AISTATS 2023 pp. 6443-6462

/aistats/2023/yang2023aistats-distributionally/

Abstract

Learning an optimal policy from offline data is notoriously challenging, which requires the evaluation of the learning policy using data pre-collected from a static logging policy. We study the policy optimization problem in offline contextual bandits using policy gradient methods. We employ a distributionally robust policy gradient method, DROPO, to account for the distributional shift between the static logging policy and the learning policy in policy gradient. Our approach conservatively estimates the conditional reward distributional and updates the policy accordingly. We show that our algorithm converges to a stationary point with rate $O(1/T)$, where $T$ is the number of time steps. We conduct experiments on real-world datasets under various scenarios of logging policies to compare our proposed algorithm with baseline methods in offline contextual bandits. We also propose a variant of our algorithm, DROPO-exp, to further improve the performance when a limited amount of online interaction is allowed. Our results demonstrate the effectiveness and robustness of the proposed algorithms, especially under heavily biased offline data.

PDF AISTATS Semantic Scholar

Cite

Text

Yang et al. "Distributionally Robust Policy Gradient for Offline Contextual Bandits." Artificial Intelligence and Statistics, 2023.

Markdown

[Yang et al. "Distributionally Robust Policy Gradient for Offline Contextual Bandits." Artificial Intelligence and Statistics, 2023.](https://mlanthology.org/aistats/2023/yang2023aistats-distributionally/)

BibTeX

@inproceedings{yang2023aistats-distributionally,
  title     = {{Distributionally Robust Policy Gradient for Offline Contextual Bandits}},
  author    = {Yang, Zhouhao and Guo, Yihong and Xu, Pan and Liu, Anqi and Anandkumar, Animashree},
  booktitle = {Artificial Intelligence and Statistics},
  year      = {2023},
  pages     = {6443-6462},
  volume    = {206},
  url       = {https://mlanthology.org/aistats/2023/yang2023aistats-distributionally/}
}