Policy gradients for CVaR-constrained MDPs

L. A. Prashanth

doi:10.1007/978-3-319-11662-4_12

Profiles Research Units Publications

Journal

Policy gradients for CVaR-constrained MDPs

Published in Springer Verlag

2014

DOI: 10.1007/978-3-319-11662-4_12

Volume: 8776

Pages: 155 - 169

Abstract

We study a risk-constrained version of the stochastic shortest path (SSP) problem, where the risk measure considered is Conditional Value-at-Risk (CVaR). We propose two algorithms that obtain a locally risk-optimal policy by employing four tools: stochastic approximation, mini batches, policy gradients and importance sampling. Both the algorithms incorporate a CVaR estimation procedure, along the lines of [3], which in turn is based on Rockafellar-Uryasev’s representation for CVaR and utilize the likelihood ratio principle for estimating the gradient of the sum of one cost function (objective of the SSP) and the gradient of the CVaR of the sum of another cost function (constraint of the SSP). The algorithms differ in the manner in which they approximate the CVaR estimates/necessary gradients - the first algorithm uses stochastic approximation, while the second employs mini-batches in the spirit of Monte Carlo methods. We establish asymptotic convergence of both the algorithms. Further, since estimating CVaR is related to rare-event simulation, we incorporate an importance sampling based variance reduction scheme into our proposed algorithms. © Springer International Publishing Switzerland 2014.

Topics: CVAR (69)%, Stochastic approximation (54)%, Importance sampling (53)% and Variance reduction (52)%

View more info for "Policy Gradients for CVaR-Constrained MDPs"

About the journal

Journal	Data powered by TypesetLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Publisher	Data powered by TypesetSpringer Verlag
ISSN	03029743
Open Access	No

Authors (1)

L. A. Prashanth
- Department of Computer Science and Engineering

ABOUT IIT MADRAS

R & D

RANKINGS & ACHIEVEMENTS

QUICK FIND