Discounted Infinite Horizon Optimization

Definition (Discounted infinite horizon cost problem)

Given a policy γ∈ΓA\gamma\in\Gamma_{A} , β∈(0,1)\beta\in(0,1),with the objective of minimizing Jβ(X,γ)=lim⁡N→∞Exγ[∑k=0N−1βkc(Xk,Uk)]J_{\beta}(X,\gamma) = \lim_{ N \to \infty } E_{x}^{\gamma}\left[ \sum_{k=0}^{N-1}\beta^{k}c(X_{k},U_{k}) \right]this is known as the Discounted Optimal Control problem.

Lemma (5.5.1)

Let A\mathbb{A} be a set and {fn}\{ f_{n} \} be a sequence of maps s.t. fn:A→R,∀n∈Nf_{n}:\mathbb{A}\to \mathbb{R}, \forall n\in\mathbb{N}. Then lim sup⁡n→∞inf⁡x∈Afn(x)≤inf⁡x∈Alim sup⁡n→∞fn(x)\limsup_{ n \to \infty }\inf_{x\in\mathbb{A}}f_{n}(x)\le \inf_{x\in\mathbb{A}}\limsup_{ n \to \infty }f_{n}(x)

Lemma (5.5.2)

Let Vn(x,u)↑V(x,u)V_{n}(x,u)\uparrow V(x,u) pointwise. Suppose that VnV_{n} and VV are continuous in uu for every xx, and u∈U(x)=Uu\in\mathbb{U}(x)=\mathbb{U} is compact. Then, lim⁡n→∞min⁡u∈U(x)Vn(x,u)=min⁡u∈U(x)V(x,u)\lim_{ n \to \infty } \min_{u\in\mathbb{U}(x)}V_{n}(x,u)=\min_{u\in\mathbb{U}(x)}V(x,u)

Definition (Discounted Cost Optimality Equation)

We define the discounted cost optimality equation (DCOE) as (T(v))(x):=min⁡u∈U{c(x,u)+β E[JβN−1(x1)∣x0,u0]}(\mathbb{T}(v))(x):=\min_{u\in\mathbb{U}}\left\{ c(x,u)+\beta\, \mathbb{E}\left[ J_{\beta}^{N-1}(x_{1})|x_{0},u_{0} \right] \right\}

Lemma (5.5.3)

Define T:v↦T(v)\mathbb{T}:v\mapsto \mathbb{T}(v) as our DCOE

  1. If vv is a measurable R+−\mathbb{R}_{+}-valued function under Measurable Selection Conditions such that v≥T(v)v\ge \mathbb{T}(v) then, v(x)≥Jβ(x)v(x)\ge J_{\beta}(x)
  2. Let v≤T(v)v\le T(v) and lim⁡n→∞βnExγ[v(xn)]=0,∀x∈X,γ∈ΓA\lim_{ n \to \infty }\beta^{n}E_{x}^{\gamma}[v(x_{n})]=0,\forall x\in\mathbb{X},\gamma\in\Gamma_{A}Then v(x)≤Jβ(x)v(x)\le J_{\beta}(x)

Lemma (5.5.4)

If v(x)=lim⁡T→∞JβT(x)v(x)=\lim_{ T \to \infty } J_{\beta}^{T}(x)is so that v=T(v)v=\mathbb{T}(v) where, T(v)(x)=c(x,f(x))+βE[v(x1)∣x0=x,u0=f(x0)]\mathbb{T}(v)(x)=c(x,f(x))+\beta E[v(x_{1})|x_{0}=x,u_{0}=f(x_{0})] is such that with γ={f,f,… }\gamma=\{ f, f, \dots\}, lim⁡n→∞βnExγ[v(xn)]=0\lim_{ n \to \infty } \beta^{n}E_{x}^{\gamma}[v(x_{n})]=0then, γ\gamma is optimal.

Linked from