XiangSiqi’s Blog
文章
56
分类
3
标签
14
学习笔记
Reinforcement Learning
发布于: 2025-12-18
最后更新: 2026-4-30
次查看
强化学习
目录
0%
Chapter 1 Basic Concepts in Reinforcement Learning
Chapter 2: Bellman Equation
Motivating Examples
State Value
Derivation of Bellman Equation
Matrix-vector Form of Bellman Equation
Action Value
Chapter 3 Bellman Optimality Equation
Motivating Example
Optimal Policy
Contraction Mapping Theorem
Bellman Optimality Equation
Iteration Algorithm
Chapter 4 Monte Carlo Learning
MC-Base RL Algorithm
MC with Exploring Starts
MC Epsilon Greedy
Chapter 5 Stochastic Approximation
Motivation Examples
Robbins-Monro Algorithm
Chapter 6 Temporal Difference Learning
TD Algorithm of State Values
Derivation of TD Algorithm
TD Algorithm of Action Values
Q-Learning
Chapter 7 Value Function Approximation
State Value Estimation
Saras and Q-Learning with Function Approximation
Deep Q-Learning Network
Chapter 8 Policy Gradient Methods
Objective Function
REINFORCE
Actor Critic Methods
Simpest AC
Advantage actor-critic
Importance Sampling
Deterministic AC
向思齐
JianXian
文章
56
分类
3
标签
14
最新发布
CS336: Basics
2026-9-22
Chapter 5 Transformer
2026-8-22
Phase Transition Theory
2026-7-20
Numerical Renormalization Group
2026-7-18
Renormalization Group
2026-7-15
Chapter 4 循环神经网络
2026-5-12
公告
目录
0%
Chapter 1 Basic Concepts in Reinforcement Learning
Chapter 2: Bellman Equation
Motivating Examples
State Value
Derivation of Bellman Equation
Matrix-vector Form of Bellman Equation
Action Value
Chapter 3 Bellman Optimality Equation
Motivating Example
Optimal Policy
Contraction Mapping Theorem
Bellman Optimality Equation
Iteration Algorithm
Chapter 4 Monte Carlo Learning
MC-Base RL Algorithm
MC with Exploring Starts
MC Epsilon Greedy
Chapter 5 Stochastic Approximation
Motivation Examples
Robbins-Monro Algorithm
Chapter 6 Temporal Difference Learning
TD Algorithm of State Values
Derivation of TD Algorithm
TD Algorithm of Action Values
Q-Learning
Chapter 7 Value Function Approximation
State Value Estimation
Saras and Q-Learning with Function Approximation
Deep Q-Learning Network
Chapter 8 Policy Gradient Methods
Objective Function
REINFORCE
Actor Critic Methods
Simpest AC
Advantage actor-critic
Importance Sampling
Deterministic AC