Написано Димом Сафиным, комментарии приветствуются.
Рекомендую почитать хэндбук яндекса.
1. RL problem statement. MDP formalism. Crossentropy method.
2. Model-based RL. Bellman equations. Policy iteration with dynamic programming.
3. Value-based RL. Model-free prediction (Monte-Carlo vs. Temporal Difference).
4. Value-based RL. Model-free control. Q-Learning, SARSA, EV-SARSA.
5. Approximate value-based methods. DQN.
6. Policy-based RL. REINFORCE (with loss derivation).
7. Actor-critic policy gradient. Baselines, Advantage, A2C.
8. Exploration strategies. Eps-greedy, UCB, Thompson sampling
9. Reinforcement learning for seq2seq. Self-critical Sequence Training.