Написано Димом Сафиным, комментарии приветствуются. Рекомендую почитать хэндбук яндекса.

1. RL problem statement. MDP formalism. Crossentropy method.

2. Model-based RL. Bellman equations. Policy iteration with dynamic programming.

3. Value-based RL. Model-free prediction (Monte-Carlo vs. Temporal Difference).

4. Value-based RL. Model-free control. Q-Learning, SARSA, EV-SARSA.

5. Approximate value-based methods. DQN.

6. Policy-based RL. REINFORCE (with loss derivation).

7. Actor-critic policy gradient. Baselines, Advantage, A2C.

8. Exploration strategies. Eps-greedy, UCB, Thompson sampling

9. Reinforcement learning for seq2seq. Self-critical Sequence Training.