Reinforcement learning from scratch: bandits, dynamic programming, Monte Carlo and TD methods (2018).