Would also be interesting to see more treatises on tranformer(-like) forecasting. Some discussion here: https://www.reddit.com/r/MachineLearning/comments/102mf6v/d_...