Dear TA,
Consider the following two cases:
- Training up to an episode (e.g. 600) when the average reward of the past 100 episodes is already 200.0
- Training up to the final episode (1000) when the average reward of the past 100 episodes might be 197.0
In general, which model should we use for testing?
Thank you very much!
Dear TA,
Consider the following two cases:
In general, which model should we use for testing?
Thank you very much!