官术网_书友最值得收藏!

Identifying episodes

We mentioned earlier that the agent explores the environment in numerous trials-and-errors before it can learn to maximize its goals. Each such trial from start to finish is called an episode. The start location may or may not always be from the same location. Likewise, the finish or end of the episode can be a happy or sad ending.

A happy, or good, ending can be when the agent accomplishes its pre-defined goal, which could be successfully navigating to a final destination for a mobile robot, or successfully picking up a peg and placing it in a hole for an industrial robot arm, and so on. Episodes can also have a sad ending, where the agent crashes into obstacles or gets trapped in a maze, unable to get out of it, and so on.

In many RL problems, an upper bound in the form of a fixed number of time steps is generally specified for terminating an episode, although in others, no such bound exists and the episode can last for a very long time, ending with the accomplishment of a goal or by crashing into obstacles or falling off a cliff, or something similar. The Voyager spacecraft was launched by NASA in 1977, and has traveled outside our solar system – this is an example of a system with an infinite time episode.

We will next find out what a reward function is and why we need to discount future rewards. This reward function is the key, as it is the signal for the agent to learn.

主站蜘蛛池模板: 济阳县| 获嘉县| 无锡市| 赤水市| 杂多县| 木兰县| 伊金霍洛旗| 宣恩县| 武宁县| 祁东县| 浮梁县| 乌审旗| 镇雄县| 皋兰县| 崇州市| 原阳县| 静安区| 汪清县| 澜沧| 沧州市| 昭苏县| 阳谷县| 礼泉县| 资兴市| 石楼县| 大竹县| 奉化市| 阿勒泰市| 石河子市| 宁都县| 仁布县| 镇赉县| 潮州市| 广平县| 建水县| 石台县| 嘉峪关市| 临泽县| 永安市| 平果县| 江城|