题目 Question
JXUFE-THU-B-PRG-003|实现一次Q学习更新。输入old_q、reward、max_next为数值,alpha和gamma在0到1之间;返回target=reward+gamma*max
Subscribe to view the full content.
内容信息
- 内容类型
- 题目 / Question
- 题型
- OJ_SCRIPT
- 难度
- NORMAL
- 更新时间
- 2026/07/23 19:37
Subscribe to view the full content.