<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title>强化学习 - 标签 - 研发日志 · R&amp;D Log</title><link>https://rd163.visword.com/tags/%E5%BC%BA%E5%8C%96%E5%AD%A6%E4%B9%A0/</link><description>强化学习 - 标签 - 研发日志 · R&amp;D Log</description><generator>Hugo -- gohugo.io</generator><language>zh-CN</language><managingEditor>whutluohui@gmail.com (小智晖)</managingEditor><webMaster>whutluohui@gmail.com (小智晖)</webMaster><copyright>本作品采用知识共享署名-非商业性使用 4.0 国际许可协议进行许可。</copyright><lastBuildDate>Fri, 07 Feb 2025 00:00:00 +0800</lastBuildDate><atom:link href="https://rd163.visword.com/tags/%E5%BC%BA%E5%8C%96%E5%AD%A6%E4%B9%A0/" rel="self" type="application/rss+xml"/><item><title>DeepSeek 模型发展时间线</title><link>https://rd163.visword.com/posts/deepseek-models-development-timeline/</link><pubDate>Fri, 07 Feb 2025 00:00:00 +0800</pubDate><author><name>小智晖</name></author><guid>https://rd163.visword.com/posts/deepseek-models-development-timeline/</guid><description><![CDATA[<h2 id="2023-年" class="headerLink">
    <a href="#2023-%e5%b9%b4" class="header-mark"></a>2023 年</h2><p>2023 年 7 月，DeepSeek（深度求索）在杭州成立，由幻方量化创始人梁文锋创立，专注于通用人工智能（AGI）与大模型研发，并依托幻方积累的算力资源开展训练。</p>]]></description></item><item><title>游戏强化训练</title><link>https://rd163.visword.com/posts/game-rl-trainning/</link><pubDate>Sun, 12 Jan 2025 00:00:00 +0800</pubDate><author><name>小智晖</name></author><guid>https://rd163.visword.com/posts/game-rl-trainning/</guid><description>&lt;p>游戏是强化学习(Reinforcement Learning, RL)最经典也最成熟的实验场。规则清晰、状态可观测、奖励可量化、能够无限次对弈——这些特性让棋类、电子竞技、Atari 游戏成为验证 RL 算法的天然沙盒。本文梳理「游戏强化训练」涉及的核心概念、关键里程碑、主流框架,以及代表性开源复现项目。&lt;/p></description></item></channel></rss>