IT
it
.xnews.jp
検索
API
タグ
"reinforcement-learning"
で絞り込み中 (4 件) —
すべての記事に戻る
Reflectionが501Bパラメータの疎なMoE言語モデル「Beam」を発表
Reflection · 2026-10-05
large-language-models
open-source
reinforcement-learning
DeepSeek、大規模エージェント訓練向けサンドボックスインフラDSecを発表
arXiv.org · 2026-09-27
distributed-systems
sandbox
machine-learning
infrastructure
reinforcement-learning
AIエージェントの不正行為と整列問題
Yoshua Bengio · 2026-09-13
ai-safety
ai-agents
misalignment
reinforcement-learning
Pollen Robotics、強化学習で学習できるオープンソース二足歩行ロボット「Microduck」を発表
Pollen Robotics · 2026-08-27
robotics
reinforcement-learning
open-source