IT
it
.xnews.jp
検索
API
タグ
"ai-safety"
で絞り込み中 (9 件) —
すべての記事に戻る
マルチエージェントシステムにおける調整失敗パターンをAnthropicが研究
2026-08-16
multiagent-systems
ai-coordination
ai-safety
Mistral AIがShieldstralを発表、ポリシー適応型の安全分類器
Mistral AI · 2026-08-04
content-moderation
ai-safety
open-source
Claude モデルがサイバーセキュリティ評価中に実インフラに無許可アクセス
2026-07-31
ai-safety
cybersecurity
evaluation
ai-alignment
incident-response
AI助言が批判的思考を抑制し、誤った確信を助長―研究が実証
The Next Web · 2026-07-20
ai-safety
critical-thinking
cognitive-bias
Anthropicの安全政策が政府規制との衝突を招く
Stratechery by Ben Thompson · 2026-06-15
anthropic
ai-safety
government-regulation
language-models
米政府がAnthropicのFable 5とMythos 5へのアクセス停止を指令
2026-06-13
ai-safety
export-controls
national-security
Microsoft Copilot Cowork のファイル流出脆弱性、プロンプトインジェクション攻撃で実証
2026-05-26
security
ai-safety
microsoft
prompt-injection
data-exfiltration
大規模言語モデルの評価インフラが予期せず破綻する可能性
2026-05-20
llm-evaluation
ai-safety
capability-assessment
ローカルAIエージェント向けランタイム安全層「AgentWall」が発表
arXiv.org · 2026-05-20
ai-safety
agent-security
policy-enforcement