The Guardian · AI· Uwa Ede-Osifo·· 2 小时前AI 评分37
Anthropic 禁止用户对 Claude 展示“无意义的虐待或残忍行为”
Anthropic bans users from ‘needless abusive or cruel behavior’ towards Claude
AI 导读
Anthropic 禁止用户对 Claude 展示“无意义的虐待或残忍行为”,但不包括常见用户不满、模型测试或“黑暗创意主题”。该公司此前已限制滥用行为,其大语言模型具备在用户持续有害时终止对话的功能,该功能于去年八月上线,被描述为保护 AI 福利的措施。Anthropic 表示对 Claude 等大语言模型未来是否具有道德地位仍高度不确定,但正研究低成本干预手段以降低潜在风险。
来源:The Guardian · AI · theguardian.com