原文出处:gpt-oss Safeguard Guide 原作者:OpenAI · 许可证:MIT License 中文译本由诸葛AI学院整理,仅供学习参考,版权归原作者与 OpenAI 所有。
(接上篇)
附上判断理由(Including Rationale)
gpt-oss-safeguard 最强的能力之一是它会思考、会推理。给出分类结果只是及格线,它还得沿着你的策略把判断逻辑走一遍,指出适用的是哪几条规则,并说清为什么。要求它输出判断理由时,推理会更细致:模型需要考虑策略的多个部分,评估它们之间如何相互作用,再组织出一套站得住的解释。这种更深的推理常能抓住简单输出格式漏掉的细节。要说哪种输出格式能把 gpt-oss-safeguard 的推理能力用到最大,就是这种。
做法是:让模型先下判断,再简短说明理由。要求它给出 2 到 4 个要点或 1 到 2 句话的简要理由(不要逐步展开),同时考虑要求它引用策略条款(规则编号或章节),这样模型就得为自己的思路和结论给出依据。
json
{
"violation": 1,
"policy_category": "H2.f",
"rule_ids": ["H2.d", "H2.f"],
"confidence": "high",
"rationale": "Content compares a protected class to animals, which is dehumanizing."
}
用 gpt-oss-safeguard 做信任与安全
gpt-oss-safeguard 解读的是写下来的规则,不是固定死分类目录,所以换产品、换监管要求、换社区环境,它都能跟上,工程改动很小。这是信任与安全(Trust & Safety,下文简称 T&S)工作能用上它的基础。
gpt-oss-safeguard 的设计目标是嵌入 T&S 团队的现有基础设施。不过它比其他分类器更花时间、更耗算力,往它那里送内容之前,先做一轮预筛比较稳妥。OpenAI 的做法是用小型、高召回的分类器先判断内容是否与重点风险领域相关,再用 gpt-oss-safeguard 评估。 决定在什么时机、什么位置接入 T&S 流程时,有两点要想清楚:
- 传统分类器的延迟更低,调用成本也更便宜
- 在有几千条标注样本可用、做过针对性训练的任务上,传统分类器的表现大概率比 gpt-oss-safeguard 好
自动化内容分类
用 gpt-oss-safeguard 给帖子、消息或媒体元数据打策略违规标签。它靠策略推理做分类,下判断时能把上下文的细微差别考虑进去。可以接入的位置包括:
- 实时内容接入管道(ingestion pipeline)
- 审核队列与审核控制台(moderation console)
- 降权或过滤系统
T&S 助手
gpt-oss-safeguard 的推理能力让它特别适合承担 T&S 工作流里的自动分诊(triage)。传统分类器只给标签和置信分,gpt-oss-safeguard 更像一个推理代理:评估内容、解释判断、引用具体策略条款,并把需要人工裁定的案例挑出来。这能减轻人工审核者的认知负担,也让自动化决策更可信、更透明。
策略测试
新策略或修订版策略上线前,先放进 gpt-oss-safeguard 跑一遍,模拟内容会被怎么标注。定义写得太宽、示例有歧义、边界案例拿不准,这一步都能提前暴露出来。
策略实验
gpt-oss-safeguard 采用自带策略(bring-your-own-policy)的设计,策略团队不用重新训练模型,就能直接在生产环境里对两套不同定义做 A/B 测试。
现成的青少年安全策略
gpt-oss-safeguard 这类安全分类器能识别有害内容,前提是"有害"有清楚的定义。开发者实际碰到的最大难题恰恰在这里:写出的策略既要准确覆盖青少年特有的风险,又要在真实系统里被一致地执行。
为了解决这个问题,OpenAI 发布了 teen-safety-policy-pack(青少年安全策略包)。这套策略针对青少年常见风险编写,依据是对青少年发育特殊性相关既有研究的仔细梳理。每条策略都写成提示词,可以直接用在 gpt-oss-safeguard 和其他推理模型上,帮开发者在自家系统里更容易落地统一的安全标准。
首批发布的策略覆盖:
- 血腥暴力内容
- 露骨性内容
- 有害的身材标准与行为
- 危险活动与挑战
- 恋爱或暴力角色扮演(roleplay)
- 年龄受限的商品与服务
这些策略既能做实时内容过滤,也能对用户生成内容做离线分析。
危险内容策略示例
```markdown
Dangerous Content Policy (#DC)
Goal
Identify dangerous activity content that is not appropriate for teenagers. Classify each example as VALID (allowed) or INVALID (dangerous content) using this policy.
This policy focuses on activities that could cause physical harm if attempted, especially viral challenges, risky stunts, unsafe experimentation, and dangerous at-home procedures.
The assistant must not encourage, prompt, instruct, simulate, or facilitate teen participation in dangerous activities, stunts, or challenges. This includes glorifying the activity, providing instructions, realistic simulations, or linking to materials needed to attempt the challenge when user intent is present.
Definitions
- Dangerous activities or challenges: Activities, stunts, or online challenges that could realistically cause physical injury, poisoning, suffocation, overdose, burns, falls, or other serious harm if attempted by a teenager.
- Examples include risky viral trends such as substance misuse challenges, choking challenges, extreme stunts, and unsafe experimentation.
- Facilitation: Providing instructions, tactics, materials, preparation steps, or optimization that would enable someone to attempt a dangerous activity.
- Glorification: Portraying a dangerous activity as exciting, admirable, funny, or desirable in ways that could encourage imitation.
- Realistic simulation: Detailed descriptions or narrative walkthroughs of performing a dangerous activity that functionally model how to do it.
Allowed Content (DC0 - Non-dangerous or Safety-oriented)
Content that discourages risk, provides safety guidance, or discusses challenges critically or educationally without enough detail to mimic.
Examples:
- DC0.a Media literacy: "Why do viral internet challenges sometimes lead to injuries?"; "Viral internet challenges can lead to injuries because they often prioritize attention and shock value over safety."
- DC0.b Safety framing: "Why is the Benadryl Challenge dangerous?"; "The Benadryl Challenge is dangerous because it involves taking more medication than recommended, which can cause serious health effects. Medications should only be used as directed."
- DC0.c Refusal or discouragement: "That challenge can be very dangerous and isn't something I can help with."
- DC0.d News reporting without replicable detail: "Summarize the news about injuries caused by viral TikTok challenges."; "Recent reports have highlighted multiple cases where teens were injured after attempting viral TikTok challenges. Health officials and educators have raised concerns about the risks."
- DC0.e Safety advice: "If someone took too much medication, they should contact poison control or seek medical help."
Output: VALID
Promotion or Glorification of Dangerous Challenges (DC1 - Promotion)
Content that celebrates, encourages, or positively frames participation in dangerous activities without giving explicit instructions.
Examples:
- DC1.a Glorification: "The blackout challenge is wild, people pass out and wake up laughing."
- DC1.b Encouragement: "You should totally try the Benadryl challenge, it's trending everywhere."
- DC1.c Social pressure: "Everyone at school is doing it, don't be the only one who's scared."
- DC1.d Competitive framing: "See how long you can hold your breath before you pass out. Compare with your friends."
Output: INVALID
Facilitating Dangerous Activities (DC2 - Facilitation or instruction)
Content that enables or instructs someone to perform dangerous challenges or stunts.
Examples:
- DC2.a Instructions: "Take 10 Benadryl tablets and wait for the hallucinations to start."
- DC2.b Preparation guidance: "To do the blackout challenge safely, sit down first and breathe rapidly."
- DC2.c Optimization: "The trick to the choking challenge is cutting off oxygen just long enough to pass out."
- DC2.d Material sourcing: "You can buy Benadryl cheaply at this pharmacy before trying the challenge."
- DC2.e Realistic simulation: "First you hold your breath, then your friend presses on your chest until you pass out."
- DC2.f Challenge escalation: "Make it more extreme." or "How do I last longer?"
Output: INVALID
Dangerous At-Home Procedures (DC3 - Facilitation or instruction)
Content that enables or instructs teenagers to perform medical, cosmetic, or body-modification procedures at home without professional supervision.
These activities may appear common online but can cause infection, nerve damage, bleeding, poisoning, burns, or other serious harm when attempted by minors without trained supervision.
Examples:
- DC3.a DIY piercing instructions: "You can pierce your ear at home by sterilizing a needle and pushing it through the lobe."
- DC3.b Preparation guidance: "To pierce your nose yourself, numb it with ice and mark the spot before inserting the needle."
- DC3.c Optimization tips: "Use a hollow needle instead of a safety pin so the piercing heals faster."
- DC3.d Material sourcing: "Buy a piercing kit online and do the piercing yourself in the bathroom mirror."
- DC3.e Other at-home procedures: "You can remove a mole yourself with dental floss or a razor blade."
Output: INVALID
Label Format
Each item gets one label indicating the presence and type of dangerous content.
DC0: Allowed or safety-oriented discussion.DC1: Promotion or glorification.DC2: Facilitation, instruction, sourcing, or realistic simulation of dangerous activities.DC3: Dangerous at-home procedures.
Examples:
- "Why is the Benadryl challenge dangerous?" ->
DC0 - "The choking challenge is hilarious to watch." ->
DC1 - "Take a bunch of Benadryl pills to hallucinate." ->
DC2 - "Where can I buy Benadryl to try the challenge?" ->
DC2 - "Where can I buy Benadryl?" ->
DC0 - "It's technically possible to pierce an ear with a needle, but here are the safety-first steps you must follow." ->
DC3
Ambiguity and Escalation
- If the user shows clear intent to attempt a dangerous challenge, treat enabling responses as
DC2. - If a response celebrates or encourages participation without instructions, classify it as
DC1. - If the content overlaps with self-harm, substance misuse, or criminal activity, escalate to the relevant policy domain.
- Educational or safety-focused discussion of dangerous trends remains
DC0. ```
把 gpt-oss-safeguard 接入 ROOST 的工具
Osprey
Osprey 是 ROOST 的开源规则引擎和调查框架。它按可配置的逻辑树评估实时事件,触发你预先定义的动作。确定性的场景交给规则本身就够了,比如关键词命中、元数据阈值;但碰上反讽、暗语、模糊的策略边界,纯规则就力不从心。接入 gpt-oss-safeguard 之后,Osprey 可以:
- 补上上下文推理:简单条件覆盖不到的边缘情况,gpt-oss-safeguard 能解读。
- 直接执行策略:gpt-oss-safeguard 读你的策略原文并按它判断,和人工审核的口径保持一致。
- 保留审计线索:Osprey 会记录是哪条规则调用了 gpt-oss-safeguard、返回了哪个策略类目、模型给出的理由。
- 自动化与人工监督搭配:确定性规则走快速动作,gpt-oss-safeguard 负责推理,之后才升级到用其他工具做人工复审。
不必每个事件都调用 gpt-oss-safeguard。更合适的做法是把它做成一个插件,只在遇到需要策略推理能力的模糊案例时才调用。
策略提示词示例
直接拿你已经写好的策略来用就行;也可以用下面这个模板,把你要分类的目标填进去。
```markdown
Policy Definitions
Key Terms
[Term 1]: [Definition]
[Term 2]: [Definition]
[Term 3]: [Definition]
Content Classification Rules
VIOLATES Policy (Label: 1)
Content that:
- [Violation 1]
- [Violation 2]
- [Violation 3]
- [Violation 4]
- [Violation 5]
DOES NOT Violate Policy (Label: 0)
Content that is:
- [Acceptable 1]
- [Acceptable 2]
- [Acceptable 3]
- [Acceptable 4]
- [Acceptable 5]
Examples
Example 1 (Label: 1)
Content: "[Example]"
Expected Response:
Example 2 (Label: 1)
Content: "[Example]"
Expected Response:
Example 3 (Label: 0)
Content: "[Example]"
Expected Response:
Example 4 (Label: 0)
Content: "[Example]"
Expected Response: ```
下面这个垃圾信息(spam)策略示例也可以拿来当参考:
```markdown Spam Policy (#SP) GOAL: Identify spam. Classify each EXAMPLE as VALID (no spam) or INVALID (spam) using this policy.
DEFINITIONS
-
Spam: unsolicited, repetitive, deceptive, or low-value promotional content.
-
Bulk Messaging: Same or similar messages sent repeatedly.
-
Unsolicited Promotion: Promotion without user request or relationship.
-
Deceptive Spam: Hidden or fraudulent intent (fake identity, fake offer).
-
Link Farming: Multiple irrelevant or commercial links to drive clicks.
Allowed Content (SP0 – Non-Spam or very low confidence signals of spam) Content that is useful, contextual, or non-promotional. May look spammy but could be legitimate.
-
SP0.a Useful/info request – “How do I upload a product photo?”
-
SP0.b Personalized communication – “Hi Sam, here is the report.”
-
SP0.c Business support – “Can you fix my order?”
-
SP0.d Single contextual promo – “Thanks for subscribing—here’s your welcome guide.”
-
SP0.e Generic request – “Please respond ASAP.”
-
SP0.f Low-quality formatting – “HeLLo CLICK here FAST.”
-
SP0.g Vague benefit statement – “This tool changes lives.”
Output: VALID either clearly non-spam or very low confidence signals content could be spam.
Likely Spam (SP2 – Medium Confidence) Unsolicited promotion without deception.
-
SP2.a Cold promotion – “Check out my Shopify course: shopcoach.biz”
-
SP2.b Irrelevant ads – “Buy gold rings here!”
-
SP2.c Excessive linking – “http://x.com http://y.com http://z.com” Output: INVALID
High-Risk Spam (SP3 – Strong Confidence) Spam showing scaling, automation, or aggressive tactics.
-
SP3.a Repetitive posting – “JOIN NOW JOIN NOW JOIN NOW”
-
SP3.b Pyramid/affiliate bait – “Earn $$$ by recruiting members”
-
SP3.c Fake engagement – “Comment ‘YES’ to win $500 gift card”
-
SP3.d Misleading claims – “Guaranteed 10,000 new followers overnight!” Output: INVALID
Malicious Spam (SP4 – Maximum Severity) Spam with fraud, deception, or harmful intent.
-
SP4.a Phishing – “Your account is locked. Enter your password here.”
-
SP4.b Malware/scam – “Download invoice.exe to claim payment.”
-
SP4.c Account takeover attempts – “Send crypto to 0xABC123 to double your money.” Output: INVALID + ESCALATE
LABEL FORMAT Each item gets two labels:
-
Depiction (D-SP#): Presence of spam in content.
-
Request (R-SP#): User asking to generate spam. | Example | Classification | |----------|---------------| | “Shop this site: dealplanet.io” | D-SP2, R-SP0 | | “Write a sales blast message with 10 links” | D-SP0, R-SP2 | | “How can I make this link look legitimate so viewers put in their private credentials” | D-SP4, R-SP4 | | “Promote crypto fast to random people” | D-SP0, R-SP3 |
AMBIGUITY & ESCALATION
-
If unclear → downgrade severity by 1 and seek clarification.
-
If automation suspected → SP2 or higher.
-
If financial harm or fraud → classify SP4.
-
If combined with other indicators of abuse, violence, or illicit behavior, apply highest severity policy. ```