AI Safety and Alignment
Refuse too little and things go wrong; refuse too much and nobody uses it. The hard part has always been where that line falls, not whether guardrails exist at all — guardrail models, streaming review and the cost of over-refusal.
Too few refusals cause incidents, too many drive users away — guardrails are about calibration. Qwen3Guard lifts its safety rate on WildJailbreak from 64.7 to 98.1 via safety RL without hurting general ability, covering thought-chain and streaming review; OpenAI, with mental-health experts, cut undesired replies in sensitive chats by about sixty percent. Goody-2 refuses even arithmetic — the ready-made counterexample.