Welcoming Musubi's PolicyLM, and a Guide to Choosing Open Safety Models
Written by
Jeremie Ponak, Alice Hunsberger, and Juliet Jonak
Published
Musubi's PolicyLM-1.7B is joining the ROOST Model Community (RMC). PolicyLM is a 1.7B-parameter decision model that reads your policy and returns a 0–1 score for multiple categories in one pass, fast enough for live chat.
To mark the launch, ROOST and Musubi are publishing Choosing and Routing Open Safety Models, an open source guide on how to build and select the right models for each task in safety.
You can read the full version at the link above, including reference architectures and worked examples. But if you want a sneak peak - and a link to our workshop later this month where we explore some of these themes, read-on.
What we cover
In 1996, residents of Scunthorpe, England, couldn’t sign up for AOL because a profanity filter caught four letters inside their town's name. Thirty years later, safety teams have far better options! But with a proliferation of open source safety models, finding the right tool for the job isn’t always obvious. The goal of the report is to help builders avoid a Scunthorpe 2.0: a model that ignores your community's policies, or one whose cost and speed let you review only a fraction of your content.
The guide answers this practical question: which model should handle which decisions when it comes to safety? Part 1 covers selection through four questions: how accurate a model is on your policy, what it costs at your scale, how fast it is, and how easily you can steer it as your policy changes.
Part 2 covers routing: when to cascade from a small model to a larger one, how to set thresholds and escalation paths, and when a single model is the better call.
Where PolicyLM fits
Live chat has long relied on fixed classifiers, because generative models take seconds per check. PolicyLM is part of an emerging family of models called “decision models” that return scores for predefined answers rather than generating a written response token by token. Because of this, PolicyLM can read your own policy at classifier speed, so teams can enforce custom rules on chat, game lobbies and DMs. The guide shows how it pairs with other RMC models, such as Roblox's PII classifier and Sentinel, in a full live-chat stack.
Find out if a model follows your rules
Models and their behavior change fast, and no single team's testing covers every platform. The patterns in this report came from running models on real content and comparing notes, and we want to keep doing that in the open.
On October 27, ROOST and Musubi are hosting a workshop on safety models: “Does a Model Follow Your Rules?”. Bring a policy you work with and run it across RMC models, including PolicyLM, and we'll walk through where models agree, where they diverge, and how to set thresholds and routing for your own surfaces. Findings from the session will feed back into the RMC!
Sign up for the October 27 workshop
Read the full report, explore the models in the ROOST Model Community, and tell us what else you're finding about Open Safety models!