Online Workshop

Does a model follow your rules?

Safety models are everywhere in software. But do they listen?

October 27, 2026 12 PM ET • Hosted by ROOST & Musubi

Register Now

Language models are increasingly being used to help moderate content: reading a set of rules, looking at a post or conversation, and deciding whether it crosses the line.

But how well does that actually work? What happens when a rule is ambiguous? When context matters? When two reasonable people might read the same policy differently? And what gets lost when instructions written for humans are handed to a model?

This is a small working session to test those questions directly!

Bring one section of rules you know well and a few examples where you have a view on the right answer. We’ll run the same rules and cases across several open models, including Musubi’s PolicyLM, and look at where the models agree with you, disagree with you, or disagree with each other.

No machine learning background is needed, and there is nothing to install. If you do not have a policy to bring, we’ll provide one, or you can choose some from our policy packs.

Agenda
10 minutes on how these models are being used for moderation; 25 minutes testing your rules and examples; 15 minutes examining disagreements and edge cases; 10 minutes comparing what we learned.
Bring
One topic’s worth of rules and 3-5 examples.
Afterwards
With participant consent, ROOST will synthesize anonymized observations from the session.