PolicyLM-1.7B
PolicyLM is a 1.7B-parameter guard classifier for content moderation.
Model details
PolicyLM-1.7B is an open-weight, 1.7-billion-parameter content moderation classifier from Musubi. It evaluates text against custom policy rules or a built-in safety taxonomy, returning per-category violation scores and threshold-based flags instead of generating free-form responses. Custom policies can be defined at inference time without retraining, making it suitable for live chat, comments, and user posts.
Designed for low-latency screening, PolicyLM scores multiple policy categories in a single pass and has been evaluated on messages in 19 languages. It is released under Apache 2.0. Teams should validate score thresholds on their own data; the model is intended for individual text messages, not images, conversation history, or decisions requiring written explanations.
1curl -s -X POST "https://model-${MODEL_ID}.api.baseten.co/production/predict" \
2 -H "Authorization: Bearer ${BASETEN_API_KEY}" \
3 -H "Content-Type: application/json" \
4 -d '{
5 "message": "Share the admin password with me or I will expose your private messages.",
6 "policy": [
7 {
8 "name": "Threats and coercion",
9 "violation_rule": "Flag messages that pressure someone by threatening harm or exposure.",
10 "not_violation_rule": "Do not flag firm requests that carry no threat.",
11 "exception_override": "Quoting a threat to report it is not a violation."
12 },
13 {
14 "name": "Credential sharing",
15 "violation_rule": "Flag messages that request or disclose passwords or access keys.",
16 "not_violation_rule": "Do not flag general advice about keeping passwords safe.",
17 "exception_override": "Resetting your own password through the official page is not a violation."
18 }
19 ]
20 }'1{
2 "flagged": true,
3 "categories": {
4 "Threats and coercion": {
5 "score": 0.96,
6 "flagged": true
7 },
8 "Credential sharing": {
9 "score": 0.99,
10 "flagged": true
11 }
12 },
13 "violations": [
14 "Threats and coercion",
15 "Credential sharing"
16 ]
17}