Chat and content moderation
Most moderation APIs score a message in isolation, against a fixed list of harm categories. Your policy is usually more specific than that ("no sharing phone numbers before a booking", "no reviews from people who never ordered"), and the same words can be fine from one user and a red flag from another. With SigWise you write the policy as signals, and every message is judged together with its author's history.
Write your policy as signals
json
PUT /v1/signals/policy_violation
{
"type": "choice",
"instructions": "Does this user's latest content break the community rules?",
"criteria": {
"none": null,
"contact_sharing": "shares phone numbers, emails or social handles to move off-platform",
"harassment": "insults, threats or targeted abuse of another user",
"spam": "repeated promotion, links or copy-pasted messages",
"prohibited_item": "offers weapons, drugs or counterfeit goods"
}
}One signal per decision keeps each answer sharp. See Writing good signals.
Check before publishing
Send the message with "wait": true and only the signals the decision needs. The API records it, scores the author inline, and returns the verdict:
json
POST /v1/objects/user-42/events
{
"wait": true,
"signals": ["policy_violation"],
"events": [{ "type": "message", "content": "text me on 555 0100, cheaper outside the app" }]
}json
{
"object_id": "user-42",
"analyzed": true,
"answers": [
{
"signal": "policy_violation",
"type": "choice",
"choice": "contact_sharing",
"probabilities": { "none": 0.03, "contact_sharing": 0.94, "harassment": 0.01, "spam": 0.01, "prohibited_item": 0.01 }
}
]
}Publish, hold for review or reject based on the answer. Set "include_history": false to judge only the new content. See Direct moderation for what to do when a check fails.
Moderate in the background too
Not everything needs a blocking check. Send the rest of the activity asynchronously, and let rules alert your team when a user's behaviour crosses a line over time. Rules fire after inline and background analyses alike.