- The X Activity API delivers a
post.reply.createevent to your webhook each time someone replies to one of the user’s Posts. - The xAI API asks Grok to judge the reply against a moderation policy you write in plain language, and to return its verdict as structured JSON.
- The Hide Replies endpoint hides the reply when Grok’s verdict crosses your confidence threshold.
An earlier version of this guide used Google Jigsaw’s Perspective API for toxicity scoring and the Account Activity API for delivery. Perspective has been deprecated, and the Account Activity API is being replaced by the X Activity API. The approach below uses Grok, which lets you describe your own policy and have the reply judged in the context of the Post it answers.
Prerequisites
- A developer account with a Project and App that has the X Activity API and Webhooks enabled.
- An xAI API key from console.x.ai. Purchases of X API credits earn free xAI API credits.
- A publicly reachable HTTPS endpoint to receive webhook events. See the Webhooks quickstart for URL requirements.
How the app works
Ask the user for permission
Authorize the user with the OAuth 2.0 Authorization Code Flow with PKCE and request these scopes:
Store the user’s access token, refresh token, and user ID. You can read the ID from
GET /2/users/me:Register a webhook
The X Activity API delivers events to a webhook you host. Register it once per environment with your app’s Bearer Token and keep the returned X immediately sends a Challenge-Response Check (CRC) to the URL and repeats it hourly. Your endpoint must answer the CRC and should verify the
webhook_id.X-Twitter-Webhooks-Signature-OAuth2 header on every event. Both are implemented in the full example below; the Webhooks quickstart explains them in detail.Subscribe to the user's replies
Create one From now on, each direct reply to one of the user’s Posts arrives at your webhook as a
post.reply.create subscription per user. This is a private event, so the request uses the user’s access token, and filter.user_id must be that user’s ID.Response
post.reply.create event. The payload is the reply Post; includes.tweets may contain the Post being replied to.post.reply.create
Ask Grok whether the reply breaks the policy
Send the reply, together with the Post it answers, to the xAI Responses API. The system prompt carries the moderation policy. A structured output schema guarantees that the answer is JSON your code can act on without parsing free text.The verdict is the Three request fields matter for this use case:
output_text item of the message in output. Other fields in the response are omitted here.Response (abbreviated)
store: falsetells the xAI API not to retain the reply text and verdict for later retrieval. Responses are otherwise stored for 30 days.reasoning.effort: "low"keeps latency and token usage down for a short classification task.text.formatwithstrict: trueconstrains the output to the schema, so a malformed answer cannot slip through to the hide step.
Hide the reply when confidence is high
Hide only when Record every verdict, including the ones you did not act on, so the user can see what the app decided and reverse it. Unhiding is the same request with
violates_policy is true and confidence meets a high threshold (this guide uses 0.9). The request uses the subscribed user’s access token because only the author of the conversation can hide replies in it.cURL
Response
"hidden": false.Put it together
The examples below are complete webhook servers. Each one answers the CRC, verifies the event signature, deduplicates events byevent_uuid, classifies the reply with Grok, and hides it when the verdict crosses the threshold. The in-memory stores stand in for the database you would use in production.
Python (Flask)
Tune the policy
The system prompt is where your app’s judgment lives, and it is the main thing that changes from one deployment to the next.- Write the policy for the account it protects. A brand account might hide competitor spam and off-topic promotion; an individual might only want slurs and threats gone. Say what to hide and, just as importantly, what to leave alone.
- Give Grok the conversation. Including the original Post lets the model tell a hostile reply from a blunt but relevant one. Use the Post from
includes.tweetswhen it is present, or fetch it with Post lookup when it is not. - Keep the schema small. The fields in
reply_moderationare enough to decide, explain, and audit. Add a category to theenumwhen the policy grows a new rule; avoid open-ended fields that your code cannot act on. - Set the threshold high and review the rest. Hide automatically only at high confidence. Route lower-confidence violations to a review queue in your app so a person makes the call, and use those decisions to refine the prompt.
- Test with real replies before switching on auto-hide. Run the classifier in log-only mode for a while, compare verdicts against what the user would have done, and adjust the policy text until the two agree.
Keep the user in control
Regardless of the model or the approach you use, make the best possible effort to ensure that your users understand what your app has hidden and can change it.- Trust the user and give them full control over their decisions. Your interface should list every reply the app hid, show the
reasonGrok returned, and offer a one-click undo that calls Hide Replies with"hidden": false. - Hide only at a very high confidence threshold. A reply left visible for a moment costs little; a reply hidden by mistake can silence a legitimate voice.
- Not everybody uses the same words. Reclaimed words, slang, and in-group humor can look like violations out of context. Tell the model about them in the policy, and let users add their own exceptions.
- Be clear with reply authors and readers. Hidden replies are still reachable through “View hidden replies”, and the reply author is not notified. Your app should not imply otherwise.
Operational notes
- Answer the webhook quickly. Return
200 OKas soon as you have verified the signature and queued the event. Call the xAI API and Hide Replies in the background, as the examples do. - Deduplicate. Keep the
event_uuidof processed events and skip repeats. - Direct replies only.
post.reply.createfires for direct replies to the user’s Posts, and does not fire for replies to replies. Hide Replies can hide any reply in a conversation the user started, so pair this guide with Manage replies by topic to sweep deeper threads with recent search andconversation_id. - Protected accounts. Posts from protected accounts are not delivered through the X Activity API, so replies from protected accounts will not reach your webhook.
- Token lifetime. OAuth 2.0 access tokens expire after two hours. Use the refresh token from
offline.accessto obtain a new one before calling Hide Replies. - Revocation. A user-context subscription is removed when the user revokes your app. Subscribe to the
oauth.revokeevent to clean up your stored tokens and decisions at the same time. - Billing. Each delivered
post.reply.createevent is billed as a Post under your X API plan, and each classification consumes xAI API tokens; see xAI pricing. X API credit purchases earn free xAI API credits.
Next steps
Manage replies by topic
Moderate an existing conversation with recent search and Post annotations
X Activity API
Event types, filters, and authentication for real-time delivery
Webhooks quickstart
CRC validation, signature verification, and webhook registration
xAI structured outputs
Schema rules and SDK helpers for JSON responses from Grok