Skip to main content
With the Hide Replies endpoint, you can build a workflow that hides replies to a user’s Posts as soon as they arrive, when there is a very high probability that the reply breaks the rules the user set for their conversation. This guide combines three APIs:
  1. The X Activity API delivers a post.reply.create event to your webhook each time someone replies to one of the user’s Posts.
  2. The xAI API asks Grok to judge the reply against a moderation policy you write in plain language, and to return its verdict as structured JSON.
  3. The Hide Replies endpoint hides the reply when Grok’s verdict crosses your confidence threshold.
An earlier version of this guide used Google Jigsaw’s Perspective API for toxicity scoring and the Account Activity API for delivery. Perspective has been deprecated, and the Account Activity API is being replaced by the X Activity API. The approach below uses Grok, which lets you describe your own policy and have the reply judged in the context of the Post it answers.

Prerequisites


How the app works

Ask the user for permission

Authorize the user with the OAuth 2.0 Authorization Code Flow with PKCE and request these scopes:Store the user’s access token, refresh token, and user ID. You can read the ID from GET /2/users/me:

Register a webhook

The X Activity API delivers events to a webhook you host. Register it once per environment with your app’s Bearer Token and keep the returned webhook_id.
X immediately sends a Challenge-Response Check (CRC) to the URL and repeats it hourly. Your endpoint must answer the CRC and should verify the X-Twitter-Webhooks-Signature-OAuth2 header on every event. Both are implemented in the full example below; the Webhooks quickstart explains them in detail.

Subscribe to the user's replies

Create one post.reply.create subscription per user. This is a private event, so the request uses the user’s access token, and filter.user_id must be that user’s ID.
Response
From now on, each direct reply to one of the user’s Posts arrives at your webhook as a post.reply.create event. The payload is the reply Post; includes.tweets may contain the Post being replied to.
post.reply.create

Ask Grok whether the reply breaks the policy

Send the reply, together with the Post it answers, to the xAI Responses API. The system prompt carries the moderation policy. A structured output schema guarantees that the answer is JSON your code can act on without parsing free text.
The verdict is the output_text item of the message in output. Other fields in the response are omitted here.
Response (abbreviated)
Three request fields matter for this use case:
  • store: false tells the xAI API not to retain the reply text and verdict for later retrieval. Responses are otherwise stored for 30 days.
  • reasoning.effort: "low" keeps latency and token usage down for a short classification task.
  • text.format with strict: true constrains the output to the schema, so a malformed answer cannot slip through to the hide step.

Hide the reply when confidence is high

Hide only when violates_policy is true and confidence meets a high threshold (this guide uses 0.9). The request uses the subscribed user’s access token because only the author of the conversation can hide replies in it.
cURL
Response
Record every verdict, including the ones you did not act on, so the user can see what the app decided and reverse it. Unhiding is the same request with "hidden": false.

Put it together

The examples below are complete webhook servers. Each one answers the CRC, verifies the event signature, deduplicates events by event_uuid, classifies the reply with Grok, and hides it when the verdict crosses the threshold. The in-memory stores stand in for the database you would use in production.
Python (Flask)

Tune the policy

The system prompt is where your app’s judgment lives, and it is the main thing that changes from one deployment to the next.
  • Write the policy for the account it protects. A brand account might hide competitor spam and off-topic promotion; an individual might only want slurs and threats gone. Say what to hide and, just as importantly, what to leave alone.
  • Give Grok the conversation. Including the original Post lets the model tell a hostile reply from a blunt but relevant one. Use the Post from includes.tweets when it is present, or fetch it with Post lookup when it is not.
  • Keep the schema small. The fields in reply_moderation are enough to decide, explain, and audit. Add a category to the enum when the policy grows a new rule; avoid open-ended fields that your code cannot act on.
  • Set the threshold high and review the rest. Hide automatically only at high confidence. Route lower-confidence violations to a review queue in your app so a person makes the call, and use those decisions to refine the prompt.
  • Test with real replies before switching on auto-hide. Run the classifier in log-only mode for a while, compare verdicts against what the user would have done, and adjust the policy text until the two agree.

Keep the user in control

Regardless of the model or the approach you use, make the best possible effort to ensure that your users understand what your app has hidden and can change it.
  • Trust the user and give them full control over their decisions. Your interface should list every reply the app hid, show the reason Grok returned, and offer a one-click undo that calls Hide Replies with "hidden": false.
  • Hide only at a very high confidence threshold. A reply left visible for a moment costs little; a reply hidden by mistake can silence a legitimate voice.
  • Not everybody uses the same words. Reclaimed words, slang, and in-group humor can look like violations out of context. Tell the model about them in the policy, and let users add their own exceptions.
  • Be clear with reply authors and readers. Hidden replies are still reachable through “View hidden replies”, and the reply author is not notified. Your app should not imply otherwise.

Operational notes

  • Answer the webhook quickly. Return 200 OK as soon as you have verified the signature and queued the event. Call the xAI API and Hide Replies in the background, as the examples do.
  • Deduplicate. Keep the event_uuid of processed events and skip repeats.
  • Direct replies only. post.reply.create fires for direct replies to the user’s Posts, and does not fire for replies to replies. Hide Replies can hide any reply in a conversation the user started, so pair this guide with Manage replies by topic to sweep deeper threads with recent search and conversation_id.
  • Protected accounts. Posts from protected accounts are not delivered through the X Activity API, so replies from protected accounts will not reach your webhook.
  • Token lifetime. OAuth 2.0 access tokens expire after two hours. Use the refresh token from offline.access to obtain a new one before calling Hide Replies.
  • Revocation. A user-context subscription is removed when the user revokes your app. Subscribe to the oauth.revoke event to clean up your stored tokens and decisions at the same time.
  • Billing. Each delivered post.reply.create event is billed as a Post under your X API plan, and each classification consumes xAI API tokens; see xAI pricing. X API credit purchases earn free xAI API credits.

Next steps

Manage replies by topic

Moderate an existing conversation with recent search and Post annotations

X Activity API

Event types, filters, and authentication for real-time delivery

Webhooks quickstart

CRC validation, signature verification, and webhook registration

xAI structured outputs

Schema rules and SDK helpers for JSON responses from Grok