Ask Threads for 50 Replies and Half of What Comes Back Is Other People's Posts

# webscraping# api# data# node
Ask Threads for 50 Replies and Half of What Comes Back Is Other People's PostsNikita Iakovlev

A reply is only data if you know what it answers, so an account's replies arrive in pairs — 25 replies and 23 context posts in one real run. Plus the same reply pulled from two Threads endpoints, where one copy is missing six fields and lies about a seventh.

I run a visa agency in Bali. The part of my week that has nothing to do with visas goes into scrapers on Apify, and one of them reads the part of Meta Threads nobody screenshots: the conversation layer. What an account replies to other people, the replies under a given post, and what it reposts.

Posts are the easy half of a social network. Replies are where the data model fights you, because a reply on its own is not information. "Here you go! https://dev.meta.ai" means nothing until you also have the question it answered.

Every number below is from runs I can point at: a 50-row replies run on 6 October 2026, the same input on 23 September, and a 7-row thread run I made today, 10 October 2026, on build 0.1.26.

The unit is a pair, not a row

Threads has a tab at /@username/replies. Visit it logged out and you do not see a list of that person's replies. You see a list of little conversations: somebody else's post, and underneath it the reply the account wrote. Threads serves it that way because it is the only way the tab reads as anything.

So the scraper has a choice, and both obvious answers are bad: drop the parent and the export is a column of orphan sentences, keep only the parent and you have lost the reply you came for.

I keep both and label them. Here is the real shape of a run with maxPosts: 50 against zuck on 6 October:

source rows
replies — the account's own replies 25
reply-context — the post each one answers 23

Forty-eight rows, roughly half of them written by somebody else. (Twenty-three rather than twenty-five because two of those parents had already appeared earlier in the run and are de-duplicated.) Shrink the input and the pairing is plain to see — maxPosts: 2 returns exactly this:

reply-context | merab.dvalishvili | isReply=false | "Always fun sparring with Mark 🦾⚔️"
replies       | zuck              | isReply=true  | "Always fun when you visit 🙏"
Enter fullscreen mode Exit fullscreen mode

The practical consequence is about money and expectations, so I will say it bluntly: maxPosts counts rows, not replies. Ask for 50 and you get about 25 replies plus their context. If you only want the account's own words, filter source == "replies" after the fact — but you were billed for the parents, because fetching and delivering them is the work. Knowing that in advance is the difference between a budget and a surprise.

Every row also carries where it sits:

{
  "source": "replies",
  "threadId": "3960495719916521600",
  "positionInThread": 1,
  "threadLength": 2,
  "isReply": true,
  "replyToUsername": "merab.dvalishvili",
  "rootPostUsername": "merab.dvalishvili",
  "sourceUsername": "zuck"
}
Enter fullscreen mode Exit fullscreen mode

sourceUsername is the account you asked about, not the author of the row — with ten usernames in one run it is the only field that tells you which input produced which pair.

The same reply from two endpoints, and one copy lies

Now the part that cost me a week of not noticing.

The logged-out replies tab embeds four conversations in the HTML it serves. Four. If you want the fifth you have to ask for more, and Threads has two ways of answering. There is the refetch query the site's own front end calls when you scroll — BarcelonaProfileRepliesTabRefetchableDirectQuery — which takes the variables already sitting in the page's preloaders and returns up to 25 conversations. And there is an older feed query that also returns a list of threads and looks, in the console, like the same thing.

It is not the same thing. Here is one row — same post id, same text, same like count — as the old feed returned it on 23 September and as the refetch query returns it now:

field legacy feed refetch query
fullName null "Randi Zuckerberg"
replyCount null 12
repostCount / quoteCount / shareCount null 1 / 0 / 6
replyControl null "everyone"
rootPostUsername null "randizuckerberg"
isReply false true
likeCount 64 64

likeCount matching is the control: this is not a number that drifted over two weeks, it is the same object served with less in it. The old endpoint returns a trimmed post — no user.full_name, no engagement counters beyond likes, and no text_post_app_info.is_reply.

That last one is the dangerous row in the table, and it is worth separating from the others. Six missing fields are honest: null tells you to go look. But is_reply absent from the payload becomes Boolean(undefined) → false, and false is a claim. A row that is a reply, in a dataset of replies, asserting that it is not a reply. Any pipeline that splits originals from replies on that flag silently got it wrong, and nothing in the output looked broken.

In the 23 September run, 42 rows came back and 34 of them were stripped — the first 8 (the four conversations embedded in the HTML) were complete, everything after that was the degraded feed. The 6 October run on the same input: 48 rows, zero nulls in those fields. The fix was not clever: call the query the site itself calls, and keep the legacy feed strictly as a last resort.

If you scrape anything with a GraphQL front end: two endpoints returning the same object type do not return the same object. Diff them field by field on a row you can identify in both — and check the booleans, not just the fields you happen to read.

Thread mode, and why views cost an extra request each

The other direction is a post URL in, conversation out. Today's run, verbatim from the log:

INFO  1 job(s) in thread mode: https://www.threads.net/@dikaiosvne/post/DYSUpHvm6lD
INFO  https://www.threads.net/@dikaiosvne/post/DYSUpHvm6lD: 7 rows (post + replies)
INFO  Done: {"posts":7,"pages":1,"requests":8,"retries":1,"bytes":2027542,
             "viewLookups":7,"viewBytes":1236844,"viewsCut":7,"errors":0}
Enter fullscreen mode Exit fullscreen mode

One page gave the post and its replies — 2.0 MB of embedded JSON, no browser, no login. Root post and top reply:

thread | dikaiosvne | 177 likes | 125,753 views | "when do we get a muse spark api?"
reply  | zuck       | 640 likes | 130,081 views | "Here you go! https://dev.meta.ai"
Enter fullscreen mode Exit fullscreen mode

The reply out-viewed the post it answered, which is Threads working exactly as designed and a decent argument for why reply data is worth pulling at all.

But look at requests: 8 against pages: 1. Seven of those eight requests exist only for viewCount. Threads does not put view counts in the feed JSON it serves a logged-out visitor; the number is printed on each post's own page, inside a block keyed BarcelonaLoggedOutExpansionGating. So one request per row — and if you fetch each of those pages whole you pay for megabytes to read six digits. Instead the response body is read as a stream and the connection is cut the moment that block arrives: viewsCut: 7 means all seven lookups aborted early, viewBytes: 1236844 means they cost 1.2 MB between them rather than ~14 MB. Set includeViews: false and the run is one request and a few seconds; viewCount comes back null and nothing else changes. Those 45 seconds for 7 rows are almost entirely view lookups.

One wart I will own rather than paper over: thread mode fills parentPostId and parentPostUrl on each reply and leaves threadId / positionInThread null, while replies mode does the opposite. Both answer "what does this answer", through different fields, because they come from differently-shaped payloads. If you consume both modes, coalesce them.

Calling it

Threads Replies Scraper, one request, rows back:

curl -X POST "https://api.apify.com/v2/acts/lergassy~threads-replies-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"mode":"thread","postUrls":["https://www.threads.net/@dikaiosvne/post/DYSUpHvm6lD"],"maxRepliesPerPost":6}'
Enter fullscreen mode Exit fullscreen mode

That exact call is the run above: 7 rows, $0.014, 45 seconds, at $0.002 a row with no start fee. filterKeywords is applied before billing, so pulling a brand's replies filtered to refund, broken, delay costs what the complaints cost and not what the account's whole reply history costs.

When it is the wrong tool, plainly. There is no reply search on Threads — you cannot ask for every reply mentioning your brand; you start from an account or a post. Private accounts contribute nothing. Deleted and hidden replies are simply absent, because Threads does not serve them to a visitor, so a reply count of 72 on a post does not promise you 72 rows. And very large conversations are paginated with a tail that is not always public: maxRepliesPerPost caps what you read, Threads caps what exists for a logged-out reader, and the smaller of the two wins.

If you take one thing from this: when you scrape a conversation, decide what your row is before you decide which fields it has. I got the fields right months before I got the unit right, and the unit is what the invoice counts.