AI Weekly Malaysia

Back to items Summaries

Does Reddit have an astroturfing problem? What the data suggests

ID
29948
Status
summarized
Published
28 Sep 2026, 9:30 PM
Fetched
30 Sep 2026, 2:08 AM
Provider
Hacker News
Category
dev-community
Original URL
https://www.petervijeh.com/projects/reddit-astroturf
Source URL
https://hnrss.org/best

Summary

Score
7.5
Created
30 Sep 2026, 2:10 AM
Tags
Audience
developersai_ml_learnerssaas_founders

What happened

Peter Vijeh fine-tuned a small GLiNER named-entity model to extract brands, models and steels from knife comments across six subreddits (r/knives, r/knifeclub, r/chefknives, r/japaneseknives, r/FixedBladeEdc, r/KnifeSteels), then asked who does the recommending in 'what should I buy' threads. The headline finding: one chef's-knife brand gets 31% of its buying-thread mentions from 5% of the accounts, four times what chance would predict. He says the buying-thread numbers can be recomputed from the published data with one script, but the account-history comparison cannot, because it rests on usernames he will not publish; the post drew 276 points and 359 comments on Hacker News.

Why it matters

If you use the 'append reddit to a Google search' trick for product or tooling research — or if your growth plan is seeding Reddit comments — this gives you a concrete number to reason about: one brand taking 31% of recommendation mentions from 5% of accounts. Note what you can and cannot verify: the 4x concentration is recomputable from the published data, the account-history evidence is not, so treat the second claim as unverified and the first as a measurable pattern you could run on your own category. Vijeh also states the post was drafted with AI from his outline and run logs before editing, which is worth knowing when you weigh the prose against the code.

Discussion angle

Is 4x concentration in 'what should I buy' threads actually evidence of paid astroturfing, or what a small subreddit with a handful of loud regulars looks like anyway — and how would you design the analysis (control brand, account-age baseline, comment-diversity scoring) to tell the two apart?

Top