A Much Deeper Look into Duplicate Comments on HIVE

Words
714
Reading
4 min
Listen
Play
1y

This work is a follow up to recently published investigations:

This post examines data from January 2025 through to the conclusion of April 2025, to enable broader analysis of the type of comments published on the platform.

The data set examined in this post has 1,758,556 rows. This covers 118 days of full data, representing an average of 14,903 comments a day.

Through this post, I will focus on a few interesting topics, as I refine my investigations into a deeper analysis leading down to the comment quality.

Before we look at comment quality, I want to revisit the comment "uniqueness" path.

Comment Uniqueness / Duplicates


TLDR: 27.1% (476,937) of comments are not unique.

Of the rows examined, 1,281,619 contain unique comments, which means that we are left with 476,937 comments which are repeated. Some are repeated by the same user, while others are comments left by different users, that contain the same comments.

There are 44,876 rows of duplicate comments after they are grouped into the format:

| Comment Text | Frequency | Unique Authors | Percentage | Word Count |

What causes not unique comments?

Examination of the duplicate rows reveals the following (and a lot of this is not new information) Note that all percentages expressed here are the percentage of the non unique comments, not the total percentage of comments on the chain for the period:

Notable Curation Projects leaving duplicate comments

  • Curation Projects Leaving a Comment with the "Felt" Sentiment token making up some 16.4% (7401) of the period's non unique comments
  • Splinterboost upvoting posts about Splinterlands and leaving a comment, 13.4% (6039) of duplicate comments for a 12% upvote, (3.2%, 1478 comments) for 5% upvotes and (3%, 1327 comments) for 15% upvotes. There's probably more. By my count, this makes splinterboost the most prolific in publishing duplicate comments to hive.
  • Pandex (2333 comments, by 1 user, for 5.19%)

Bot Calls for Tokens

Bot Calls, including calls

bbh (6651 comments, by 1223 users) 14.82%
lolz (5252 comments by 129 users) 11.7%
hiqvote (3145 comments by 88 users) 7%
(beer 2556 comments by 73 users) 5.69% (nice?)
lady (1939 comments by 2 users) 4.32%
pizza (1855 comments by 102 users) 4.13%
hbit (1418 comments by 91 users) 3.16%

And these are only the top 25 duplicate comments being examined! 22,816 comments are ONLY calls to these bots, without any further remark.

Common Phrases

Humans say the same thing to humans.

The fifth most repeated, non-unique comment is "thank you", with it being made 5,027 times by 919 users, for a total of 11% of the duplicate comments. Variations such as "thanks" appear 2970 times, by 700 users, for 6.6%, then we get "thank you!" 2563 times by 523 enthusiastic users. We say thank you a lot, here's some examples, and their incidence:

CommentFrequencyNumber of Users% Non UniqueWord Count
thanks29707006.6182371
thank you!25635235.7112932
thanks!21513384.7932081
thank you.15873113.5364112
thank you so much11243822.504684
you're welcome10831412.4133172
gracias9172532.0434091
thank you so much!6011821.3392464
thank you so much.4701571.047334
welcome450681.0027631
you are always welcome41680.9269994
muchas gracias4121590.9180852
your welcome198300.4412162
thank64310.1426151
thank youu25180.0557092
thank you)1150.0245122
graciasss970.0200551
thank you,950.0200552
graciass650.013371
thankssss550.0111421
graciassss440.0089131
thanksss440.0089131
gracia330.0066851
thank you so much,220.0044574

There are probably more variations that my regex didn't catch, but the totals here are 15,084 comments, by 3,341 users.

Now I'm interested in commonly appearing terms in the comments:

First, let's start with some words:

Duplicate Comment ContainsAppears in How Many Comments
hive168093
thank92066
vote74882
HP46458
delegate30425
curate27366
witness16147
delegation12540
splinterlands10474
splinterboost10435
shit9200
fuck1157
nice post157

image.png

Next, lets look at some web activity:

Duplicate Comment ContainsAppears in How Many Comments
"http" (Contains some sort of link)188203
".png" (Contains a PNG)83684
".gif" (Contains a GIF)33247
".jpg" (Contains a JPG)26893
"youtube" (mentions YouTube)4985
"twitter" (mentions twitter)4603
".webp" (Contains a webp)1331
"Reddit" (mentons reddit)524
"facebook" (mentions facebook)368
".jpeg" (Contains a JPEG)286
ChatGPT13
"github" (mentions github)7

image.png

The next part of my study, which I'll do at some point in the future, will focus on "Ranking" user comments in terms of complexity. This will take into account length, word count, sentence counts, and comment depth. I will come up with some sort of scoring algo, and see who comes out on the top.

As that will be based on a "user" dimension, we'll finally find out who swore the most in the study period ;)

A Much Deeper Look into Duplicate Comments on HIVE | Ecency