Over the last weeks, I've been thinking about the curation-reward approach of steemit. I was intrigued by how large human curated networks like curie are trying to discover new aspiring authors. You certainly get a lot of high-quality content discovered by this. But somehow I can't shake off the feeling that you also miss a lot of new authors.
The Lost Authors
That's what I'd call those people who haven't been discovered by large curation networks and get discouraged by the lack of interaction. Sure, there's no such thing as instant gratification. You won't get thousands of votes even if you deliver Pulitzer prize-winning blog posts. You need to be persistent and keep on writing. Otherwise, you won't build a following. But wouldn't it be great to catch every promising author? We do live in the age of artificial intelligence and machine learning. Why not use the extensive knowledge of those curation networks to discover them.
Approaches on Finding Them
Finding an aspiring author consists of multiple aspects.
1. User Interaction
This includes how many times a user writes blog posts, comments on other users posts, upvotes and resteems. But how do you detect whether that interaction is a high-quality interaction? A bot which upvotes and resteems every post is certainly high on the interaction scale, but the quality of the interaction is deficient. One would need to detect whether the author strikes up conversations with other users, not only on his posts but on others as well. Maybe a combination of word count and conversation length?
2. Text Quality
The most prominent part of the blog post: letters and words. There are specific readability metrics to aide this. This includes average sentence length, word length etc. On the other hand, there are things like word variability, the part of speech and the words themselves. A simple n-gram bag of words model should uncover at least something on the text quality, aside from that, trial and error. And of course, reading some natural language processing papers.
Images and Links
One thing which this post is missing is images and links. But for a high-quality blog post, those two things are essential. High-quality photos make a post more interesting. I think advances in the field of image processing might help in detecting high-quality images.
Conclusion
Those are my thoughts on the subject of automated curation. It's going to be certainly an exciting process. For a start, I've extracted a couple of thousand posts from steemit which were curated by curie. This should help train a first machine learning model. Onto the feature engineering part!