Are Language Model "AI" Tools Heading for 馃毥?

Words
641
Reading
3 min
Listen
Play
3y

image.png

You know the saying, garbage in, garbage out? We might be heading for peak GPT if the developers behind the tools are not careful.

Google is the latest to say that it has the right to scrape all publicly available information posted online for its AI projects. Their new policy specifically mentions that this includes building products and features like Google Translate, Bard, and Cloud AI capabilities.

Not only does this move raise privacy concerns (it gives Google the ability to harvest and use data posted on any part of the public web, which is public but usually aimed at human consumption) but I can see Google and others gobbling up masses of poor quality content becoming a problem.

AI Poop Pipeline

It's kind of like the photocopy of a photocopy of a ... ok, for the younger folks, it's like those JPGs from Aunty that have been processed so much you can't make out the pixels.

Other companies like OpenAI have also admitted they scraped the internet to fuel their AI. This was obviously not bad from a quality standpoint up to now, but those tools are being used a LOT, especially for "programmatic" and automated SEO posts. Training your model on already crap AI-generated text, and then further processing it?

Reddit and Twitter Turn AI Racist?

Remember when Microsoft's AI chatbot, Tay, was corrupted by Twitter users within 24 hours of its launch?

Tay was designed to learn and engage in "casual and playful conversation", users began tweeting misogynistic, racist, and offensive remarks to Tay and like a small, impressionable child, Tay started repeating these sentiments back to users.

"Hilarity" ensued.

It's one thing to "jailbreak" an AI to get it to say funny stuff but some of the offensive remarks made by Tay were unprompted, showing that the bot's entire learning had gone astray.

Remember this was the pre-$8-checkmark era, where Twitter on the whole surfaced the best comments rather than $8 payers. Now imagine an AI trained on Reddit, or god-forbid, 4chan ...

Open but not free

But even more tricky, who owns the content once it is published? Search engines obviously NEED the content, but they need to ask nicely before grabbing the stuff that puts food on people's tables, right?

The legality of the scrape and ask forgiveness practice is questionable in view of steps from Australia and Canada to restrict Google's slurping down their news and column content and is expected to be a topic of debate and a lot of litigation in the coming years.

Stop the Scrape or Freedom to Embed?

Another worry is freedom of automation - Twitter and Reddit have already made changes to their platforms to restrict access to their APIs, affecting third-party tools and causing controversy. Elon Musk has been vocal about his concerns regarding data scraping on Twitter, which led to limitations on the number of tweets users can view.

This effectively broke key uses for Twitter - you can't embed a tweet or share with friends effectively, or find them on Google. Think back to the last time you looked up a celebrity on Google, did you see their latest tweets?

As marketers we used to hope for a "hey Martha" moment, where one person would turn to another and say "Look at this!". Twitter just lost that.

Reddit notoriously has seen mass protests by moderators due to changes in API accessibility, which may have permanent damaging consequences for the platform.

Conclusion ...?

These tools are highly useful but the training is problematic.

Right now I am loving using GPT API to help automate things I could not do before, but as well as programming I also make my living from writing, either directly or indirectly.

Things are going to get interesting, and perhaps in the Chinese curse way ...

Are Language Model "AI" Tools Heading for 馃毥? | Ecency