Natural language generation: text plans and semantic networks

Words
602
Reading
3 min
Listen
Play
9y

Hi all, for those who want to understand a bit better on the subject of AI generation of posts without those posts being gibberish (as seen from markov chain outputs), I thought i'd explore it more.

Within the field of computational linguistics there's the subfield of discourse analysis and generation that study how to analyse and generate larger documents, as opposed to single sentences and the concept of text plans comes from this field.

A text plan is simply an outline of a document, like a template into which the real substance of the document can be placed.  For an example, here's the (manually specified) text plan behind my recent hacked together post generator's first post:

Title: Facebook zero and net neutrality
What is net neutrality?
[Net Neutrality]
[Zero Rating]
[Net Neutrality and Zero Rating]

What is facebook zero?
[Facebook zero and zero rating]
[Facebook zero and Wireless]

Why facebook zero is relevant to Net Neutrality
[Facebook Zero and Net Neutrality]

The formatting is a bit messed up unfortunately, but it's clear how this forms a basic outline of a post about facebook zero and net neutrality. We can derive such text plans from rulesets (in the above case I cheated - once I had the script generate the related concepts, I tweaked the output and selected concepts that it made sense to write a post about, I then specified a text plan by hand).

In the above post we have a set of concepts under discussion: Facebook Zero, Net Neutrality, Zero rating, Wireless. We then simply look to the semantic network of our choice (I used a mix of dbpedia, MIT's ConceptNet and OpenCyc) and pull out relations between the concepts in order to flesh out the text plan.

I also cheated by not bothering with WSD (Word Sense Disambiguation) - I manually selected the sense of the phrase "Zero rating".

After importing the knowledge into an inference engine like OpenCog's atomspace, we can run queries on it.

After running these queries, a library like MontyLingua can be used for sentence generation to fill out the text plan with actual natural language.

Under "What is Net Neutrality?" we query the semantic network for "IsA" relations and get back a machine-readable description of the concept, which i'll summarise here:

Net Neutrality IsA proposed regulatory principle.

Regulatory principles state policies.

Net Neutrality IsA regulatory principle.

Thus, Net Neutrality states policies.

I used a manual tweak here (in a real post-generation system you'd build up these sorts of tweaks as a ruleset over time and manually observe the output to ensure it makes sense), at first the output I got was this:

Net Neutrality is a proposed regulatory principle that states policies

That is clearly not sufficient - what policies does it state? So I tweaked it to replace "policies" with the actual policies and selected manually the simplest such policy: the idea that ISPs should treat packets equally.

That leads us to the final output:

Net Neutrality is a proposed regulatory principle that states ISPs Should treat Packets and content equally.

The "from different networks" bit is something I added manually to slightly humanise it.

With the rest of the text plan I made similar tweaks - some just manually editing the output, some to the actual code - the final result was a post that is mostly software generated. I intend to clean up the system a bit and use it for some of my own posts as a source of inspiration and ideas, slowly making it more advanced over time - and then, and only then, will I release the source.