While the current coronavirus pandemic has removed a necessary aspect of in-person communication from our daily lives, it has also emphasized the need for human-computer interaction, especially in regard to the use of computational techniques to better understand linguistics and speech. As multilingual translation services and speech recognition AI continue to etch themselves within our everyday routines, technology plays a larger role in the way we virtually communicate and interact with each other, despite being miles apart in real life. In our modern society, which places a heavy emphasis on technology, vast volumes of text data are created daily – from long conversations containing lengthy text messages to billions of google searches and online queries. However, with all of this unstructured text data piling up unnoticed, combined with the astonishingly diverse human language, which contains unique syntax, grammatical rules, and slang that are incredibly difficult to understand, Natural Language Processing (NLP) poses significant challenges to digital marketing agencies and major corporations looking to analyze speech data. Fortunately, this subfield also offers an intriguing insight into the intricate interplay between technology and linguistics.
The fascinating scope of human-computer interaction!
The main drawback to processing text with data-driven programming languages like Python and R is that understanding and manipulating the human language is fundamentally difficult; accounting for syntactical rules and exceptions proves to be an overly complex challenge. However, NLP’s novel use of algorithmic text processing techniques makes this task slightly less daunting. Let’s take a look at four main concepts that underlie this mechanism:
- Tokenization – The process of segmenting long lines of text into individual words and sentences that can easily be understood and isolated for further analysis. By cutting up drawn-out, running text into small pieces called tokens, tokenization breaks down text data in a way machinery can meaningfully comprehend and process.
- Stemming – After volumes of text data have been successfully processed, stemming is performed, a concept that refers to slicing the starting and ending points of a given word to remove affixes or lexical additions to a word’s base root. Crudely chopping off the beginning and end of an inflectional word to its common root form, stemming offers a computationally efficient solution for making text processing more convenient within data-driven technology.
- Lemmatization – Similar to the concept of stemming, lemmatization reduces a word to its dictionary base form and groups together different forms of the same word. Unlike stemming, however, lemmatization relies on resolving words to their morphological origins rather than rudimentarily chopping off their lexical endings (e.g. “runs”, “running”, and “ran” would translate to the common base origin of “run”).
- Stop Word Removal – Amidst the issues that computationally-deriving meaning from text pushes forward, common articles and prepositions such as ‘the’ and ‘to’ do not make things any easier. This is where Stop Word Removal comes in – a filter that directly removes widespread and commonly used terms that aren’t necessarily meaningful to the text being analyzed.
While these concepts only graze the surface of the expansive realm of Natural Language Processing, they importantly reveal the inner abstraction that comes with developing a reliable computational understanding of language-based data. The fascinating scope of human-computer interaction allows us to both humanize complex computational tools and familiarize ourselves with their extensive capabilities – a duality that strongly echoes the importance of technology in our rapidly evolving world.
Works Cited
1. Garbade, Dr. Michael J. “A Simple Introduction to Natural Language Processing.” Medium,
Becoming Human: Artificial Intelligence Magazine, 15 Oct. 2018,
becominghuman.ai/a-simple-introduction-to-natural-language-processing-ea66a1747b32.
2. “What Is Text Mining, Text Analytics and Natural Language Processing?” What Is Text Mining,
Text Analytics and Natural Language Processing? Linguamatics,
http://www.linguamatics.com/what-text-mining-text-analytics-and-natural-language-processing
3. Yse, Diego Lopez. “Your Guide to Natural Language Processing (NLP).” Medium, Towards Data
Science, 30 Apr. 2019,
towardsdatascience.com/your-guide-to-natural-language-processing-nlp-48ea2511f6e1.