Concept Analytics Lab

Text is data too: what a day of training on corpus methods showed me

by Justyna Robinson
Last week the Concept Analytics Lab team ran a one-day course for the National Centre for Research Methods on meaning extraction from large text data. We taught how thematic analysis can be led by the empirical methods of corpus linguistics, showed people how to build their own corpora for exploration, and demonstrated Concept Cruncher, a new meaning extraction tool developed in the Lab.

Three things in particular stayed with me.

The first observation is how much empirical techniques of language analysis are needed beyond academia. The room held academics, public servants and people from non-governmental and charitable organisations. Their data was just as varied: surveys with a free-text element that somebody has to make sense of, recordings and transcriptions of focus groups, conversations between customers and providers, and material published on the internet. Some of it was spoken, some written. The common thread was that everyone in that room was holding more language than a person can read.
The second observation is the appetite for learning how to treat language as data. While there are public conversations about advanced numeracy, the equivalent skill for texts gets far less attention. Sometimes this is down to a disciplinary bias, visible in the habit of calling language unstructured data. The term comes from database management, where it means information without a predefined format (IBM), the kind that fits into rows, columns and schemas. Language does not work that way, since the same meaning can be expressed in many different ways, and so it gets classed as unstructured alongside images, audio and video. Language is nonetheless an organised system. Making that existing structure tractable is precisely what corpus linguistics does. Participants wanted the skills to find patterns in language data, take control of large amounts of text, and evaluate it critically.

That request makes more sense once you hear how people had been working before. And this is where my third observation comes in. Almost everyone had tried to analyse their text themselves, and run into the same difficulty. The material was too large to annotate or colour code through close reading. Turning to generative artificial intelligence did not help either, since these models tended to force new data into pre-existing formats. If you are in the business of knowledge making, that is a serious constraint.

Whichever method people used to read the data, they soon began to doubt their own reading. Prior knowledge of a field helped them interpret text, but it also shaped what they noticed, and a theme in the data started to look like a theme they had expected to find. Soon the whole task became overwhelming, and it stopped being empirical.
Certainty about what data tells you is crucial when writing up research conclusions, policy recommendations, or a new strategy. Participants left the training day with exactly that skill: an empirical way of reading large amounts of text that they could stand behind.