Computational linguistics is the scientific study of human language, examining the rules, structures and interrelated systems that humans use to communicate. While Large Language Models generate remarkably fluent text, they operate on fundamentally different principles from human language: statistical patterns between tokens rather than the layered syntactic, semantic and pragmatic structures that linguists study. Understanding this distinction is essential for anyone working with AI language systems, as it reveals where interventions can be made to improve model output.
This article serves as a guide to navigate these differences, offering a clearer framework for thinking about language itself and what you can, and cannot, expect from current state-of-the-art language models.
“If you are going to be a teacher of English, you need a scientific understanding of it.” Wilson Miller, Alexa Home Assistant team, Amazon
Your flywheel is only as fast as your annotation pipeline.
LXT embeds into your retraining loop, delivering labeled data at the cadence your model demands, not months later.
What Does ‘Language’ Mean?
What is a language? And what is ‘language’? They seem like simple enough questions, but they are anything but simple to answer. The obvious answers are: ‘language is the word we use to describe how humans communicate’, and ‘a language is the specific way a group of humans communicate, using particular words and grammatical rules’.
But you can get different answers depending on who you ask. A primary school teacher might focus on a student’s ability to understand an author’s and illustrator’s description of a story. A computer scientist might focus on programming languages and sorting algorithm logic. A company selling an AI tool might focus on interpreting requests and providing correct products or information.
‘Language’ is not a single, universal concept. In the age of Large Language Models, the differences in how we use the term can make it difficult to understand one another, even when we are all using ‘English’.
Before we can talk sensibly about what LLMs can and cannot do, we need to examine what human language actually is: a layered, structured, deeply social system involving syntax, meaning (both semantic and pragmatic), intention, identity and context.
What Linguists Mean by ‘Language’
Linguists study humans to figure out the rules of the structured, rule-governed human behaviour that is language. In the 1800s, English ‘language rules’ were prescriptive lists showing education through adherence. These included ‘do not split infinitives’ (But I like to occasionally split an infinitive!), ‘do not end a clause with a preposition’ (A preposition is something you must never end a sentence with), and ‘avoid double negatives’. Most were based on languages that educated English speakers considered ‘better’ than their own tongue, often Latin.
During the 20th century, linguists began conducting empirical research into the rules that govern language, rather than defining it. Modern linguists are interested in describing rather than prescribing language use.
Language is a fascinating, structured, multilayered system that people wield with skill and care, even when following rules unconsciously, even when they think they ‘only speak slang’. When a linguist uses the word ‘language’, we refer to the rules, structures and interrelated systems that humans use to communicate and interpret communication.
Syntax
Syntacticians study what you probably think of as ‘grammar’: why words go in the order they do, which words ‘belong’ together, why different languages put words together differently, and whether any true language universals exist.
Surprisingly few things are common to all human languages, but many things are common to many languages. Joseph Greenberg’s famous book Universals of Language proposed several universals, all ‘if-then’ universals. For example: ‘If a language has dominant order VSO in declarative sentences, it always puts interrogative words or phrases first in interrogative word questions; if it has dominant order SOV in declarative sentences, there is never such an invariant rule.’
Most languages build up meaning in nested chunks called phrases. Thus we have noun phrases, verb phrases, adjective phrases and so on. A noun phrase is a ‘head noun’ plus all words that ‘belong’ with that noun. The simplest test: see if you can replace the whole chunk with a pronoun like it.
The happy wombat ate lots of grass. → It ate lots of grass.
Since it cannot replace ‘happy wombat’ alone but can replace ‘the happy wombat’, this tells us that the happy wombat is a chunk (a noun phrase), while happy wombat is not complete. Noun phrases can act as grammatical subjects (the ‘doer’ of a verb) or grammatical objects (the ‘do-ee’ of a verb).
People tend to answer questions with full noun phrases rather than parts of noun phrases. Replying with just part of a noun phrase is something a child or non-native speaker might do, and not just in English. Humans seem to think of the world in chunks corresponding to full syntactic phrases.
Even in languages with extremely free word order, like many Australian languages including Warlpiri, words that we would consider part of a syntactic noun phrase normally appear together (see research on Warlpiri word order).
Syntacticians often use ‘phrase-structure trees’ to illustrate the chunking or nesting of syntactic structure. The fundamental design of a phrase structure tree indicates ‘phrase structure’: all words occurring under a single node such as S, VP or NP are constituents, meaning these words behave as a unit syntactically and often semantically. This nesting of phrases is considered critical to understanding language and human cognition.
Morphology
Morphology is the study of how words are made up of ‘bits of meaning’. Some languages, like English, can put two or more morphemes, or ‘bits of meaning’, into a single word. A simple example: the word cats is made up of CAT + PLURAL. The word feet is made up of FOOT + PLURAL.
Some changes are purely grammatical: with a third-person subject like she, verbs in English must end in -s in the present tense (she laughs versus I laugh). Other changes create new words: adding -ment to verbs like excite, enjoy, abash, fulfil creates nouns: excitement, enjoyment, abashment, fulfilment. Morphologists ask questions like: why don’t words like *surprisement, *watchment, *sleepment exist?
Semantics
Semantics is the study of meaning. As you may have guessed, there are many different types of meaning, so there are many branches of semantics. Semanticists study:
Logic, such as if-then statements. ‘If all roads lead to Rome, and this is a road, then it must lead to Rome’ is logically coherent and correct. Contrast with ‘If all roads lead to Rome, and this leads to Rome, then it must be a road’, which is logically incoherent and incorrect. Formal Semanticists study all the complicated ways statements can be True, False, or Unprovable.
Entailment: some words carry underlying assumptions that are true if the word is true. ‘When did you stop spending three hours on Reddit every day?’ can only be understood if it is in fact true that you have, at some point, spent three hours on Reddit every day. Contrast with ‘I tried spending three hours on Reddit every day’, where it is NOT necessarily true that I did spend three hours on Reddit every day.
Implicature: some words carry underlying assumptions that are not necessarily true, but heavily implied. ‘I used to spend three hours on Reddit every day’ implies, but does not entail, ‘three and only three’, because you can continue with something that defeats the implicature, like ‘and actually, some days I spent five or more hours on Reddit’.
Reference and Sense: There is a difference between a thing that exists and the representation of that thing in your head. A tree you can touch in the real world versus the representation you conjure when thinking of it. For many people, the mental representation includes a mental image, textures of bark and leaves, the smell of leaves after rain, shadows trees cast on solar panels, a tree with a tyre swing over a creek, the spelling ‘T-R-E-E’, and many other associations. There are interesting questions about words with the same referent, such as ‘The morning star is the evening star’, where both referents are the planet Venus, but there is somehow ‘more’ to the meaning than just the referent.
Truth: Some statements can have their truth verified through direct observation, such as ‘The carpet I am standing on is green’. Some can be verified through a firsthand account: you believe me when I tell you, so now you have secondhand verification. Some languages even force encoding of ‘evidentiality’: whether the speaker knows because they themselves saw or heard it, or heard it from someone else.
It is currently true to say that language models only have access to second- and thirdhand verification of truth, and cannot access firsthand verification.
Examples of evidentiality encoding in different languages:
English: If it were to rain entails that the speaker cannot personally verify whether ‘it will rain’ is true, though this is seen as archaic language in the 21st century.
Icelandic: Hún sagði hún væri hrifin af þér ‘She said she was-subjunctive enamoured of you’, where the subjunctive is part of grammatically encoding ‘she said it, but I cannot personally confirm the truth’.
Cherokee: wesa u-tlis-ʌʔi ‘A cat ran (I saw it running)’
Cherokee: u-wonis-eʔi ‘He spoke (someone told me)’ (Aikhenvald 2024 Evidentiality, p26)
Context and identity: Nontextual context refers to any relevant information not included in words spoken or written. It can be temperature, seeing someone walk into the room, the smell of baking bread, the frustration of asking your teenager the same question five times without getting a decent answer, accompanied by memories of how much they loved telling you about their day ten years ago. These multimodal meanings are part of our lived experience, stored in our brains as an intrinsic part of how we use and understand language.
Every time we speak, write, or interpret language, we apply our own beliefs and assumptions as underlying context. We all have an agenda each time we use language, including ‘being polite’, ‘being funny’, ‘demonstrating or pretending familiarity’, or just wanting to spend time with our interlocutor.
Phonetics
Phonetics, Phonology and Prosody are aspects of how we produce sounds which encode meaning. Phonetics is the study of sounds used in human languages and how these sounds are created. Phonology is the study of sound and sound changes within a single language. Prosody is the study of how we ‘sing’ our language, our intonation.
Phoneticians are often great fun at parties, not only for their ability to closely mimic different accents, but for their ability to teach others how to move their lips and tongue in new ways to mimic different accents too.
Pragmatics
Pragmatics is the study of implied meanings (and is therefore related to Semantics). Your friend saying ‘Do you have plans tonight?’ might really be asking ‘Do you want to go out together?’ or ‘Do you want to hang out at my place tonight?’ or ‘You should stop seeing that person you’re with because they’re bad for you’, or any number of other things, which the two of you know because of your shared history.
Discourse Analysis
Discourse Analysts study how conversation works in dyads, triads and larger groups. There are scores of fascinating questions to ask and find answers to.
How do you know when it’s your turn to talk? Either you self-select and just talk, or the current speaker selects you by looking at you expectantly. The least common way is for a non-speaker to select the next speaker using the same technique. This is a fun experiment to try at home, with friends, or at work: see if you can get a particular person to speak next without saying ‘You speak next’.
Experiment: Where do you look when talking with someone? It’s common NOT to look at your conversation partner until it’s time to swap speakers, or until you need to judge the listener’s reaction. Notice how power dynamics play into this, and how some people control the flow more than others. Remote work has changed some methods for controlling conversations, but one-on-one video chats display much of the same behaviour as in-person conversations.
If someone is telling an anecdote, how do you know which bits are supposed to be funny? Usually, the speaker laughs or otherwise lets listeners know what reaction they should give.
How do you know if someone wants sympathy or a solution? Usually, the speaker gives signals. Why do we get this wrong so often? Because we don’t all have the same ‘internal language’, so we each interpret signals from others, and nearly everything has more than one plausible interpretation, depending on backgrounds of speaker and listener. In studying discourse and conversation, context and participants’ personal backgrounds become critical.
Critical Discourse Analysis
Critical Discourse Analysis is the microscopic study of conversation and texts, evaluating the author’s background, the author’s intention in writing at a particular point in time, the intended target audience, and why the author selected that audience.
Critical Discourse Analysts look at vocabulary choice (why did the speaker say ‘The cake was very good’ and not ‘The cake was heaps good’?), syntactic choice (why did the speaker say ‘The woman was raped’ and not ‘Someone raped the woman’?), and other choices made consciously and subconsciously.
Critical Discourse Analysis reveals unspoken assumptions, beliefs and implications in a given text, grounding it firmly and concretely in the real world.
Language and Identity
One fascinating area of linguistics is studying how we use language to perform our identity. The words we choose, the accent we lay on, the sentence structures we prefer, the meanings we emphasise, and even conversational norms all signal who we are and how we see the world. We reaffirm our identity through our manipulation of language every time we speak.
In addition to using language to perform our own identities, our choices, assumptions and beliefs shape how we interpret others’ words. Do we listen ready to believe? Why? Is it because their accent or word choice tells us they are inherently trustworthy or untrustworthy? Do we listen ready to disagree because they use complicated grammatical structures that take effort to parse, or does your identity make you more inclined to agree with such speakers?
Langue vs Parole
The distinction between langue and parole refers to the difference between internal or potential capabilities versus external or performed capabilities. Langue refers to underlying, shared knowledge of a language, the rules and structures making language possible. Parole refers to actual language use as produced in real situations.
In natural language, humans are capable of producing utterances they will never produce, and also capable of producing utterances that fail to adhere to their own internal rules. This latter mismatch can be called a slip of the tongue, and is usually something the speaker themselves acknowledges as incorrect.
Different fields of linguistics rely more or less heavily on the langue or the parole. Chomsky famously described syntax as the study of ‘linguistic competence’, the rules of language, and specifically not of ‘linguistic performance’ with its production errors, arguing that linguistic theory should describe and explain possible grammatical structures, not everyday speech errors.
What Do Linguists Actually Study?
Linguists study anything to do with human language, human communication and human understanding, both internal structures and external performances. In the same way physiotherapists understand how muscles, joints and neural control interact to produce effective movement, linguists understand how language works at a structural level. We are experts of the systems involved in language and human communication.
What ‘Language’ Means in Large Language Models
Language models, large, tiny, and various sizes in-between, are built from complicated algorithms that have found patterns between ‘tokens’. Tokens are sequences of letters, numbers and other characters forming words or parts of words, stored as high-dimensionality vectors of numbers or groups of numbers, called embeddings.
Generative language models select from groups of likely tokens the next most-likely token in a sequence using probability distribution, and the result looks like human language.
As noted in this analysis of GPT and ‘wide AI’: ‘ChatGPT doesn’t understand language, grammar, or meaning. What it does “understand” is statistical relations between meaningless character-strings (tokens).’
The technology is amazing, but the strings produced by ‘language models’ are not language in the way a linguist would usually use the word. There is no intentionality, social identity or social obligations being upheld when generating tokens, nor the same syntactic, semantic, morphological or pragmatic structures humans use to build utterances.
Rather, as each token extends existing output, it is selected because, among training data, fine-tuning data and contextual (prompt) data, this token is in the small group identified as best fitting the preceding linear pattern. The result differs quite from the linguistics discussed above.
Syntax, Semantics and Language Models
Data scientists use ‘semantics’ and ‘syntax’ to describe generative language model output, but these differ quite from what these words mean when applied to human language.
Syntactic trees versus syntactic templates: The ‘syntactic templates’ learned by language models reflect recurring patterns in how parts of speech are sequenced in real-world text. These patterns are linear, rather than hierarchically nested as in human-generated sentences.
What is a ‘part-of-speech’? ‘Part-of-speech’ is the name given to categories of words humans use when speaking and writing: nouns, verbs, adjectives, determiners, adverbs. Sometimes quite specific categories are used: copulas (often corresponding in English to ‘be’ and ‘become’), serial verbs, auxiliary verbs (‘to be’, ‘to have’, ‘to get’), modal verbs, proper nouns, common nouns, count nouns, and so on.
For example, the part-of-speech template for ‘The happy wombat ate lots of grass’ is: Determiner, Adjective, Noun, Verb, Quantifier, Preposition, Noun. This differs fundamentally from the syntactic structure from a human perspective represented by phrase structure trees, where hierarchical relationships between words and meaningful groups are explicitly defined by nodes and branches, rather than emerging as a surface-level sequential pattern.
There is a fundamental difference between linear syntactic templates of language models and the two-dimensional hierarchical syntactic structure of natural language.
Semantic Representation, Truth and Context
The semantics that language models use are also ‘sensuously and operationally different’ (quoting Pullum’s analysis of Whorf 1940) from those of natural language as recognised by linguists.
LLM Semantics: The semantics of language models generally refers to groups of tokens that co-occur within a specific distance. Clusters of tokens with similar patterns are considered semantically related. The string ‘Paris is the capital of France’ and ‘Oslo is the capital of Norway’ end up semantically similar, causing ‘Paris’ and ‘Oslo’ to be seen as semantically similar because of co-occurrence with tokens representing ‘is the capital of’. ‘France’ and ‘Norway’ are seen as semantically similar for the same reasons.
Linguists’ Semantics: When linguists talk about semantics, starting with Russell, Frege and Wittgenstein in the 19th century, not only does ‘similarity’ matter, but, depending on the subfield, a range of factors including truth/false values, textual and nontextual context (physical, historical), and assumptions about reality are critically involved in ‘meaning’. Beyond word and sentence level, at language-in-interaction, meanings such as intention, social identity, and beliefs about interlocutors’ intentions, identities and beliefs also come into play.
The interplay of LLM Semantic and Syntax: As described in the paper Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models, because syntax and semantics of language models are contained in the same embeddings, non-human errors can arise in generated output. Their findings highlight the need to ensure syntactic diversity in training data within each semantic domain to reduce spurious correlations.
Language and Metaphor
With this understanding that ‘language model’ output looks like, but is not ‘language’, we turn to metaphors often used when talking about language model output, specifically semantic and cognitive metaphors such as ‘understand’, ‘the model thinks’, ‘the model hallucinates’, and so on. But firstly, why should we care?
It has been shown that metaphors shape our thoughts and opinions, as demonstrated in George Lakoff and Mark Johnson’s classic Metaphors We Live By, and Lakoff’s Women, Fire and Dangerous Things: What Categories Reveal about the Mind (about the Dyirbal noun class system). LOVE IS WAR means we ‘fight for our loved ones’. SICKNESS IS WAR means we ‘fight until the very end’, but sadly ‘lose the battle against cancer’. LIFE IS A JOURNEY and LIFE IS A COMPETITION produce quite different worldviews.
LLMs, Metaphors and Dangerous Things
Like language models, humans excel at finding patterns and using comparisons, which is why we are so fond of metaphors. But it behoves us to be aware of the metaphors we use and their implications, because they can lead us in both useful and potentially harmful directions.
Using the metaphor LLMS ARE SENTIENT (LIKE US) leads to expressions such as ‘ChatGPT understands…’, ‘Claude reasons through problems’, ‘Copilot can offer to help you’, all implying LLMs are a lot like us, offering thoughtful responses. These phrases hide, or obfuscate, the actual processes of generative language models, which is a series of probabilistic calculations moderated by other probabilistic and hard-coded restrictions we call ‘guardrails’ (another metaphor: maybe LLMS ARE AN AMUSEMENT PARK RIDE, LLMS ARE A SAFE FOOTPATH, or from the other side: LLMS ARE A STEEP CLIFF). In the physical world, ‘guardrails’ really just prevent accidental falls; they do not protect against deliberate breaches.
By using the LLMS ARE SENTIENT metaphor, it becomes easy to overlook actual processes, and the framing can make it difficult to identify how we can update models to produce output closer to what we expect and require.
While there are many specific implementations, some fundamental components are common: language model output is generated based on high-dimensionality vectors and probabilistic combinations, created via original training data and fine-tuning, with guardrails monitoring final output. Therefore, these are points that can be adjusted.
One glaring problem with the LLMS ARE SENTIENT LIKE US metaphor is that it implies the ‘language’ language models produce is like ours. But it really isn’t. Language models do not use human syntactic structures, nor do semantic representations closely resemble human semantic representations except at the most basic level. Language model semantics are not rooted in human-defined categories the same way ours are, not rooted in identity and performing a personal agenda, nor shaped by personal goals or lived experience.
While this seems like an interesting philosophical discussion, it has practical implications. Misleading metaphors can obscure where actual control over a language model lies: through changes to training data, modifications to fine-tuning datasets, adjustments to system or task prompts, and design of guardrails constraining output. Recognising these intervention points allows developers and users to shape model output more effectively, without relying on the misleading notion that the model ‘understands’ or ‘intends’ like a human.
For a great explanation of why the LLMS ARE SENTIENT LIKE US metaphor is problematic, and why ‘hallucination’ is misleading, see the section ‘The hidden message when we say something is an “error”‘ in this article on the hidden meaning of ChatGPT’s errors.
Where Language Stops and Human Identity Begins
Throughout this essay, I have tried to show that language, as linguists study it and humans use it for communication, is embedded in identity and the ethnographic performance of that identity. The choices we make when speaking and writing signal our beliefs, stances, assumptions and judgements. The language we use reflects semantic and syntactic categories our experiences have led us to create, as well as social and cultural contexts in which we participate.
The fields of linguistics that have arisen over the past century give even a casual observer insight into aspects of language that are intrinsically human: implied meanings in pragmatics, social dynamics of turn-taking and interaction in discourse analysis, truth conditions and concept integration in semantics, units of meaning in morphology and syntax, and overlaps with psychology, sociology and biology in psycholinguistics, sociolinguistics and neurolinguistics — the very structures that natural language processing research draws on.
This linguistic conceptualisation reveals a much more complete and interlinked interpretation compared with generative language models, where produced ‘language’ results from statistically probable continuations. A language model cannot change its mind halfway through a sentence, as it does not have a mind. It cannot produce false starts where it does not ‘know’ how to continue, because its algorithm states that it keeps producing tokens until reaching an ‘END’ token.
Why Does Any of This Matter?
The goal of this essay is to spark reflection and discussion concerning what ‘language’ means and how it means different things in different contexts. My primary goal is to get readers to examine not only how they think about language, but also how they think about the process language models use to generate output.
My secondary goal is to support the growing number of people trying to change how we talk about processing and output of generative language models, moving away from misleading metaphors toward more precise, actionable language.
With a clearer conceptual toolkit for thinking and talking about LLMs, it becomes easier to identify where interventions can be made, through training data, fine-tuning, system prompts, or output constraints, and to guide these models toward output that better meets human expectations.
Ultimately, a more accurate understanding of both human and machine language can only improve the reliability, safety, and usefulness of AI systems.




