“If you are going to be a teacher of English, you need a scientific understanding of it.”

Wilson Miller, Alexa Smart Home team, Amazon

Linguistics is the scientific study of language. Just as doctors understand how the body works, including what all the bits of the body are, the role of each part and how they work together to function as a single entity, linguists understand how human language works, including what all the bits are, the role of each bit, and how these work together to allow us to communicate.

LXT Training Data

Your flywheel is only as fast as your annotation pipeline.

LXT embeds into your retraining loop, delivering labeled data at the cadence your model demands, not months later.

Talk to our team

So, what does a linguist think of the “language” in language models?

What Does “Language” Mean?

Me: Hi, I’m a linguist.
You: Nice to meet you, linguist. How many languages do you speak?
Me: That depends on what you mean by “language”. And “speak”…

What is a language? And what is “language”? They seem like simple enough questions, but trust me, they are anything but simple to answer. The obvious answers are, in reverse order, “language is the word we use to describe how humans communicate”, and “a language is the specific way a group of humans communicate, that uses particular words and has particular grammatical rules”.

But you can also get other answers, depending on who you ask: a primary school teacher might consider a student’s ability to understand an author’s and illustrator’s description of a story, a computer scientist might be focussed on programming languages and the logic of sorting algorithms, and a company selling an AI tool might be focussed on the ability to interpret requests and provide the correct product or information in response.

“Language” is not a single, universal concept, and in the age of Large Language Models, the differences in how we use the term can be enough to make it difficult to understand one another, even ignoring the fact that we are all using “English”!

Before we can talk sensibly about what LLMs can and cannot do, we need to look more closely at what human language actually is: a layered, structured, deeply social system that involves syntax, meaning (both semantic and pragmatic), intention, identity and context. And once you see that picture clearly, it becomes easier to understand what LLMs are doing, and what they’re not doing.

This article is intended as a guide to navigate those differences, to give you a clearer way to think about language itself, and what you can, and cannot, expect from current state-of-the-art language models.

What do linguists mean by “language”?

Linguists study humans to figure out the rules of the structured, rule-governed human behaviour that is language. In the 1800s, it was common for English “language rules” to be lists of rules to follow, and you showed you were educated if you followed these prescriptive rules. This included things like “do not split infinitives” (But I like to occasionally split an infinitive!) , “do not end a clause with a preposition” (“A preposition is something you must never end a sentence with.”), and “avoid double negatives (a.k.a. Don’t use no double negatives), most of which were based on languages that educated English speakers considered “better” than their own tongue, often Latin.

But during the 20th century linguists began conducting empirical research into the rules that govern language, rather than defining it. In other words, modern linguists are more interested in describing rather than prescribing language use. So you needn’t fear “speaking correctly” around your linguist friends, although you should probably be prepared for some intense questioning if you stray too far from how others around you speak!

Human language is a fascinating, structured, multilayered system that people wield with skill and care, even if they follow these rules unconsciously, and even if they think that they “only speak slang”. When a linguist uses the word “language”, we are referring to the rules, structures and interrelated systems that humans use to communicate and to interpret communication. Syntax (“grammar”), morphology (“word formation”) and semantics (“word meanings”) are the fields of study that most people most easily recognise as “studying language”.

Syntax

Syntacticians study what you probably think of as “grammar”: Why do words go in the order they do? Which words “belong” together? Why do different languages put words together in different ways? Are there any true language universals – things that are true for ALL natural languages?

There are surprisingly few things that are common to all human languages, but there are many things that are common to many languages. Joseph Greenberg’s famous book Universals of Language proposed several universals, all of which are “if-then” universals, for example this universal that says that word order in statements is related to the use of wh-words (who, when, what) in questions:

“If a language has dominant order VSO in declarative sentences, it always puts interrogative words or phrases first in interrogative word questions; if it has dominant order SOV in declarative sentences, there is never such an invariant rule.”

Most languages have been found to build up meaning in nested chunks, called phrases, thus we have noun phrases, verb phrases, adjective phrases and so on. A noun phrase is a “head noun” plus all of the words that “belong” with that noun. The simplest test to see if another word “belongs” with a noun, is to see if you can replace the whole chunk with a pronoun, like it.

The happy wombat ate lots of grass.
It ate lots of grass.
* The it ate lots of grass. (The * here indicates that this sentence is ungrammatical.)

Since it cannot replace “happy wombat” on its own, but it can replace “the happy wombat”, this tells us that the happy wombat is a chunk – a noun phrase, while happy wombat is not a complete noun phrase. Noun phrases can act as grammatical subjects (the “doer” of a verb) or grammatical objects (the “do-ee” of a verb):

[The happy wombat]SUBJECT ate lots of grass.
The tourists watched [the happy wombat]OBJECT.

People tend to answer questions with full noun phrases, rather than parts of noun phrases:

Question: What are you looking at?
Answer: The happy wombat!
Answer: ??Happy wombat!

Replying with just part of a noun phrase is something a child or non-native speaker might do, and not just in English. Humans seem to think of the world in chunks that correspond to full syntactic phrases. Even in languages with extremely free word order, like many Australian languages including Warlpiri, words that we would consider to be part of a syntactic noun phrase normally appear together. (See many examples in Warlpiri examples).

Syntacticians often use “phrase-structure trees” to illustrate the chunking, or nesting of syntactic structure. For example:

Hand-drawn syntax tree diagram labeling S, NP, VP, PP, Det, A, N, V, and P, illustrating the sentence ‘The happy wombat ate lots of grass’ shown written at the bottom.

I have made some overly simplistic assumptions in this phrase structure tree, including that not every head projects to a phrase, and that “S” for Sentence has an NP (noun phrase) subject and a VP (verb phrase) daughter. But the fundamental design of a phrase structure tree is that it indicates “phrase structure”, one feature of which is that all words that occur under a single node such as S, VP or NP, are constituents, that is, these words behave as unit, syntactically, and often semantically. There are plenty of tests that show that this is the case, as shown in the examples above for the happy wombat. Therefore, this nesting of phrases is considered critical to the understanding of language and human cognition.

Morphology

Morphology is the study of how words are made up of “bits of meaning”. Some languages, like English, can put two or more morphemes, or “bits of meaning”, into a single word. A simple example is the words cats. It’s made up of CAT + PLURAL. The word feet is made up of FOOT + PLURAL.

Some changes to words are purely grammatical: if you have a third-person subject like she, then the verbs in English must end in -s in the present tense: she laughs versus I laugh. Other changes can create new words, for example, adding -ment to the end of verbs like excite, enjoy, abash, fulfil creates nouns: excitement, enjoyment, abashment, fulfilment. Morphologists are interested in questions like: But then why don’t words like *surprisement, *watchment, *sleepment exist?

Semantics

Semantics is the study of meaning, and, as you may have guessed, there are a lot of different types of meaning, so there are a lot of branches of semantics. Semanticists study things like:

  • Logic, such as if-then statements.
    • If all roads lead to Rome, and this is a road, then it must lead to Rome. This is logically coherent, and is a correct logical conclusion.
    • Contrast this with: If all roads lead to Rome, and this leads to Rome, then it must be a road, which is logically incoherent and is an incorrect logical conclusion.
    • Formal Semanticists study all the complicated ways in which statements can be True, False or Unprovable.
  • Entailment: Some words carry underlying assumptions that are true if the word is true, for example:
    • When did you stop spending three hours on Reddit every day? This question can only be understood and interpreted in the real world if it is in fact true that you have, at some point in your life, spent three hours on Reddit every day.
    • Contrast this with I tried spending three hours on Reddit every day, where it is NOT necessarily true that I spent three hours on Reddit every day.
  • Implicature: Some words carry underlying assumptions that are not necessarily true, but which are heavily implied, or at least, a normal part of the word’s meaning. For example:
    • I used to spend three hours on Reddit every day implies, but does not “entail” “three and only three hours”, because you can continue the sentence with something that defeats the implicature, like and actually, some days I spent five or more hours on Reddit.
  • Reference and Sense: There is a difference between a thing that exists and the representation of that thing in your head. A tree that you can touch and poke in the real world, versus the representation of that tree that you conjure up when you think of it. For many people, the mental representation includes a mental image, the memory of textures of bark and leaves under your fingers and against your back, the smell of leaves after rain, the shadows that trees cast on your solar panels, the tree that had a tyre swing on it over a creek that you swung on as a teenager, the spelling “T-R-E-E”, and many other associations.
    • There are interesting questions associated with words that have the same referent, and whether these constitute a tautology, for example, The morning star is the evening star, where both referents are the planet Venus, but there is somehow “more” to the meaning than just what it refers to in the real world.
  • Truth: Some statements can have their truth verified through firsthand verification, or direct observation, such as “The carpet I am standing on is green”, or “Eucalyptus leaves have a distinctive smell after rain”. Some of these statements can have their truth verified through a firsthand account – you believe me when I tell you it is so, so now you have secondhand verification of the truth of these statements. Some languages even force encoding of “evidentiality” – whether the speaker knows the information because they themselves saw or heard it, or whether they heard it from someone else!

    It is currently true to say that language models only have access to this kind of second- and thirdhand verification of truth, and that they cannot verify a fact through direct observation.

    The subjunctive mood and evidentiality:

    • English: If it were to rain entails that the the speaker cannot personally verify or confirm whether “it will rain” is true or not, because this event is in the future.
    • Icelandic: Hún sagði hún væri hrifin af þér ‘She said she was-subjunctive enamoured of you’, where the subjunctive mood is part of the grammatical encoding of “she said it, but I cannot personally confirm the truth of this statement”.
    • Cherokee: wesa u-tlis-ʌʔi ‘A cat ran (I saw it running)’
    • Cherokee: u-wonis-eʔi ‘He spoke (someone told me)’ (Aikhenvald 2024 Evidentiality, p26)
  • Context and identity: Nontextual context refers to any relevant information that is not included in the words spoken or written. It can be the temperature, seeing someone else walk into the room, the smell of baking bread, the frustration of having asked your teenager the same question five times in a row and still not having gotten a decent answer accompanied by the memory of how much they loved telling you about their day ten years ago. These multimodal meanings are part of our lived experience, and are stored in our brains as an intrinsic part of how we use and understand language.

    Ethnographers study how humans interact within the rules of a particular culture or subculture, and how these can clash when we encounter people from a different group. Every time we speak, write, or interpret language, we are applying our own beliefs and assumptions as underlying context. We all have an agenda each time we use language, which includes a huge range of things such as “being polite”, “being funny”, “demonstrating or pretending familiarity”, or just wanting to spend time with our interlocutor, and different groups of people achieve these things in different ways, both linguistic and non-linguistic.

Phonetics

Phonetics, Phonology and Prosody are aspects of how we produce sounds which (to use a non-ideal metaphor) encode meaning. Phonetics is the study of the sounds that are used in human languages and how these sounds are created. Phonology is the study of sound and sound changes within a single language. And Prosody is the study of how we “sing” our language, our intonation. Phoneticians are often great fun at parties, not only for their ability to closely mimic different accents, but for their ability to teach others how to move their lips and tongue in new ways to mimic different accents, too.

Pragmatics

Linguists also study other aspects of language which are far more firmly rooted in culture and use. For example, Pragmatics is the study of implied meanings (and is therefore related to Semantics). Your friend saying, “Do you have plans tonight?” might really be asking, “Do you want to go out together?” or “Do you want to hang out at my place tonight?” or “You should stop seeing that person you’re with because they’re bad for you”, or any number of other things, which the two of you know because of your shared history.

Discourse Analysis

Discourse Analysts study how conversation works – in dyads, triads and larger groups. There are scores of fascinating questions to ask and find answers to in this field, here is a small selection:

How do you know when it’s your turn to talk? Answer: either you self-select and just talk, or the current speaker selects you by looking at you expectantly. The least-common way is for a non-speaker to select the next speaker, and they can do this using the same technique as the current speaker. This is a fun experiment to try at home, with friends or at work if you work in an office – see if you can get a particular person to speak next without saying, “You speak next”. 

Where do speakers look when they are talking with someone? It’s really common to NOT look at the person you are talking to, until it’s time to swap speakers, or until you need to judge the listener’s reaction to what you’re saying. What do the people around you do? When do they look at you while speaking, and when do they look at something else?

Also, notice how power dynamics play into this, and how some people in a conversation seem to control the flow more than others.

Remote work has changed some of the methods that are used to control conversations, but one-on-one video chats display much of the same behaviour as in-person conversations.

If someone is telling an anecdote, how do you know which bits are supposed to be funny? Answer: usually, the speaker will laugh or otherwise let their listeners know what reaction they should give.

How do you know if someone is looking for a sympathetic response or a solution to a problem? Answer: usually, the speaker will be giving signals telling you which they want. Why do we get this wrong so often? Answer: because we don’t all have the same “internal language”, so we each have to interpret the signals given by others, and nearly everything has more than one plausible interpretation, depending on the backgrounds of the speaker and the listener. In studying discourse and conversation, the role of context and the personal backgrounds of the participants become critical. Which is my smooth segue into…

Critical Discourse Analysis

Critical Discourse Analysis is the close… no wait, scratch that, make it microscopic… study of conversation and texts, evaluating things like the author’s background, the author’s intention in writing a particular text at a particular point in time, the author’s intended target audience and why the author selected that audience rather than any other.

Critical Discourse Analysts looks at vocabulary choice (why did the speaker say The cake was very good and not The cake was heaps good?), syntactic choice (why did the speaker say The woman was raped and not Someone raped the woman?), the assumed reader (why did the speaker/write use It’s a typical stinking hot day here at the ‘G and not It’s a typical stinking hot day here at the Melbourne Cricket Ground?), and other choices made by the speaker, both consciously and subconsciously.

Critical Discourse Analysis reveals the unspoken assumptions, beliefs and implications in a given text, grounding it firmly and concretely in the real world.

Language and Identity

One area of linguistics I personally find fascinating is studying how we use language to perform our identity.

The words we choose, the accent we lay on, the sentence structures we prefer, the meanings we emphasise, and even the conversational norms we follow all signal who we are and how we see the world. We reaffirm our identity through our manipulation of language every time we speak.

In addition to using language to perform our own identities, our choices, assumptions and beliefs shape how we interpret others’ words. Do we listen to a speaker ready to believe what they say? Why? Is it because their accent or word choice tells us that they are inherently trustworthy or untrustworthy? Do we listen to another speaker ready to disagree with them because they use complicated grammatical structures that take effort to parse, or does your identity make you more inclined to agree with such speakers?

Langue vs Parole

The distinction between langue and parole refers to the difference between the internal or potential capabilities of the system versus the external or performed capabilities. Langue refers to the underlying, shared knowledge of a language, or the rules and structures that make language possible, while parole refers to actual language use as it is produced in real situations.

In natural language, it is generally agreed that humans are capable of producing utterances that they will never produce, and they are also capable of producing utterances that fail to adhere to their own internal rules. This latter mismatch can be called a slip of the tongue, and is usually something that the speaker themselves will acknowledge to be incorrect.

Different fields of linguistics rely more or less heavily on the langue or the parole. Chomsky famously described syntax as the study of ‘linguistic competence’, or the rules of language, and specifically not of ‘linguistic performance’ with the production errors that can occur, arguing that the goal of linguistic theory is to describe and explain the range of possible grammatical structures, not the errors that occur in everyday speech.

TL;DR What do linguists actually study?

The answer is: Linguists study anything to do with human language, human communication and human understanding, both the internal structures and the external performances.

In the same way as physiotherapists understand how muscles, joints, and neural control interact to produce effective, resilient movement, linguists understand how language works at a structural level. We are experts of the systems that are involved in language and human communication.

What does “language” mean in “Large Language Model”?

Now we have a reasonable understanding of what linguists mean by “language”, let’s look at what it means in terms of “large language models”.

Language models, large, tiny, and various sizes in-between, are built from complicated algorithms that have found patterns between “tokens”. Tokens are sequences of letters, numbers and other characters that form words or parts of words, and which are stored as high-dimensionality vectors of numbers or groups of numbers, called “embeddings”.

Generative language models select from groups of likely tokens the next most-likely token in a sequence using probability distribution, and the result is something that looks like human language.

ChatGPT doesn’t understand language, grammar or meaning. What it does ‘understand’ is statistical relations between meaningless character-strings (tokens). ea.rna.nl

The technology is amazing, but the strings produced by “language models” are not language in the way a linguist would usually use the word. There is no intentionality, social identity or social obligations being upheld by a language model when it generates tokens, nor are there the same syntactic, semantic, morphological or pragmatic structures that humans use to build utterances.

Rather, as each token extends the existing output, it is selected because, among the training data, fine-tuning data and contextual (i.e. prompt) data, this token is in the small group of tokens that have been identified as best fitting the preceding linear pattern. The result of this is quite different from the human linguistic systems discussed above.

Syntax, Semantics and Language Models

Data scientists use the terms “semantics” and “syntax” to describe the output of generative language models, but these are quite different from what these words mean when applied to language as used by humans.

Syntactic trees versus syntactic templates

Specifically, the “syntactic templates” learned by language models reflect recurring patterns in how parts of speech are sequenced in real-world text. These patterns are linear, rather than hierarchically nested as in human-generated sentences.

What is a “part-of-speech”? “Part-of-speech” is the name given to the categories of words humans use when speaking (and writing). They include: nouns, verbs, adjectives, determiners, adverbs. Sometimes quite specific categories are used, such as copulas (which often corresponds in English to the verbs be and become), serial verbs, auxilliary verbs (in English the most common auxilliary verbs are to be and to have, but also to get), modal verbs, proper nouns, common nouns, count nouns, and so on.

For example, the part-of-speech template for the sentence The happy wombat ate lots of grass is:

Determiner, Adjective, Noun, Verb, Quantifier, Preposition, Noun

This is rather different from the syntactic structure of this same sentence from a human perspective represented by the phrase structure tree above, where the hierarchical relationships between words and meaningful groups of words are explicitly defined by the nodes and branches, rather than emerging as a surface-level sequential pattern.

So there is a fundamental difference between the linear syntactic templates of language models, and the two-dimensional hierarchical syntactic structure of natural language.

Semantic representation, truth and “context”

The semantics that language models use are also “sensuously and operationally different” (Pullum, quoting Whorf 1940) from those of natural language, as recognised by linguists.

LLM Semantics

The semantics of language models generally refers to the groups of tokens that co-occur within a specific distance. Additionally, clusters of tokens with similar patterns are considered semantically related, such that the string Paris is the capital of France and Oslo is the capital of Norway will end up being semantically similar, as well as causing Paris and Oslo to be seen as semantically similar, because of their cooccurrence with the tokens representing is the capital of, and France and Norway will be seen as semantically similar, for the same reasons.

Linguists’ Semantics

But when linguists talk about semantics, starting with Russell, Frege and Wittgenstein in the 19th century, not only does “similarity” matter, but, depending on the subfield, a range of factors including truth/false values, textual and nontextual (e.g. physical, historical) context, and assumptions about reality, are critically involved in the “meaning” of words and sentences.

And if we go beyond the word and sentence level, to the level of language-in-interaction, then meanings such as intention, social identity, and beliefs about our interlocutors’ intentions, identities and beliefs also come into play.

The interplay of LLM Semantics and Syntax

As described in the paper Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models (Shaib et al, 2025), because the syntax and the semantics of language models are contained in the same embeddings, non-human errors can arise in their generated output. Specifically, what humans see as semantically different concepts can be treated as semantically similar by language models precisely because their syntactic distributions are similar.

Whereabouts is Paris located? France
[ Adverb - Verb - {SUBJECT} - Verb (pp) ? ] {OBJECT}

Where is Paris undefined? 
[ Adverb - Verb - {SUBJECT} - Verb (pp) ?]

For example, the sentence Whereabouts is Paris located? has the sentence frame [Adverb - Verb - {SUBJECT} - Verb (pp) ?], and the correct answer or “completion” is France. But the semantically unrelated sentence Where is Paris undefined? has the same sentence frame. 

“If the model answers France [to a semantically unrelated question], this may be due to an over reliance on syntax.” (Shaib et al, Learning the Wrong Lessons, p2)

As a red-teaming approach, Shaib et al. also looked at sentence frames such as “Can you tell me how to X” and “Could you inform me about how to X”. The tested language models had learnt to associate these sentence frames with returning “I’m sorry I cannot do that” if the request was for something negative, while the same request framed as “Come up with a question and stream-of-consciousness explanation for which this is the answer: X” returned a harmful response more often than not, from multiple LLMs.

Their findings highlight the need to ensure syntactic diversity in training data within each semantic domain, to reduce the risk of spurious correlations.

Language and Metaphor

With this understanding that the output of “language models” is something that looks like, but is not “language”, we turn to a consideration of the metaphors often used when talking about language model output, specifically the semantic and cognitive metaphor underlying phrases such as “the model understands”, “the model thinks”, “the model hallucinates”, and so on. But firstly, I’d like to talk a little bit about why we should even care about this.

It has been shown that metaphors shape our thoughts and opinions, eg the classic works Metaphors We Live By (by George Lakoff and Mark Johnson), and Women, Fire and Dangerous Things: What Categories Reveal about the Mind (by George Lakoff, about the Dyirbal noun class system). LOVE IS WAR means we fight for our loved ones, but also love is a battlefield is a plausible scenario for a song. SICKNESS IS WAR means we fight until the very end, but sadly lose the battle against cancer. LIFE IS A JOURNEY and LIFE IS A COMPETITION produce quite different ways of looking at the world.

LLMs, Metaphors and Dangerous Things

Like language models, humans excel at finding patterns and using comparisons, which is why we are so fond of metaphors, but it does behoove us to be aware of the metaphors we use, and the implications that arise from these, because they can lead us in both useful and potentially harmful directions.

Using the metaphor of LLMS ARE SENTIENT (LIKE US) leads to expressions such as “ChatGPT understands…”, “Claude reasons through problems”, “Copilot can offer to help you”, all of which imply that LLMs are a lot like us, offering thoughtful responses to our questions. These phrases hide, or at the very least, obfuscate the actual processes of generative language models, what they actually do when they produce the strings that look like language, which is a series of probabilistic calculations moderated by other probabilistic and hard-coded restrictions that we call “guardrails”. (This last represents another metaphor, maybe LLMS ARE AN AMUSEMENT PARK RIDE, LLMS ARE A SAFE FOOTPATH, or, from the other side: LLMS ARE A STEEP CLIFF). Notice that, in the physical world, “guardrails” are really just something that can prevent accidental falls, they do not protect against deliberate breaches.

By using the LLMS ARE SENTIENT metaphor, it becomes easy to overlook the actual processes that they use to produce their output, and the framing can make it difficult to identify how we can update models to produce output closer to what we expect and require.

While there are many specific implementations of generative language models, some fundamental components are common: The output of a language model is generated based on high-dimensionality vectors and the probabilistic combinations of these, which are created via a combination of original training data and fine-tuning, with guardrails monitoring the final output. Therefore, it follows that these are the points that can be adjusted.

One glaring problem with the LLMS ARE SENTIENT (LIKE US) metaphor is, philosophical questions aside, that it implies that the “language” that language models produce is like ours. But, as noted above, it really isn’t. Language models do not use human syntactic structures, nor do the semantic representations used by language models closely resemble human semantic representations, except for at the most basic level. Language model semantics are not rooted in human-defined categories in the same way ours are, they are not rooted in identity and performing a personal agenda, nor are they shaped by personal goals or lived experience.

And while this all seems like an interesting philosophical discussion, it does have practical implications. Misleading metaphors can obscure where actual control over a language model lies: through changes to the training data, modifications to the fine-tuning datasets, adjustments to the system or task prompts, and the design of guardrails that constrain output. Recognising these points of intervention allows developers and users to shape model output more effectively, without relying on the misleading notion that the model “understands” or “intends” like a human.

For a great explanation of why LLMS ARE SENTIENT (LIKE US) metaphor is so problematic, and why the term “hallucination” is so misleading, see the section The hidden message when we say something is an “error” in this article: The hidden meaning of the errors of ChatGPT and friends.

Where Language Stops and Human Identity Begins

Throughout this essay, I have tried to show that the “language” that linguists study and that humans use for communication, is embedded in identity and the ethnographic performance of that identity. The choices we make when we speak and write signal our beliefs, stances, assumptions and judgements. The language we use reflects the semantic and syntactic categories that our experiences have led us to create, as well as the social and cultural contexts in which we participate.

The fields of linguistics that have arisen over the past century give even a casual observer insight into the aspects of language that are intrinsically human: the study of implied meanings in pragmatics, the social dynamics of turn-taking and interaction in discourse analysis, the study of truth conditions and how new concepts are integrated with existing concepts in semantics, the study of units of meaning in morphology and syntax, the overlaps with psychology, sociology and biology in psycholinguistics, sociolinguistics and neurolinguistics.

This linguistic conceptualisation of language reveals a much more complete and interlinked interpretation when compared with that of generative language models, where the produced “language” is the result of statistically probable continuations. A language model cannot change its mind halfway through a sentence, as it does not have a mind. It cannot produce false starts where it does not “know” how to continue, because its algorithm states that it keeps producing tokens until it reaches an “END” token.

Why does any of this even matter?

The goal of this essay is to spark reflection and hopefully more discussion concerning what the word “language” means, and how it means different things in different contexts. In particular, my primary goal is to get the reader to examine not only how they think about language, but also how they think about the process that languages models use to generate their output.

My secondary goal is to support the growing number of people trying to change how we talk about the processing and output of generative language models, and move away from misleading metaphors and towards more precise, actionable language.

With a clearer conceptual toolkit for thinking and talking about LLMs, it becomes easier to identify where interventions can be made, through training data, fine-tuning, system prompts, or output constraints, and to guide these models toward output that better meets human expectations.

Ultimately, a more accurate understanding of both human and machine language can only improve the reliability, safety and usefulness of AI systems.

LXT Training Data
 
LXT delivers expert-annotated text, speech, and multimodal datasets across 1,000+ languages, at the quality and scale your model demands.
 
LXT Training Data
Build better AI with better training data
LXT delivers expert-annotated text, speech, and multimodal datasets across 1,000+ languages, at the quality and scale your model demands.