Helping Hands
What Is a Token? AI Basics Explained Simply
What is a token? How AI models break text, images and sound into tokens, what a context window is, and why understanding isn't the same as generating.

ChatGPT, Claude and the like seem like true all-rounders at first glance. They answer questions, write texts, summarise documents and, on top of that, operate other programs quietly in the background.
But how exactly does a language model actually do this? What basic principle is behind it – and why can AI models now even create images or generate audio?
In this article, we explain the basic principle behind AI models: tokens. What they are, how a model works with them, and why it can process not only text but also images and spoken language. Anyone who knows these basics will also have a much better understanding of what AI can achieve in customer service – and where its limits lie.
This is Part 1 of our fundamentals series, created in collaboration with our Machine Learning Engineer Dr. Tae-Gil Noh.
Key Takeaways
- One principle: An AI model calculates which token comes next. Every capability we see emerges from this single task.
- Tokens: A model reads neither whole words nor individual letters, but word fragments. “Working”, for example, breaks down into “work” and “ing”.
- Context window: A model can only process a limited amount at once. That’s why the right knowledge for each question has to be specifically retrieved and supplied.
- More than text: Images and sound can also be converted into token sequences. This lets a model “see” and “hear”, even though it is really only reading.
- Understanding is not generating: What a model can take in and what it can produce itself are two different things. Phone calls require models that can understand and speak – in real time.
Initialisation: “Take a scenario and spin it further”
What exactly are you doing when you communicate with a language model? You ask a question… or do you? Essentially, you’re doing something else: you give the model a situation or scenario and expect it to “spin it further” – in other words, to continue it.
Before it answers, the model reads what you’ve given it. That can be instructions, a customer message, or any kind of knowledge in the form of attached images or PDFs – all in one and the same message. You can think of this as a kind of initialisation: the model takes this scenario and brings it into a unified state so it can answer the question: “What comes next?” Then the model produces the continuation of this thought piece by piece. True to the motto: take a deep breath in and out, step by step.
Text, audio or image — everything is translated into the same sequence of tokens before the model reads it.
A model can only take in a limited amount at once. That is why the right knowledge has to be found and handed over with every question — that is RAG
Why this continuation is not just plausible, but also useful
The model learned by reading an enormous amount of human-written text: large parts of the internet, books, documents, conversations. It didn’t memorise them, but rather “absorbed” them as patterns. And one pattern dominates human texts: a question is followed by an answer, a request is followed by its execution, half a sentence is followed by a fitting ending. So when the model looks for the most natural continuation of a customer question, the most natural thing – based on everything it has read – is a good answer. The usefulness comes from the principle of “just continue the text”, because in the texts it learned from, a question was followed by a helpful continuation.
Reading alone isn’t enough, though. It makes a model plausible, but not helpful, honest or safe. That only comes with a second training phase: people show the model which continuations are actually good, and it learns to prefer them. One sentence sums it up:
First read everything, then learn what’s good.Dr. Tae-Gil Noh, Machine Learning Engineer at OMQ
With every model generation, both aspects improve.
What is a token?
These “bits” or “pieces” are called token. In a text, a token is usually a piece smaller than a word. Some words are token in their own right; other, longer words get split. The English word “working”, for example, becomes “work” + “ing”. Spaces and punctuation marks are tokens too.
A token is usually smaller than a word: ‘confirmation’”’ splits into two fragments, and the spaces and the full stop are tokens of their own.
But why do words need to be split up instead of being used whole?
Tokenisation allows the language model to handle anything: every language, meanings, topics, spelling mistakes – and even new product names, which the model assembles from pieces it already knows.
A language model does not answer a question. It calculates which token is most likely to come next — step by step.
The downside is that the model doesn’t perceive words and letters the way we humans do. Language models see, so to speak, a stream of tokens. That’s why many language models also struggle to count the letters in a word – they simply don’t see them.
Language models see a stream of tokens, not letters. That is why they stumble over questions we find trivial.
In the beginning was the word
All the language models we know and use today started with words – more precisely, with enormous amounts of text data. And text is a medium that is especially interesting for at least two reasons. Text contains content (meaning: “Your order was shipped yesterday”), but also instructions (“Open order #12345”).
A model that speaks “language” fluently can both pass on content and give instructions – and thereby use tools. The model writes: “Give me the shipping information for order #12345”, this instruction is carried out, and the result flows back into the conversation as a new sequence of tokens. So “text” is both the meaning and the control system!
The token as an element in an ordered sequence
If you look at the token more closely, it isn’t really about “text” at all. A token is simply an element in an ordered sequence that the model reads. Which raises the question: what else can be used as an element in an ordered sequence?
- Replace “text” with “sound”, and the model can “hear”.
- Replace it with “image”, and it can “see”.
Sound is a wave: air pressure moving up and down over time. And “over time” is the key point here – it means sound has a built-in order, from beginning to end.
When you give an AI model “sound”, the recording has to be split into very small segments – really small ones – around a hundred segments per second. Each of these segments is then converted into a token that captures what the sound is doing at that moment. String these tokens together and you get the complete expression (the “utterance”) as a sequence – which is exactly the form a model can read. So the model “hears” by reading the sound as a sentence.
Sound already runs from beginning to end. Cut it into roughly a hundred segments per second and it becomes a sequence of tokens — so the model hears by reading.
Images are a far more interesting case, because they work differently from “text” and “sound”. An image has no natural sequence and no natural order. Where does an image begin? Top left? In the middle? There’s no first word and no reading direction dictated by the image itself.
So how is it possible that AI-generated images and other AI tools for image editing exist?
The model makes up an order! It lays a grid over the image, cuts it into a mosaic of small “tiles” and reads these pieces in a fixed order – while also storing exactly where each “tile” was originally located (top left, in the middle, …). Each of these “tiles” becomes, in a sense, a token*; the position data makes it possible to reconstruct the spatial relationships. An image becomes a sequence because the model defines it as one and records the “coordinates”.
An image gives no reading direction. The model lays a grid over it, reads the tiles in a fixed order and remembers where each one sat.
* An image “token” is a collection of numbers describing a “tile”. As described above, these are turned into a unified sequence that is then made readable. Strictly speaking, however, this isn’t a “token” in the proper sense – only the mechanism is similar. The term is used here as a simplified illustration.
Reading and writing are different skills
The model reads and writes a sequence – but the input is not the same as the output (in ≠ out). Reading a medium and producing one are different abilities, and they are learned in different ways.
This means:
- A model can read in an image but only output text. That’s exactly what Claude currently is: show it a screenshot of a broken checkout page and it will describe in words what’s wrong – but it won’t give you an image back.
- Some models can generate images. Different abilities, different mechanisms.
- Audio has the same distinction: some models can listen (audio input), others can also speak (audio output).
- A model that takes in and puts out audio – and is fast enough to keep up with a live conversation – is the special one. And that is exactly what a phone call requires.
| Combination | What it means |
|---|---|
| Understand images, generate text | The model describes a screenshot in words but can’t deliver an image itself. |
| Understand text, generate images | Image generation – a separate ability with its own technology. |
| Understand speech, generate text | The model listens but only responds in writing. |
| Understand speech and speak, in real time | The most demanding combination, and the prerequisite for a phone conversation. |
Conclusion
An AI model is a giant knowledge engine that takes in an ordered sequence of tokens – text, sound or image – and continues it with an ordered sequence of tokens; where what it can take in and what it can put out are two different lists.
Once you understand that input medium, output medium and a model’s speed are separate dials, you start to wonder how exactly the right information ultimately reaches users – or, in customer service, customers. The knowledge reaches them through channels – by phone, email, chat – and making sure the answer actually reaches customers in a form they can use is an entirely different problem, which we’ll tackle in our follow-up article.



