How Models Talk: Tensors and Tokenizers
A short discussion and interactive lab going through how models see and process input
One of my personal a-ha moments in understanding LLMs is understanding that at a base level, a model is a file on a filesystem and as part of its invocation process input is passed to that model. These files are “weights” and are loaded by inference servers or other similar systems.
When people in application security think about inputs into a system, they think about strings and payloads - things you can fuzz. A second a-ha moment was understanding what a tokenizer is and how it works.
I’ve added two models (a text model and an image model). You can try them below. They run entirely in your browser - the models are fetched from Hugging Face when you press the button, and nothing you type or upload is sent anywhere.
How models parse text input
When you send in a sentence into a system, like:
Hello there, my name is PratikThat string isn’t what the model sees, instead it sees a tensor, like this:
tensor([[ 0, 31414, 89, 6, 127, 766, 16, 2869, 415, 967, 2]])At first glance, you might look at this and want to walk away from your computer but its not that complicated once you understand how it works! First, a tensor is just a multi-dimensional array, for now just think of it as a list of numbers. A model does not consume text directly, instead it consumes a tensor and returns a tensor back.
In order to convert text into a tensor you would use a tokenizer, this is done transparently in pretty much any AI interaction you have ever done but it is a critical step. Different models use different types of tokenizers, for the example I have loaded onto this page using Transformers.js it is the one that ships with RoBERTa - the same tokenizer used by the spam classifier from my Hackfest talk.
As an example of a simple to understand tokenizer, you might have a “list of words” tokenizer, where the tokenizer has a large JSON file that maps wordparts to numbers (here is RoBERTa’s. The tokenizer looks up each piece of your text in that table and constructs a tensor out of the numbers it gets back. For our sample string above, here is how it all breaks down:
| Piece | Token | ID |
|---|---|---|
<s> |
0 | |
| Hello | Hello |
31414 |
| there | Ġthere |
89 |
| , | , |
6 |
| my | Ġmy |
127 |
| name | Ġname |
766 |
| is | Ġis |
16 |
| Pr | ĠPr |
2869 |
| at | at |
415 |
| ik | ik |
967 |
</s> |
2 |
The Ġ marks “there was a space before this”. Spaces are part of the token, which is why Ġthere and there are two different entries with two different numbers.
My own name is an interesting row. “Pratik” is not in the vocabulary, so it gets broken into three pieces: ĠPr, at, ik. Nothing in the model ever sees “Pratik” as a unit - it sees 2869, 415, 967.
How models parse images
That might be a bit of a simple example, how does it work if you upload an image? You can’t match a picture of a cat to a list in a JSON file. Firstly, the type of model that would be used is different, this might be an image classification model - for example MobileViT, a small one from Apple that sorts pictures into a fixed list of 1,000 categories.
There is no vocabulary to look things up in, but the job is the same: get to a tensor. An image is already numbers - three colour values per pixel - so the work is mostly making those numbers a consistent shape and size. Each model wants the image in a different format, but all the images it sees need to be in that format for it to be accurate. The requirements for this model can be found in the config file that ships with it:
- Resize the image so its shortest edge is 288 pixels
- Crop the middle 256 × 256 out of it
- Divide every colour value by 255, so they all land between 0 and 1
- Reorder the colour channels from RGB to BGR
What comes out is a tensor with the shape [1, 3, 256, 256] - one image, three colour channels, 256 rows, 256 columns. That is 196,608 numbers. This is the “multi-dimensional” part of the tensor.
The model gets this tensor, performs its task and returns a tensor. That response is a list of all its 1,000 categories with a value associated with each one. Here is what actually comes back when an image of a tiger is sent:
tensor([[0.6880, 1.1420, 0.9890, 0.8280, 0.3550, -0.1250, 0.7890, 1.1680, ...]])
# shape: [1, 1000]One number per category, in the order the model’s label list defines. Notice some of them are negative - so these are not probabilities yet. They are raw scores, called logits. To turn them into percentages that add up to 100 you run them through a softmax, which is what the widget above is doing before it shows you a result.
Sorted, the top five and bottom five look like this:
| Category | Logit | After softmax | |
|---|---|---|---|
| 292 | tiger, Panthera tigris | 10.438 | 71.796% |
| 282 | tiger cat | 9.201 | 20.845% |
| 290 | jaguar, panther, Panthera onca, Felis onca | 3.952 | 0.109% |
| 281 | tabby, tabby cat | 3.090 | 0.046% |
| 288 | leopard, Panthera pardus | 3.051 | 0.044% |
| … | |||
| 740 | power drill | -0.608 | 0.001% |
| 998 | ear, spike, capitulum | -0.612 | 0.001% |
| 956 | custard apple | -0.706 | 0.001% |
| 845 | syringe | -0.822 | 0.001% |
| 941 | acorn squash | -1.155 | 0.001% |
The model did not decide “this is a tiger”. It scored all thousand of its categories and tiger came out highest. Acorn squash also got a score. So did syringe. Every image you ever send it gets a number against all 1,000, whether or not any of them are remotely right.
Other formats and security
At the core, models need input to be normalized in some way and turned into a tensor. Audio processing, for example, goes through a similar but different process. When you are thinking about how to target a model that is performing a task, an average pentester might think their string with some {} or ; characters will somehow trigger an issue within the model’s internal logic, causing it to go rogue - but as you now know, those values are actually just a bunch of numbers that have math operations performed on them.
You can, however, manipulate the steps leading up to it. In the image example above, a product that accepts an image upload for classification needs to do those steps itself, and in that flow there may be input validation issues or library vulnerabilities in something like ImageMagick. Context wise, whatever comes out of that model, especially if it is chained into further tool usage, can also be an interesting avenue for attack. Manipulated tensors are an attack path in their own right - that is what adversarial ML attacks do.
At least for me, understanding how these flows work helps remove some of the magic from it all, and that helps you think clearer about what problems to tackle.