Tokens for Tokens
This is all and everything a token is to me. You put it into the machine and you get to play a video game.
This is me describing Aladdin’s Castle, my digital playground as an 80s kid. This arcade used tokens instead of quarters, and those tokens represented a quarter for the purposes of playing a single game.
A token is simply a symbol for something else. It can look like the thing it’s supposed to represent, or it can appear totally different—the important thing is that it stands in for another thing.
This was an ultra-simple stand-in for a quarter, but you can have stand-ins for just about anything. When we’re talking about generative AI, tokens represent bits of words, or even very common words themselves.
By assigning a preset value for specific combinations of letters, AI is able to do its thing. This tokenization process is crucial to getting a good outcome and having a powerful and effective large language model. Instead of looking at all possible combinations of letters and deciding which individual letter makes sense next, the model just picks the next word or word fragment.
Thank you, tokenization.
Have you noticed that there are two types of tokenization here? One is a stand-in for ownership or value, and the other is a stand-in for a combination of letters.
Both use a simpler or cheaper trinket (token) to represent the real thing. And, their worlds are about to intersect in a big way.
That Aladdin’s Castle-style token is a representation of value, while the computer/LLM token represents language. At their core, both are just representations of information, packaged in a more useful way. Are all language tokens the same, though? In other words, are they fungible?
This turns out to have a complicated answer. A two-letter bit of text that shows up all the time is a lot more useful than one that doesn’t. Still, the thing to do if you’re an AI company is to keep track of all the tokens a particular model uses, and charge you on that basis.
The tokens aren’t really apples to apples, but if you add them all up, it makes perfect sense in aggregate. If we each use 100 million tokens, they are very likely to use roughly the same amount of compute—assuming we are using the same model, mind you.
The individual word fragments start out like morsels of food in a stew, but once they’ve cooked for a bit, you just sell the soup by the ladleful. Similarly, you have a certain accounting fungibility now that your LLM model is a stew.
So now, tokens are being created to buy…. tokens.
Language tokens are bought with stablecoin tokens, which are themselves tokens for dollars… which are tokens for value.


One must not also forget the vocal group The Tokens, who brought "The Lion Sleeps Tonight" to the attention of North America and operated a record label named B.T. Puppy. I'll spare you "Lion" by providing a link to their first big hit instead: https://www.youtube.com/watch?v=WRe753rvhpc
That last line makes me want to toke on something else.
80s arcades were awesome but the one closest to us was in Westwood which was too far away during the week so we went to 7-Eleven. The one on the walk home to my best friend’s house had Star Castle and we were addictted.
It was a vector graphics game and you had a little spaceship trying to destroy a cannon inside 3 rotating circles and different speeds. You had to shoot through the circles or shields and then get the cannon once there was a path.
But once the cannon had a path it would shoot you. Oh and then the shields regenerated. In the 7-Eleven versus the arcade we had Star Castle all to ourselves.