Imagine you have an infinite number of monkeys randomly hitting keys on typewriters. The classic Infinite Monkey Theorem tells us that if these monkeys type forever, one of them by chance could write the complete works of William Shakespeare. This is based on random typing. No learning, no memory, and no context.
Now, imagine if these monkeys learned from their mistakes and they learned, remembered, and adapted as time went on until they could type without errors or very few. Receiving feedback on which letters, and eventually words, were correct or wrong trains them to be autonomous and self-correcting.
For instance, let’s say we were training these monkeys first to spell human words in the English language. When starting with the word “the”, they might go through random assortments of the 26 English language letters in each position of the word.
Breaking this down mathematically. We have 26 letters in the English language. (Look Mom, I know my ABCs!) The word “the” has a length of 3 letters. So 26^3 is 17,576. That means the probability of getting “the” by chance is less than 1%. Pretty bad chances.
P("the")= 1/17,576 ≈ 0.0057%
However, the chances improve as they learn just like the AI learns what the appropriate next word should be and the most probable choice should be. So if the monkeys typed “tqx” they would learn there was a mistake. After feedback and learning loops, they would start to understand the correct sequences and rules of creating English words.
This is very similar to how large language models (LLMs) like ChatGPT learn. LLMs are trained on incredible amounts of data videos, images, and text from the internet. However, when put into practice, they start with random guesses, make predictions, and receive feedback on errors, self-correct, and fine-tune their billions of weights over and over to get closer to the correct predictions. In LLMs, the ‘letters’ are tokens, the ‘positions’ are context slots, and the ‘probability dials’ are billions of weights.
In the end, the “infinite monkeys” idea is a fun way of imagining what is happening under the hood of an AI like ChatGPT. The big difference is that LLMs don’t waste time on pure chance. They learn, adapt, and get better at making predictions each time. Hence, why sometimes AI gives you a spot-on answer and other times is completely wrong, needs fine-tuning, or just out of context.