First, I’d like to wish you all a poetic, fun, yet healthy start to the New Year. At the same time, I’d like to thank you for your comments and discussions on this blog. For four and a half years now, we’ve been talking about artificial intelligence here, especially deep learning—that’s wonderful! Thank you!
Most recently, we’ve had OpenAI’s chatbot ChatGPT write Christmas stories. Thank you so much for participating! The ChatGPT Christmas stories based on your creative ideas in the blog comments were wonderful:
Now, just before New Year’s Eve, humor is in order. Even as we reach the peak of a year marked by global and domestic turmoil: COVID-19, war, climate change, the energy crisis, rampant conspiracy theories, hate, lies, and stupidity in the media and on the streets, pseudoscience. But we’ve cried and whined far too often, and that hasn’t made the world any better. I have lived under two dictatorships. Dictators’ greatest fear is not being taken seriously: that people will laugh at them instead of fearing them. Since we’re talking about artificial intelligence here on this blog, let’s have artificial intelligence contribute to a humorous end to this year:
Can you ChatGPT get it to write the most absurd text possible? The more absurd and funny, the better.
Otherwise, I don’t want to delve too deeply into theory in this post. I’ll just briefly explain at the end how “deep learning” language models relate to facts: so we can view artificial intelligence with a bit more ease: In my opinion, “deep learning” models are the greatest technology that could have been gifted to humanity so close to the brink. But we should be aware of their shortcomings and stop anthropomorphizing this tool. So that we can finally put it to proper use to save the world. Otherwise, we’ll have to rely on Hollywood for that.
I have covered the theoretical background of generative models frequently and in detail this year in the video series AI Crime Stories on the video channel of the same name at our SRH Distance Learning University : If you want to learn about generative language models like GPT-3 and chatbots like LaMDA (or ChatGPT) in a way that’s easy to understand, I can recommend the videos from the “AI Crime Stories” series “Do Chatbots Have Feelings?” and “Can Artificial Intelligence Have Consciousness?” . I discuss generative text-to-image models like DALL·E 2 in the “K.I. Krimis” episode “Will Artificial Intelligence Take Over Art & Design?”. The latest (and rather short :-)) episode also deals with DALL·E 2 and explores the question: “Has DALL·E 2 Developed a Secret Language?” By the way, we’re thrilled to have every subscriber to the “K.I. Krimis” channel. Thank you!
The next episode of AI Crime Stories has a lot to do with the topic of this blog post: humor and deep learning. In the January video, I’ll ask (and answer) the question: “Does DALL·E 2 have a sense of humor?” Today, I’d like to use your sense of humor to explore the unintentional “humor” of ChatGPT . Just like with the Christmas story in the penultimate post on this blog, I’ll start with an example:
I asked ChatGPT to come up with a new brainteaser and its solution. The chatbot suggested a wordplay that is difficult to translate into German. So in the following screenshot, I’ve included only the original English version of the puzzle I created and ChatGPT’s solution in the following screenshot (“JA” stands for Jaromir; the OpenAI symbol indicates the responses from ChatGPT ):
At first, I was shocked when I read ChatGPT’s suggestion for a puzzle task:
“The task is to rearrange the letters in the word ‘EVALUATE’ to form a new word that is a synonym for ‘evaluate.’”
That sounds like a damn good language puzzle, doesn’t it? Can you actually rearrange the letters in the word “EVALUATE” to form a new word that is a synonym for the same word “evaluate”? But when I read the explanation of the puzzle from ChatGPT , I had to laugh: The chatbot’s answer reminded me of Karl Valentin.
“To solve this puzzle, you must first understand the meaning of the word ‘evaluate.’ This means to consider the worth or value of something, or to assess its worth. Then, you can rearrange the letters in the word ‘EVALUATE’ to form a new word that has a similar meaning. In this case, the word ‘VALUATEE’ is a synonym for ‘evaluate,’ meaning that it has the same meaning as the word ‘evaluate.’ Therefore, ‘VALUATEE’ is the correct solution to the puzzle.”
Nonsense!, I thought. But ChatGPT presented the explanation for the solution to its language puzzle with such conviction that I slowly began to doubt myself: Maybe my English isn’t that good. Maybe the claim that “VALUATEE” is a synonym for “EVALUATE” is true. I asked a friend who had lived in England for a long time: Fortunately, he confirmed what I thought: The chatbot was talking nonsense.
ChatGPT itself provided a nice punchline to this. I added the chatbot as an extension to my Chrome browser. ChatGPT seems to process language better than the AI behind Google’s search engine. Plus, ChatGPT gives you a detailed, direct answer to your question, rather than a list of websites from Wikipedia where the term appears frequently or makes semantic sense. On the right side of the screenshot below, for example, you can see how ChatGPT explains “New Year’s Eve” when I type “What is New Year’s Eve?” (modern language models can do much more with our searches when we ask proper questions, rather than just searching for terms):

So how does ChatGPT the word “VALUATEE” in the Chrome extension—a word the chatbot itself suggests as the solution to its language puzzle and as a synonym for “evaluate”?

As you can see in the screenshot above on the right, ChatGPT my question about whether “VALUATEE” is a synonym for “evaluate” with: “No, ‘valuate’ is not a synonym for ‘evaluate’. … ‘valuatee’ is not a word recognized in the English language.”
Thus, the chatbot suggests “valuatee” as a synonym for “evaluate,” but at the same time claims that “valuatee” is not an English word. Better safe than sorry, I thought, and asked ChatGPT directly whether “valuatee” is an English word:

That’s just how it is with “deep learning” language models: You can’t rely on the content of their answers. They often provide “alternative facts,” albeit eloquently and convincingly.
“Deep learning” language models are essentially optimized not to spread disturbing or biased content—so that we’ll like them. Platforms are forced to do this by the media and public opinion. Otherwise, storms would rage through the media landscape. But when do we (statistically speaking) like a chatbot’s answers the most? When they tell us what we want to hear. This was quite evident in the dialogues between Google developer Blake Lemoine and Google’s chatbot LaMDA :
Based solely on his conversations with LaMDA , Blake Lemoine attributed human emotions and even a soul to the chatbot. Without considering how a language model is structured and how it is trained. Although no architecture of a modern “deep learning” model enables human mental states: Generative models can only identify features of the data examples on which they are trained and reassemble these features—that’s all they can do!
No matter how large the models are made: Even if a language model has more parameters in a year than the human brain has synapses, it is not a human-like being or an AI (Artificial General Intelligence) on par with humans, but rather a statistical optimization process—a tool! A wonderful one, to be sure, but still a tool.
Here we see a paradox that I would like to call the paradox of the inevitable humanization of language models: We train language models—automatically supervised (self-supervised)—on vast amounts of text written by us humans. So that the models learn to mimic our language. Once they can do this perfectly, many people are so unsettled by it that they deny the machines are machines and attribute human characteristics to them.
But no matter how eloquently the language models speak, they merely reflect our language. They cannot think. They have no sensorimotor “experiences” of the world, no common sense, and no intentions. They cannot think; they can only mix features across various hierarchies and thus also various concepts—to the extent that their filters and the optimization of their objective functions allow. That is why they reflect not only soundness and factuality, but also the alternative facts of the foolish and the shameless. Of course, soundness and factuality can be trained into the models the better they are developed. And they are getting better and better: GPT-3 demonstrates more soundness than GPT-2, and ChatGPT more than GPT-3.
However, we can never achieve a situation where “deep learning” models provide only facts. To a certain extent, they will always spread nonsense. Why? First of all, the models learn from human texts, and the knowledge contained in these texts is limited. Furthermore, as already mentioned, they are optimized not to disturb us. It is also right and important that “deep learning” models do not spread racist, homophobic, and misogynistic prejudices or display brutality. But: How can “deep learning” models be made both well-founded and factually accurate on the one hand, and on the other hand, ensure they do not generate disturbing content? In my opinion, this isn’t possible: disturbing content simply exists in the world. And the knowledge we use to train the models is not perfect.
As I wrote in a response to a comment on my penultimate blog post : We could approach this much more calmly if we viewed “deep learning” models for what they really are: tools! Not “intelligent” or “conscious” machines. Just as knives are tools. We don’t expect knives to be trustworthy either. We only expect people to handle knives “responsibly.”
Despite what I wrote above, I believe that we cannot save the world without the help of the tool (the technology) known as artificial intelligence. While we cannot always make the models “trustworthy,” we still need them. We just need to know what can go wrong and when. Especially when the models are used in sensitive areas such as medicine. That’s why we must monitor the decisions made by trained models and test them extensively: just as doctors are examined before being allowed to make life-and-death decisions.
Alright! I’ve been serious enough. Let’s get started with our game. Are you in? Can you create the most absurd text possible using ChatGPT and post it in the comments below? That would be great and fun! Just as the year comes to a close. Thanks!
Featured image: ChatGPT & DALL·E 2 (both models from OpenAI) based on my prompts.