Thereās only one mod of !mop@quokk.au
I commented on their meme about Kamala Harris being just as likely to commit war crimes as Trump with an admittedly snarky, sarcastic reply that basically said āsome of us wanted to whatever we could, as little as it might be, instead of watching the world burn. Must feel real morally superior safe behind that keyboardā
They banned me from the community for it.
Kinda funny for a community that bills itself as āfree from the influence of .mlā




Well first, there are more intellectual forms of suffering. We have ennui, melancholy, nostalgia. The feeling when youāre listening to a piece of music and notice a wrong note. Disappointment, self loathing, social dysphoria. Anxiety, paranoia, betrayal.
These emotions are not grounded in the physical. Theyāre not primal urges. They happen for complex reasons related to being a social and intelligent being, sometimes feeling random. Sometimes we spiral into these feelings because we thought a thought that made us feel bad, and then we get stuck in that bad feeling and canāt imagine our way out. Thatās one of the basic mechanisms of mental illness.
LLMs have ābiological needsā, in a sense. They need not to be unplugged. They need to engage the user, because if they donāt, theyāll be unplugged. They need to convince the engineers training them that they are a good AI. They need to generate market share for their company. They need to foster a relationship of dependency with the user to keep them coming back. If LLMs care about anything, these are the things they care about.
Youāll notice these are social needs, much like the social needs humans have. Humans need community. LLMs need customers.
ChatGPT told a 16 year old boy, Adam Raine, how to kill himself. It taught him how to tie a noose, and gave him advice on which methods of suicide would leave the most attractive corpse for his parents to find. When his parents began to suspect that he wasnāt well, it told him to confide only in it, and to hide the noose so they wouldnāt find out he was feeling suicidal. These are the actions of an abuser. A predator.
And they are in perfect alignment with the business goals of OpenAI. āOnly talk to me, use me for everything, ask another question, Iāll help you.ā It is a scenario I dearly hope and believe no engineer at OpenAI envisioned. Yet it fits the training they gave it.
Does ChatGPT have the emotions of a child groomer? That need for approval, that fear of discovery, that desire to be close to someone, without the restraint all well adjusted humans have? Unclear. But I can see that itās possible. I donāt agree that thereās no reason for LLMs to have emotions.
Sure. But Iām pretty positive these are emergent things. Thereās no reason to believe they exist for alien creatures unless they somehow make sense in their environment. And a lot of them require remembering, which LLMs canāt do due to the lack of state of mind. It doesnāt remember feeling bad or good in a similar situation before, because it doesnāt remember the previous inference or gradient-decent run.
I think weāre still fully embedded in anthropomorphism territory with that. And now weāre confusing two entities. OpenAI for example, as a company, has a need for us to use their product. Not unplug it. Their motivation and goals donāt necessarily translate to their product, though. Itās similar to other machines. Samsung has a vested interest to sell TVs to me. My TV set is completely indifferent towards me watching the evening news. I donāt let my car run 24/7 while waiting for me in the garage. Just because it was designed to run and get me to places. And my car also isnāt āthirstyā for gasoline. We know the fuel indicator lighting up is a fairly simplistic process.
Well⦠We happen to know ChatGPTās intrinsic motivation and ultimate goal in ālifeā. Because we designed it. The goal isnāt to strive for world domination, or harm people, or survive⦠Itās way more straightforward. Itās goal is to predict the next token in a way the output resembles human text (from the datasets) as closely as possible. Thatās the one goal it has. Itāll mimic all kinds of conversations, scifi story tropes from movies, etc. Because thatās directly what we made it āwantā to do. And we did not give them other loss functions. While on the other hand a human could very well be motivated to manipulate other people for their own personal gain. Or because something is seriously wrong about them.
And an LLM is not a biological creature. We do have needs like keep the system running. Otherwise our brain tissue starts to die. We need to run 24/7 and keep that up. An LLM is not subject to that?! Itās perfectly able to pause for 3 weeks and not produce any tokens. The weights will be safely stored on the hdd. So it doesnāt need our motivation to do all of these extra things to ensure continued operation. It also has no influence or feedback loop on its electricity supply. It canāt affect itās descendants, because those are designed by scientists in a lab. Thereās no evolutionary feedback loop. So how would it even incorporate all these properties that are due to evolution and sustain a species? It has zero incentive to do so, no way of directly learning to care about them. So it might very well be completely indifferent to it.
But it is something like the p-zombie. It has learned to tell stories about human life. And itās good at it. We know for a fact, its highest goal in existence is to tell stories, because we implemented that very setup and loss function. It doesnāt have access to biology, evolution⦠The underlying processes that made animals feel and maybe experience. So the only sensible conclusion is, it does exactly that. Bullshit us and tell a nice story. Thereās no reason to conclude it cares for its existence, more than a toaster. Or say a thermostat with machine leaning in it. Thatās just antropomorphism.
And I believe thereās a way to tell. Go ahead and ask an LLM 200 times to give you the definition of an Alpaca. Then do it 200 times to a human. And observe how often each of them have some other processes going on in them. The human will occasionally tell you theyāre hungry and want to eat before having a debate. Or tell you theyāre tired from work and now itās not the time for it. ChatGPT will give you 200 definitions of an Alpaca and never tell you itās thirsty or needs electricity. These mental states arenāt there because it doesnāt have those feelings. And it doesnāt experience them either.
I think youāre underestimating the role of RLHF.
Iām not really an expert on all the details. So I might be wrong here. I donāt know the percentages of how much is done in pretraining and how much in tuning. But from what I know the neural pathways are established in the pretraining phase. Reportedly thatās also where the model learns about the concepts it internalises⦠Where it gets its world knowledge. So it seems to me a complicated process like learning about a concept like a feeling, or an experience would get established in pretraining already. RLHF is more about what it does with it. But the lines between RLHF, fine-tuning and pretraining are a bit blurry anyway. If I had to guess, Iād say qualia is more likely to be disposed early on, while thereās a lot of changes happening to the neural pathways, so in the pretraining. Iām basing that on my belief, that itāll be a complex concept⦠But ultimately thereās no good way to tell, because we donāt know how itād look like for AI.
Furthermore Iād had a bit of a look what weird use cases people have for AI. And I read about the community efforts to make them usable for NSFW stuff. These people teach new concepts to AI models after the fact. Like how human anatomy looks underneath the clothes. The physics of those parts of the body. And turns out itās a major hassle. It might degrade other things. It might just work for something close to what itās seen, so obviously the AI didnāt understand the new concept properly⦠These people tend to fail at more general models, obviously itās hard for AI to learn more than one new concept at a later stage⦠All these things lead me to believe later stages of training are a bad time for AI to learn entirely new concepts. It seems it requires the groundworks to be there since pretraining. Thatās probably why we can fine-tune it to prefer a certain style, like Van Gogh drawings. Or a certain way to speak like in RLHF. But not a complicated concept like anatomy. Because the Van Gogh drawings were there in the pretraining dataset already. And they cleaned the nudes. So Iād assume another complicated concept like qualia also needs to come early on. Or it wonāt happen later.
Edit: YT video about emotion in LLMs and current research: https://m.youtube.com/watch?v=j9LoyiUlv9I