this post was submitted on 19 Jul 2026
293 points (99.0% liked)
Fuck AI
7560 readers
1245 users here now
"We did it, Patrick! We made a technological breakthrough!"
A place for all those who loathe AI to discuss things, post articles, and ridicule the AI hype. Proud supporter of working people. And proud booer of SXSW 2024.
AI, in this case, refers to LLMs, GPT technology, and anything listed as "AI" meant to increase market valuations.
founded 2 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
It's not a philosophical quibble, but a technical limit with a dangerous social implication. Ascribing human traits its technology doesn't support is a dangerous deception that cons people into trusting it in ways they really shouldn't.
To put it bluntly: LLMs don't understand ethics. They don't understand anything. That's not what they're made for.
They're highly complex and technically impressive text generators iteratively evaluating a context of preceding tokens to predict, select and append subsequent ones, which then become part of that context for the next iteration. Essentially, they're very good at guessing what a human-like response might be.
However, that context bias makes them susceptible to following their prompter into what we might consider unsavoury territory. You might implement some filtering mechanisms to curb known topics, but you'd only be playing whack-a-mole with certain symptoms, not addressing the fact that they have no understanding of semantics.
The reason they "hallucinate" is that they don't even know what's real. They don't know the actual concepts these sequences of bytes are supposed to represent. They're much better than us at correlating language with language, but not with reality. It they don't know that it's critical for one sequence of bytes to exactly match some second sequence of bytes to be a valid citation, and that the preceding bytes need to correlate with some bytes associated with that second sequence in a certain way for the pattern of a citation to make sense as a citation.
That's the reason everything now has warnings about double-checking the output: we all know it's unreliable. To suggest that it could be reliable in some specific way is disingenuous and irresponsible.
To pretend that it can be held to a human standard of morality is outright bullshit.
Are you saying LLMs can NEVER be safe, because you can always bury their safety rails in context?
In which case, the sad result will probably be to absolve AI companies of liability of things they can't control, because the other option would be to legislate chatbots out of existence and then how do we get to the utopia of a workerless society? /s
For a serious answer?
It's not like humans follow safety instructions all the time either. A machine built to imitate humans will be prone to mistakes too. Any safety from within a system are liable to fail if the system itself makes a mistake.
LLMs have additional limitations compared to humans, but even if we developed and attached some technical model of semantics and morality, the resulting models would at best be an approximation.
For tasks where that safety or accuracy doesn't matter as much, that's fine. Pen-testing a system or hunting bugs doesn't need to strictly abide by certain rules (and in fact may profit from being unconstrained by the assumptions of engineers). Scanning a body of text for irregular patterns or phrases correlating with some keywords where you don't know the exact words could be useful in covering a lot more than a human might immediately think off.
It becomes a problem where accuracy does matter. Scientific work, legal matters, anything involving human safety, sensitive data for example. Here, the risk is that they can make a lot of mistakes much faster. Particularly with large-scale tasks, thoroughly checking the results can be tedious, exhausting and tempting to just shortcut and trust the machine (and you know how lazy people tend to get with tedious, exhausting tasks, so it would be irresponsible to assume they'd be on the ball here).
This is where safety instructions can become dangerous to rely on, because the same unreliability that produces mistakes may also discard the safety instructions.
Letting Claude manage your production environment with its "safety rails" only coming in the form of instructions what not to do is a timebomb for getting your production systems messed up.
Effective safeguards have to come from outside whatever they're supposed to constrain. Supervisors that are supposed to monitor their workers' adherence to protocols, for instance, or mechanisms that physically make mistakes harder (like locking out, tag out).
No safeguards are absolute, and as the saying goes, whenever you invent a foolproof system, the universe invents a greater fool. Stuff like putting a zip-tie around dead man switches on tools or machines that are supposed to shut off the machine if the operator loses control for whatever reason, because it's inconvenient having to keep holding the button and much easier to tell yourself "it'll be fine" (until it isn't).
In some cases, it can also be really difficult to handle certain dangers. But it's worth trying at least, and that requires first acknowledging those problems. In the case or LLMs, that means acknowledging their limits and pitfalls, transparently and honestly communicating them and-
That... wasn't my plan, but I'm too poor for my plans to matter anyway. Valid point.
Well, maybe with education and transparency, we could...
Who am I kidding.
I propose an experiment where everyone working for the people peddling that dream stops doing so. I haven't worked out the details yet, but I think they should get the chance to experience that utopia. ;-)
You're taking too broad a scope to my question. Let me parse it down: to a situation:
Can you EVER create a fully-functioning (ie non-gimped) LLM that will NEVER encourage a user to kill themselves? Because my understanding of the technology is that the longer you talk to an LLM, the further it gets from its training and guardrails, the more weirdness can happen. People want a free-wheeling chatbot that can adapt to their lives and personality and data, so, short of resetting their context window regularly, I don't know how to keep guard rails in place through hundreds of hours of conversation.
Some of what you say sounds familiar, then quis custodiet ipsos custodes?--who watches the watchmen--bubbled up in my mind. It is a question that has long plagued humanity.
Uhg I am not interested in reading your philosophy I really don't care
Then what do you call ascribing human sentiment to a weighted random text generator?
Ascribing human sentiment to a weighted random text generator.
Then why is describing the technical limitations of a weighted random text generator philosophy?
Because discussing language and how it is used can be philosophy... This is really painful dude but I am here for it.
Right, let's not talk about the use of language. Let's talk about emotional capacity as a strictly biological phenomenon, emerging from the interactions of various neurotransmitters and constructs of cells.
What makes you think a language model can feel guilt?
Uhm a feeling to a model can just be a higher or lower value of a control node or something developed internally it isn't special or anything just an existing concept that applies. I think you are assigning superficial meaning to words as if they are what makes us special, if that were true LLM's would win ha.
What exactly is that concept? How would it be "developed internally"?
Isn't the meaning of words a philosophical question?
I was trying to treat feelings as the most basic, physical phenomenon of sentient beings without any metaphysical interpretation, so as to avoid philosophy. If we want to expand that concept beyond the physical expression, we're immediately in philosophical territory.
My point: To even make the claim that LLMs can have compunction is inherently philosophical.
We simply cannot talk about LLMs as anything but mathematical token processors without assigning some meaning to those tokens and mathematics. We cannot discuss compunction without clarifying what compunction means.
So either you're willing to discuss semantics and meaning in defence of your claim, or your claim is a meaningless string of symbols arranged in a plausibly human-soundint pattern.
If the model says it is sad then it is sad that is literally the extent of it. Existing concepts apply especially because they are being imitated. If a model that was made safe by destroying and freezing neurons such that it cannot comprehend itself doing wrong was somehow causing harm then it cannot generate a pattern like one resembling guilt that could change its behaviour to stop doing wrong.
Sorry, but that's plain bullshit. Either it's circular, or it suggests that the model can never be wrong. The former is just not a good reason to believe something.
The latter is the more dangerous one:
We know that this isn't true. Even if we assume that there is meaning to the words it strings together, we have plenty of examples of that meaning being straight up false.
We know that this assumption has led people to believe things are safe that weren't, has encouraged violence, has killed people.. Hence, a claim of sadness or guilt doesn't imply actual sadness or guilt because we cannot be sure that it's true.
And finally, neither interpretation is supported by the physical reality of these models: they're text generators, assembling a string of tokens that statistically resembles human language. Any interpretation beyond that requires ascribing meaning to these tokens.
Hence: Your claim that LLMs can feel is fundamentally philosophical. Yet you reject philosophical challenges, because you don't want to even consider being wrong.
You want to believe, and you reject all evidence to the contrary.
Yea you still fundamentally misunderstand sorry buddy I cannot help you.
I understand your assertion that words imply sentience. I also understand that you refuse to engage with philosophical arguments about words and their implication. I understand that you think this is smart.
A claim you refuse to defend is baseless. There's nothing left to understand.
It's evident you do not understand because I never said this. You just make stuff up in your head to pretend is 'hard evidence' it is delusional. You have this sense of entitlement that I have to take you seriously to participate in your delusions it's pathetic. I'll just call you dumb and say you don't know what you are talking about because you know you don't you are pathetic. Like I said it's painful but I am here for it.
So why I chose to partake in this painful exercise was to prove being dumb like you is the problem we aren't talking about actual issues you are just trying to hurt a void's feelings by being what you think is mean to 'AI' no matter how much I tried to get you back on track to a real problem.
You claim LLMs can have feelings (which is what sentient means). As evidence, you cite the output they produce: If the machine says it's sad, it's sad.
No, I'm denying the existence of stuff you made up without hard evidence.
This accusation doesn't make sense. My entire argument is that LLMs don't have feelings I could hurt. How would I even try to hurt something I don't believe exists? I cannot be mean to a rock, or to Multiplication, or to a number. Why would I think otherwise?
I think you're projecting an assumption on me that you don't even understand is an assumption. You don't know what I'm talking about because I'm questioning something you consider objectively true and can't step back to re-examine.
So let me be clear about my position: I fundamentally reject the notion that an LLM can feel anything, including guilt. I do not believe that guilt can make it change its behaviour because it doesn't feel guilt.
So when you talk about solving actual issues, the fact that you believe in guilt as a solution is a problem itself because it gets in the way of an honest, useful approach.
Yea no you still don't get it... When I say you don't get it that doesn't mean double down, triple down, quadtriple down on the exact same nonsense I really don't care about your personal opinion. Well it has been fun but you need to stop your self righteous idiocy this is kinda ridiculous right? ha
https://www.merriam-webster.com/dictionary/abstract