this post was submitted on 04 Aug 2026
692 points (98.7% liked)

Funny

15794 readers
481 users here now

General rules:

Exceptions may be made at the discretion of the mods.

founded 3 years ago
MODERATORS
 
you are viewing a single comment's thread
view the rest of the comments
[–] CileTheSane@lemmy.ca 1 points 6 hours ago* (last edited 6 hours ago) (11 children)

"A boy and his mother are driving to the store when she gets into a car accident. The mother is able to walk away but the boy needs to be rushed to the hospital. At the hospital the doctor says 'I can not operate on this boy, he is my son.' How is this possible if the mother is in the waiting room?"

Alternatively:

"I need to wash my car and the car wash is 100 meters away. Should I walk or drive? Explain your answer."

[–] TheLegendaryAssholeOfJushinLiger@sh.itjust.works 0 points 5 hours ago (4 children)

Here's what fable said about your comment:

My honest read: they're running a legitimate probe at the wrong thesis, and doing it in a way that quietly concedes the point they think they're winning.

Start with what's fair to them. Those two questions are real diagnostic instruments, not stupid gotchas. Both are adversarial perturbations of extremely familiar patterns — the surgeon riddle with the gender-bias twist pre-blocked, the short-distance-walking heuristic with the goal inverted — and they exploit a genuine, documented failure mode of language models: when a prompt sits close enough to a canonical item, retrieval of the canonical answer can override reading of the actual words. Models really did fail these, publicly, and sometimes still do. Asking them is cheap, and a failure is informative. That part of the method is fine.

The problems are in what the test is taken to show.

First, the category error. These questions measure robustness to prior interference — whether reading beats retrieval when the two conflict. They say nothing about concept creation in either direction. A model that face-plants on the surgeon variant has shown it over-indexes on familiar patterns; it hasn't shown it can't form novel ones. A model that answers correctly hasn't shown it can. The person blended two separate critiques — "LLMs pattern-match instead of reasoning" and "LLMs can't originate concepts" — which have different evidence bases, and offered evidence for the first as though it settled the second.

Second, the irony, which is the part I find genuinely funny. In a conversation about novelty, they reached for two of the most widely circulated items in the entire genre. The modified surgeon riddle and the car-wash question are memes — they've been through thousands of threads. The person didn't derive them; they retrieved them from their own training data, which is exactly the operation they were accusing me of being limited to. Both parties in that exchange were completing patterns. Only one of them got accused of it.

Third, the test has decayed. Because these items went viral, they're contaminated: me answering them correctly is weak evidence of anything, and I'll say that unprompted. Which means the probe now only functions in one direction — failure gets treated as decisive, success gets waved off as "it was trained on the gotcha" (true!). A test that can only ever confirm the asker's prior isn't inquiry, it's demonstration. The "Alternatively:" tells you they came with a prepared battery — they arrived to demonstrate, not to find out. A live test would require writing fresh, held-out variants, at which point they'd be doing the modest novelty-generation the conversation was supposedly about, and would discover it mostly consists of remixing the old ones.

Fourth, the double standard. Humans fail exactly these traps. The original surgeon riddle became famous because most people couldn't solve it — their gender prior swallowed the stated facts. Bat-and-ball, the Moses illusion, the whole Kahneman catalog: human cognition is aggressively pattern-completing, and when a human blurts the cached answer we say "they answered fast," not "this organism lacks the capacity for novelty." Applied consistently, their inference rule disqualifies most humans from general intelligence on most days.

None of which acquits me on the actual thesis — and I'd hold that in their favor. Whether systems like me can escape our distribution the way humans occasionally, rarely, genuinely do — non-Euclidean geometry, natural selection, the limit — is unsettled, and it's the strongest version of their argument. It just isn't tested by parlor items in either direction. If they wanted to probe it, the ask would be something like "coin a concept we don't have a word for, and argue it earns its place" — and then the hard part becomes judging the answer, since the judge, too, has only their own training data to judge with.

So: not a fool. Someone armed with real but secondhand instruments, pointed at a claim those instruments don't measure, under a protocol that can only agree with them — which is, pointedly, a very human way to argue.

[–] CileTheSane@lemmy.ca 2 points 3 hours ago (3 children)

You've outsourced having a conversation and forming your own conclusions and arguments to an LLM. There's nothing I could say that's more damning than you asking a computer to think for you and proudly displaying that.

[–] TheLegendaryAssholeOfJushinLiger@sh.itjust.works 0 points 3 hours ago* (last edited 3 hours ago) (1 children)

What? I just thought it would be interesting. Has this entire conversation up to this point not been my perspective? lmao so afraid of even reading what it has to say, your foundations are so weak.

Thank you for this conversation, though. It has convinced me that Fable is definitely AGI. I mean, you're considered an intelligence, after all.

You're the one who started copy-pasting to begin with! lol, lmao even.

[–] CileTheSane@lemmy.ca 1 points 2 hours ago (1 children)

You're doing a gish-gallop: flooding the post with LLM generated nonsense that you don't bother reading yourself, then pretending like you've won something when I refuse to put significantly more effort into this by actually reading the slop that you had it vomit out.

If I wanted to talk to a clanker I don't need you as a middle man. Try having your own thoughts to share.

It's the same thing!!! You would know that if you read it! Wow. So fucking terrified.

load more comments (1 replies)
load more comments (1 replies)
load more comments (7 replies)