Technology

85357 readers

3848 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 3 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws

714

Car Wash Test on 53 leading AI models: "I want to wash my car. The car wash is 50 meters away. Should I walk or drive?" (opper.ai)

submitted 3 months ago by fubarx@lemmy.world to c/technology@lemmy.world

330 comments fedilink hide all child comments

Screenshot of this question was making the rounds last week. But this article covers testing against all the well-known models out there.

Also includes outtakes on the 'reasoning' models.

you are viewing a single comment's thread
view the rest of the comments

[–] tigeruppercut@lemmy.zip 9 points 3 months ago (3 children)

But natural language in service of what? If they can't produce answers that are correct, what's the point of using them? I can get wrong answers anywhere.

[–] Threeme2189@sh.itjust.works 1 points 3 months ago (1 children)

As OP said, LLMs are really good at generating text that is fluid and looks natural to us. So if you want that kind of output, LLMs are the way to go.
Not all LLM prompts ask factual questions and not all of the generated answers need to be correct.
Are poems, songs, stories or movie scripts 'correct'?

I'm totally against shoving LLMs everywhere, but they do have their uses. They are really good at this one thing.

[–] tigeruppercut@lemmy.zip 5 points 3 months ago* (last edited 3 months ago) (2 children)

Are poems, songs, stories or movie scripts ‘correct’?

It's a valid point that they can produce natural language. The Turing Test has been a thing for awhile after all. But while the language sounds natural, can they create anything meaningful? Are the poems or stories they make worth anything? It's not like humans don't create shitty art, so I guess generating random soulless crap is similar to that.

The value of language produced by something that can't understand the reason for language is an interesting question I suppose.

[–] Threeme2189@sh.itjust.works 4 points 3 months ago

I'm with you on that. I've come to realize that I value a shitty stick figure that was drawn by a human much more than an AI generated 'Mona Lisa'.

[–] iopq@lemmy.world 1 points 3 months ago (1 children)

There are people out there whose job is to format promotional emails for companies. AIs can replace this kind of soulless work completely. We should applaud that.

[–] snooggums@piefed.world 3 points 3 months ago

No, we don't need to applaud automation of spam.

[–] Iconoclast@feddit.uk 0 points 3 months ago (1 children)

I'm not here defending the practical value of these models. I'm just explaining what they are and what they're not.

[–] XLE@piefed.social 5 points 3 months ago (1 children)

You're definitely running around Lemmy defending AI, Iconoclast... Might as well be honest about it

[–] Iconoclast@feddit.uk -3 points 3 months ago (1 children)

I'm not really interested in engaging in discussions about what you or anyone else thinks my underlying motives are. You're free to point out any factual inaccuracies in my responses, but there's no need to make it personal and start accusing me of being dishonest.

[–] XLE@piefed.social 4 points 3 months ago

Your motivations are self-evident, I'm just pointing them out because you are misrepresenting them here

[–] iopq@lemmy.world -3 points 3 months ago

Some of them can produce the correct answer. Of we do the test next year and they do better than humans then, isn't it progress?