10 tests per model seems like way too little and they should give confidence intervals…
the 10/10 vs. 8/10 is just as likely due chance than any real difference. But some people will definitely use this to justify model choice.
This is a most excellent place for technology news and articles.
10 tests per model seems like way too little and they should give confidence intervals…
the 10/10 vs. 8/10 is just as likely due chance than any real difference. But some people will definitely use this to justify model choice.
It should get it wrong 0% of the time because it is a computer that should have predictable results about basic things like requiring a car to be present to be washed.
Did this say whether the reasoning models get this right more than the others? Was curious about that but missed it if it was mentioned.