Virginia Brassesco, Herman Schinca and Fernando Schapachnik from the University of Buenos Aires in Argentina wanted to find out if smart kids, in this case a sample of 146 first year STEM students from Argentina and China, were overestimating the capability of Generative Artificial Intelligence (GenAI) chatbots such as ChatGPT.
To find out they devised an experiment where students were first given a standardized reasoning test of ten questions and asked to complete it and assigned a score immediately thereafter. On the test 76% of students got either 9/10 or 10/10.
They were then asked what if ChatGPT were given the same test, how did they think it would score? The median response was that ChatGPT would get 7/10, in fact it got only 1/10.
There was a tiny bias (but with so small a sample it’s probably not relevant) of China students to think ChatGPT was smarter but other than that there was no real difference in terms of sex and other background variables.
The result is especially interesting as most students have first hand and day-to-day interaction with GenAI. The researchers suggest perhaps the daily tasks most assign to their AI buddies are not actually that hard centering around data retrieval, translation and writing. What may be happening then is these simple tasks are being conflated with problem solving.
Despite the small sample size and the way in which the study was slightly gamed (questions known in advance to be problematic for AI were deliberately chosen for the test) the conclusion is clear. If science and engineering students are unaware at how poor AI is at problem solving how are the rest of the general public to know any better?
You can access the paper in full via the following link Would ChatGPT outperform your outstanding performance?
Happy Sunday.