Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed, Katherine M. Collins, Yael Moros-Daval, Seraphina Zhang, Qinlin Zhao, Yitian Huang, Luning Sun, Jonathan E. Prunty, Zongqian Li, Pablo Sánchez-García, Kexin Jiang Chen, Pablo A. M. Casares, Jiyun Zu, John Burden, Behzad Mehrbakhsh, David Stillwell, Manuel Cebrian, Jindong Wang, Peter Henderson, Sherry Tongshuang Wu, Patrick C. Kyllonen, Lucy Cheke, Xing Xie, José Hernández-Orallo
Save this paper to a shelf, write a review, and keep your own notes.
Paper of the year. Actual psychometrics come to machine learning. After 60 years, item response theory (Rausch 1960) finally shows up in ML. This lets us put benchmarks on a common scale and actually estimate latent capabilities. Given an ability dimension, the ADeLe rule of thumb defines every level L as the capability held by 1 in 10^L humans on Earth. Current systems are 1 in 100,000 on some things, 1 in 1000 on others. GPT-2 to GPT-3 was a 0.6 level jump. ADeLE is fully-automated, using LLM as a judge. It explains the abilities a benchmark is actually measuring, gives you an interpretable ability profile for an AI, and predicts OOD performance on new task instances better than embedding and finetunes (AUROC=0.8). Their "volume" ability pre-dates the HCAST task horizon, and includes it as a special case among 18 other abilities. They throw in a guessability control as well! Not enough for ya????? How about “the very first scaling laws of the actual abilities of LLMs” too See here for why you can't just use CHC and compare to humans as a percentage: https://aievaluation.substack.com/p/is-the-definition-of-agi-a-percentage