AI labs often tout impressive benchmark scores when releasing new models, showing better capabilities in areas like coding, computer use, and more than their competitors. However, those benchmarks aren’t always a reliable measure of what AI can do because they’re easily beaten by exponentially improving models and can emphasize marketing over actual performance.