If you read what ARC-AGI has stated from day one, the tests are designed to stress frontier models in tasks that humans are uniquely good at. When the models get to 100%, the next set of tasks is deployed. When they run out of ideas for how to stress a model (i.e. no more tests), THAT is when AGI is achieved. It's a great definition.