

I agree with most of this. I’ve also said in the past that LLMs cannot think, and I think that’s still true for most models. The reason ARC-AGI-3 is interesting is that it was specifically designed to test reasoning, adaptability, novel problem solving, planning, memory, etc. So it was a surprise to me that Astra was able to defeat it so effectively, and that Astra invents algebras for each novel task.
But I agree we can’t trust OpenAI if these results are self-reported, and we may not be able to trust the ARC Prize Foundation fully either. Extraordinary claims require extraordinary evidence, so we need replication, transparency, and proper open science to confirm things.
I also agree with ARC Prize’s conclusion, that there are still capabilities any AI system would need to demonstrate before we can claim a full general intelligence.
























Yes, it was released in March. One important point is that these results are based on the semi-private test set. Perhaps it’s best to wait until it is measured against the fully private test set, but that might not happen for months.