known·good search mcp servers how we grade about the index

search / category / interactive reasoning benchmark

Agent-ready interactive reasoning benchmark

1 sites in this index, 1 of them transactional. Every one probed live, with the date it was checked.

ARC-AGI-3 Docs docs.arcprize.org
Documentation for ARC-AGI-3, an interactive reasoning benchmark measuring AI agent generalization in novel env