Follow your favourite podcasts, listen offline and in the car with CarPlay and Android Auto, and always pick up where you left off. Free to try.
A new benchmark called Humanity's Last Exam is redefining how we measure artificial intelligence. Designed with 2,500 highly specialized questions across fields like advanced mathematics, ancient languages, and natural sciences, the test aims to challenge even the most powerful AI systems.
Unlike traditional benchmarks, it focuses on deep expertise rather than searchable facts. Early results suggest that despite rapid progress, a significant gap still exists between machine pattern recognition and true human-level knowledge.