Read the original at HF Daily Papers
Researchers introduce AutoSciBench, a framework that automatically generates and iteratively adapts scientific agent benchmarks to maintain evaluation relevance as agent capabilities evolve.
Carried by: HF Daily Papers. First seen: .