The official GitHub repository of the paper "Recent advances in large language model benchmarks against data contamination: From static to dynamic evaluation"
-
Updated
Aug 11, 2026
The official GitHub repository of the paper "Recent advances in large language model benchmarks against data contamination: From static to dynamic evaluation"
EvaLearn is a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks.
A decentralized, adversarial + dynamic AI evaluation protocol on Bittensor. Combats benchmark saturation by measuring genuine intelligence through dynamic, zero-shot generalization tasks.
To associate your repository with the dynamic-evaluation topic, visit your repo's landing page and select "manage topics."