
MMLU
MMLU(Massive Multitask Language Understanding)是由加州大学伯克利分校于2020年9月推出的基准测试,旨在评估大型语言模型在多领域的理解和推理能力。
AGI-Eval is a large model evaluation community jointly launched by Shanghai Jiao Tong University, Tongji University, East China Normal University, and DataWhale, dedicate...
In today’s rapidly developing AI era, objectively and fairly evaluating various large models has become a focal point in the industry. AGI-Eval, jointly launched by Shanghai Jiao Tong University, Tongji University, East China Normal University, and DataWhale, is a large model evaluation community established to meet this demand.







