LLMEval

LLMEval

Comprehensive, contamination-resistant LLM evaluation for research and development.

Visit LLMEval

About LLMEval

LLMEval is an academic initiative from Fudan NLP Lab, providing comprehensive and rigorous evaluation frameworks for large language models (LLMs). It offers curated, contamination-resistant benchmarks covering over 13 academic disciplines, medical AI, and logical reasoning, with a database of 220,000+ questions and a robust longitudinal evaluation pipeline. The platform is designed for AI researchers, practitioners, and developers seeking accurate, robust, and fair performance analysis of LLMs.

Resources

Product Website

Visit LLMEval's official website for product details and getting started.

Visit website →

Documentation

Comprehensive guides and API references for utilizing LLMEval effectively.

View docs →

Blog

Insights and updates on the latest developments and research related to LLMEval.

Read blog →

Community Forum

A space for users to discuss features, share experiences, and seek assistance.

Join community →