LLMRank is an innovative framework for evaluating large language models using a peer-based cross-evaluation and PageRank algorithm. It provides transparent and flexible benchmarking, allowing for unbiased model ranking across diverse or specific prompts, and is easily scalable for varied AI systems.
LLMRank, an innovative evaluation framework, leverages PageRank in conjunction with peer-based cross-evaluation techniques to rank Large Language Models (LLMs). This system promotes fairness, dynamism, and scalability in benchmarking AI models, ensuring transparency and encouraging progress in artificial intelligence innovation.
MODEL_NAMES and raw_prompts lists for targeted evaluations.USE_SUBSET_EVALUATION feature to limit evaluation to a subset of models, saving on API costs.Contributions are welcome! Suggestions, issues, or pull requests are encouraged to further develop and enhance the LLMRank framework.
Acknowledgment goes to SimonW for the invaluable llm library, as well as the broader AI research community for their continuous support and contributions.
No comments yet.
Sign in to be the first to comment.