DreamBench++
A benchmark for evaluating the performance of large language models (LLMs) in various tasks related to both textual and visual imagination.
OPEN IT →Every llm link in the ranking, ordered by what people actually open. All checked automatically.
A benchmark for evaluating the performance of large language models (LLMs) in various tasks related to both textual and visual imagination.
OPEN IT →A large-scale question-answering benchmark focused on real-world financial data, integrating both tabular and textual information.
OPEN IT →A benchmark for evaluating AI models across multiple academic disciplines like math, physics, chemistry, biology, and more.
OPEN IT →A biomedical question-answering benchmark designed for answering research-related questions using PubMed abstracts.
OPEN IT →A benchmark that evaluates large multimodal models (LMMs) on their ability to perform human-like mathematical reasoning.
OPEN IT →A ground-truth-based dynamic benchmark derived from off-the-shelf benchmark mixtures, which evaluates LLMs with a highly capable model ranking (i.e., 0.96 correlation with Chatbot Arena) while running locally and quickly (6% the time and cost of running MMLU).
OPEN IT →Benchmark designed to evaluate large language models (LLMs) on solving complex, college-level scientific problems from domains like chemistry, physics, and mathematics.
OPEN IT →A meta-benchmark that evaluates how well factuality evaluators assess the outputs of large language models (LLMs).
OPEN IT →A benchmark designed to evaluate large language models (LLMs) specifically in their ability to answer real-world coding-related questions.
OPEN IT →An Automatic Evaluator for Instruction-following Language Models using Nous benchmark suite.
OPEN IT →A benchmark that evaluates large language models on a variety of multimodal reasoning tasks, including language, natural and social sciences, physical and social commonsense, temporal reasoning, algebra, and geometry.
OPEN IT →A benchmark that evaluates large language models' ability to answer medical questions across multiple languages.
OPEN IT →A large-scale Document Visual Question Answering (VQA) dataset designed for complex document understanding, particularly in financial reports.
OPEN IT →Exploring the Intelligence of Large-scale MoE Model.
OPEN IT →A benchmark dataset testing AI's ability to reason about visual commonsense through images that defy normal expectations.
OPEN IT →CompassRank is dedicated to exploring the most advanced language and visual models, offering a comprehensive, objective, and neutral evaluation reference for the industry and research.
OPEN IT →A paid product for tracking model training and prompt engineering experiments.
OPEN IT →A Python library for validating outputs and retrying failures. Still in alpha, so expect sharp edges and bugs.
OPEN IT →A bilingual Chinese-English knowledge extraction model with knowledge graphs and natural language processing technologies.
OPEN IT →A paid product for testing and improving prompts.
OPEN IT →