EVERYHUBToday's ranking →
20 CHECKED LINKS · UPDATED DAILY

Llm

Every llm link in the ranking, ordered by what people actually open. All checked automatically.

AI · LLM

DreamBench++

A benchmark for evaluating the performance of large language models (LLMs) in various tasks related to both textual and visual imagination.

OPEN IT →
AI · LLM

TAT-QA

A large-scale question-answering benchmark focused on real-world financial data, integrating both tabular and textual information.

OPEN IT →
AI · LLM

OlympicArena

A benchmark for evaluating AI models across multiple academic disciplines like math, physics, chemistry, biology, and more.

OPEN IT →
AI · LLM

PubMedQA

A biomedical question-answering benchmark designed for answering research-related questions using PubMed abstracts.

OPEN IT →
AI · LLM

We-Math

A benchmark that evaluates large multimodal models (LMMs) on their ability to perform human-like mathematical reasoning.

OPEN IT →
AI · LLM

MixEval

A ground-truth-based dynamic benchmark derived from off-the-shelf benchmark mixtures, which evaluates LLMs with a highly capable model ranking (i.e., 0.96 correlation with Chatbot Arena) while running locally and quickly (6% the time and cost of running MMLU).

OPEN IT →
AI · LLM

SciBench

Benchmark designed to evaluate large language models (LLMs) on solving complex, college-level scientific problems from domains like chemistry, physics, and mathematics.

OPEN IT →
AI · LLM

FELM

A meta-benchmark that evaluates how well factuality evaluators assess the outputs of large language models (LLMs).

OPEN IT →
AI · LLM

InfiBench

A benchmark designed to evaluate large language models (LLMs) specifically in their ability to answer real-world coding-related questions.

OPEN IT →
AI · LLM

AlpacaEval

An Automatic Evaluator for Instruction-following Language Models using Nous benchmark suite.

OPEN IT →
AI · LLM

M3CoT

A benchmark that evaluates large language models on a variety of multimodal reasoning tasks, including language, natural and social sciences, physical and social commonsense, temporal reasoning, algebra, and geometry.

OPEN IT →
AI · LLM

MMedBench

A benchmark that evaluates large language models' ability to answer medical questions across multiple languages.

OPEN IT →
AI · LLM

TAT-DQA

A large-scale Document Visual Question Answering (VQA) dataset designed for complex document understanding, particularly in financial reports.

OPEN IT →
AI · LLM

Qwen2.5-Max

Exploring the Intelligence of Large-scale MoE Model.

OPEN IT →
AI · LLM

WHOOPS!

A benchmark dataset testing AI's ability to reason about visual commonsense through images that defy normal expectations.

OPEN IT →
AI · LLM

CompassRank

CompassRank is dedicated to exploring the most advanced language and visual models, offering a comprehensive, objective, and neutral evaluation reference for the industry and research.

OPEN IT →
AI · LLM

Weights & Biases

A paid product for tracking model training and prompt engineering experiments.

OPEN IT →
AI · LLM

Guardrails.ai

A Python library for validating outputs and retrying failures. Still in alpha, so expect sharp edges and bugs.

OPEN IT →
AI · LLM

OneKE

A bilingual Chinese-English knowledge extraction model with knowledge graphs and natural language processing technologies.

OPEN IT →
AI · LLM

PromptPerfect

A paid product for testing and improving prompts.

OPEN IT →