Human Benchmark — Test Your Brain
Reaction time, memory, typing and aim measured against millions of people.
OPEN IT →Research-grade personality models and reaction-time benchmarks, no email required.
Reaction time, memory, typing and aim measured against millions of people.
OPEN IT →Free 16-personality type assessment taken by over 100 million people worldwide.
OPEN IT →A benchmark for evaluating the performance of large language models (LLMs) in various tasks related to both textual and visual imagination.
OPEN IT →The personality model psychologists actually use. Free, no email required.
OPEN IT →
TOOLS · PYTHON(Python standard library) Unit testing framework.
OPEN IT →The official taster test. Take it once and then argue about it forever.
OPEN IT →A large-scale question-answering benchmark focused on real-world financial data, integrating both tabular and textual information.
OPEN IT →A benchmark for evaluating AI models across multiple academic disciplines like math, physics, chemistry, biology, and more.
OPEN IT →
STRANGE TECH · SCIENCEMIT Lincoln Laboratory, in partnership with NASA Goddard Space Flight Center, built and designed a communications payload, TBIRD, which is demonstrating space-to-ground laser communications at unprecedented data rates.
OPEN IT →Research on evaluation of LLMs conducted by Microsoft Research and other collaborated institutes. (Updated at: 2023/10)
OPEN IT →A biomedical question-answering benchmark designed for answering research-related questions using PubMed abstracts.
OPEN IT →A benchmark that evaluates large multimodal models (LMMs) on their ability to perform human-like mathematical reasoning.
OPEN IT →An interactive environment to create and test Go templates.
OPEN IT →A ground-truth-based dynamic benchmark derived from off-the-shelf benchmark mixtures, which evaluates LLMs with a highly capable model ranking (i.e., 0.96 correlation with Chatbot Arena) while running locally and quickly (6% the time and cost of running MMLU).
OPEN IT →