All Datasets

Search
all
AIOZ AI
CommonGen

CommonGen

Building machines with commonsense to compose realistically plausible sentences is challenging. CommonGen is a constrained text generation task, associated with a benchmark dataset, to explicitly test machines for the ability of generative commonsense reasoning. Given a set of common concepts; the task is to generate a coherent sentence describing an everyday sce- nario using these concepts.

user-avatar
300
158
BLiMP

BLiMP

The Benchmark of Linguistic Minimal Pairs, a challenge set for evaluating the linguistic knowledge of language models (LMs) on major grammatical phenomena in English, finds that state-of-the-art models identify morphological contrasts related to agreement reliably, but they struggle with some subtle semantic and syntactic phenomena.

user-avatar
306
154
TAL-SCQ5K

TAL-SCQ5K

TAL-SCQ5K are high-quality mathematical competition datasets created by TAL Education Group.

user-avatar
282
151
X-CSR

X-CSR

To create these datasets, the authors automatically translated the original CSQA and CODAH datasets, originally available only in English, into 15 other languages.

user-avatar
263
156
KLUE

KLUE

A native-Korean NLU benchmark covering eight tasks — topic classification, semantic similarity, NLI, NER, relation extraction, dependency parsing, machine reading comprehension, and dialogue state tracking. Over 200,000 labeled examples built from real Korean corpora rather than machine translation.

user-avatar
97
34
DOCCI

DOCCI

The DOCCI dataset consists of comprehensive descriptions on 15k images specifically taken with the objective of evaluating T2I and I2T models. These cover a lot of key details in the images, as illustrated below.

user-avatar
288
154
AI2_Reasoning_Challenge

AI2 Reasoning Challenge

The ARC dataset consists of 7,787 science exam questions drawn from a variety of sources, including science questions, provided under license by a research partner affiliated with AI2.

user-avatar
294
154
LongBench

LongBench

A bilingual (English + Chinese) benchmark for long-context understanding, with 21 tasks across document QA, summarization, few-shot learning, synthetic reasoning, and code completion. 4,750 test samples averaging 5k–15k tokens, scored with a fully automated evaluation.

user-avatar
84
42
MInDS-14

MInDS-14

A multilingual spoken-language dataset for intent detection in e-banking, covering 14 intents across 14 language varieties. 8,168 labeled audio clips at 8 kHz, each paired with a transcription and English translation.

user-avatar
104
33
PLOD_An_Abbreviation_Detection_Dataset

PLOD: An Abbreviation Detection Dataset

This is the repository for PLOD Dataset subset being used for CW in NLP module 2023-2024 at University of Surrey.

user-avatar
279
155
BIG-bench

BIG-bench

A collaborative benchmark of 204 tasks built to probe large language models beyond single-metric tests, spanning reasoning, knowledge, creativity, and multilingual ability. Contributed by 450 authors across 132 institutions, with canary strings to detect training-data contamination.

user-avatar
115
35
MMLU

MMLU

A multiple-choice benchmark measuring a model's knowledge across 57 subjects in the humanities, social sciences, STEM, and professional domains. Roughly 16,000 exam-sourced questions in a consistent four-option format, the field's common reference for general-knowledge evaluation.

user-avatar
70
37
NIH_Chest_X_ray

NIH Chest X-Ray

NIH Chest X-Ray is a large dataset containing chest X-ray images of patients collected by the National Institutes of Health (NIH) of the United States.

user-avatar
378
156
MNIST

MNIST

MNIST is used to train and evaluate image classification models in complex tasks.

user-avatar
310
151
1
2