Who Cited It

GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

2018 · 4,055 citations · 11 from inside this corpus

Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, Samuel Bowman

Human ability to understand language is general, flexible, and robust. In contrast, most NLU models above the word level are designed for a specific task and struggle with out-of-domain data. If we aspire to develop models with understanding beyond the detection of superficial correspondences between inputs and outputs, then it is critical to develop a unified model that can execute a range of linguistic tasks across different domains. To facilitate research in this direction, we present the General Language Understanding Evaluation (GLUE, gluebenchmark.com): a benchmark of nine diverse NLU tasks, an auxiliary dataset for probing models for understanding of specific linguistic phenomena, and an online platform for evaluating and comparing models. For some benchmark tasks, training data is plentiful, but for others it is limited or does not match the genre of the test set. GLUE thus favors models that can represent linguistic knowledge in a way that facilitates sample-efficient learning and effective knowledge-transfer across tasks. While none of the datasets in GLUE were created from scratch for the benchmark, four of them feature privately-held test data, which is used to ensure that the benchmark is used fairly. We evaluate baselines that use ELMo (Peters et al., 2018), a powerful transfer learning technique, as well as state-of-the-art sentence representation models. The best models still achieve fairly low absolute scores. Analysis with our diagnostic dataset yields similarly weak performance over all phenomena tested, with some exceptions.

GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding (2018)GLUE: A Multi-Task Benchmark …Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank (2013)Recursive Deep Models for Sem…SQuAD: 100,000+ Questions for Machine Comprehension of Text (2016)SQuAD: 100,000+ Questions for…Supervised Learning of Universal Sentence Representations from Natural\n Language Inferen… (2017)Supervised Learning of Univer…The PASCAL Recognising Textual Entailment Challenge (2006)The PASCAL Recognising Textua…Transformers: State-of-the-Art Natural Language Processing (2020)Transformers: State-of-the-Ar…On the Dangers of Stochastic Parrots (2021)On the Dangers of Stochastic …A Survey on Evaluation of Large Language Models (2024)A Survey on Evaluation of Lar…LXMERT: Learning Cross-Modality Encoder Representations from Transformers (2019)LXMERT: Learning Cross-Modali…Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing (2021)Domain-Specific Language Mode…Language Models as Knowledge Bases? (2019)Language Models as Knowledge …TinyBERT: Distilling BERT for Natural Language Understanding (2020)TinyBERT: Distilling BERT for…Text Data Augmentation for Deep Learning (2021)Text Data Augmentation for De…Deep Learning--based Text Classification (2021)Deep Learning--based Text Cla…Pre-trained models for natural language processing: A survey (2020)Pre-trained models for natura…ERNIE: Enhanced Language Representation with Informative Entities (2019)ERNIE: Enhanced Language Repr…
15 of 15 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

What cites it, inside the corpus

Topics

Topic ModelingComputer Science
Natural Language Processing TechniquesComputer Science
Multimodal Machine Learning ApplicationsComputer Science

Is this record sound?

complete

Nothing in this record contradicts itself and no field we check is missing.

  • supports6 author record(s) attached.
  • supports19 reference(s) recorded.
  • neutralThe DOI carries no year to check against.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:46+00:00.

sha256 bba2969b3567609a…