Who Cited It

Language Models are Few-Shot Learners

2020 · arXiv (Cornell University) · 3,020 citations · 2 from inside this corpus

T. B. Brown, Benjamin Mann low, Nick Ryder low, Melanie Subbiah low, Jared Kaplan low, Prafulla Dhariwal low, Arvind Neelakantan low, Pranav Shyam low, Girish Sastry low, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss low, Gretchen Krueger low, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler low, Jeffrey Wu, Clemens Winter low, Christopher Hesse low, Mark Chen, Eric J. Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark low, Christopher Berner low, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei

The source holds an abstract for this work, but its best open-access copy is under no open licence, which does not permit us to republish the text. Read it at the source below.

Language Models are Few-Shot Learners (2020)Language Models are Few-Shot …Glove: Global Vectors for Word Representation (2014)Glove: Global Vectors for Wor…BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2019)BERT: Pre-training of Deep Bi…HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanit… (2019)HISTORIAE, History of Socio-C…Distilling the Knowledge in a Neural Network (2015)Distilling the Knowledge in a…Efficient Estimation of Word Representations in Vector Space (2013)Efficient Estimation of Word …Multitask Learning (1997)Multitask LearningDistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter (2019)DistilBERT, a distilled versi…Affordance-Compiled Intelligence: Observable-Only Cognitive Impedance Matching for No-Met… (2026)Affordance-Compiled Intellige…Domain randomization for transferring deep neural networks from simulation to the real wo… (2017)Domain randomization for tran…SentiWordNet 3.0: An Enhanced Lexical Resource for Sentiment Analysis and Opinion Mining (2010)SentiWordNet 3.0: An Enhanced…[No title in the source record — Edinburgh Research Explorer (University of Edinburgh)][No title in the source recor…Optimization as a Model for Few-Shot Learning (2017)Optimization as a Model for F…Know What You Don’t Know: Unanswerable Questions for SQuAD (2018)Know What You Don’t Know: Una…Natural Questions: A Benchmark for Question Answering Research (2019)Natural Questions: A Benchmar…Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation (2013)Estimating or Propagating Gra…NLTK: The Natural Language Toolkit (2002)NLTK: The Natural Language To…TinyBERT: Distilling BERT for Natural Language Understanding (2020)TinyBERT: Distilling BERT for…The PASCAL Recognising Textual Entailment Challenge (2006)The PASCAL Recognising Textua…Cross-lingual Language Model Pretraining (2019)Cross-lingual Language Model …Unsupervised Data Augmentation for Consistency Training (2019)Unsupervised Data Augmentatio…Semantic Parsing on Freebase from Question-Answer Pairs (2013)Semantic Parsing on Freebase …Scaling Laws for Neural Language Models (2020)Scaling Laws for Neural Langu…LoRA Fine-Tuning of a 3B Code LLM for Algorithmic Efficiency (2021)LoRA Fine-Tuning of a 3B Code…On the Opportunities and Risks of Foundation Models (2021)On the Opportunities and Risk…
24 of 24 neighbouring works in this corpus. Blue is what this paper cites; orange is what cites it, and a dashed line is one neighbour citing another. Only the largest labels are drawn — every node carries its full title on hover.
this paper works it cites works citing it node size = global citations · hover for the full title

What this paper cites, inside the corpus

PaperYearCited
Glove: Global Vectors for Word Representation201434,067
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding201933,416
HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanit…201917,489
Distilling the Knowledge in a Neural Network201514,099
Efficient Estimation of Word Representations in Vector Space201311,714
Multitask Learning19976,478
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter20194,600
Affordance-Compiled Intelligence: Observable-Only Cognitive Impedance Matching for No-Met…20263,059
Domain randomization for transferring deep neural networks from simulation to the real wo…20172,890
SentiWordNet 3.0: An Enhanced Lexical Resource for Sentiment Analysis and Opinion Mining20102,678
[No title in the source record — Edinburgh Research Explorer (University of Edinburgh)]2,485
Optimization as a Model for Few-Shot Learning20172,436
Know What You Don’t Know: Unanswerable Questions for SQuAD20182,196
Natural Questions: A Benchmark for Question Answering Research20192,095
Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation20131,996
NLTK: The Natural Language Toolkit20021,943
TinyBERT: Distilling BERT for Natural Language Understanding20201,706
The PASCAL Recognising Textual Entailment Challenge20061,672
Cross-lingual Language Model Pretraining20191,624
Unsupervised Data Augmentation for Consistency Training20191,624
Semantic Parsing on Freebase from Question-Answer Pairs20131,618
Scaling Laws for Neural Language Models20201,537

What cites it, inside the corpus

Topics

Topic ModelingComputer Science
Natural Language Processing TechniquesComputer Science
Text Readability and SimplificationComputer Science

Is this record sound?

partial

One field of this record is missing or disagrees with another. What is shown below is what the source publishes.

  • supports31 author record(s) attached.
  • supports127 reference(s) recorded.
  • weakensThe DOI names 2005 but the record dates this to 2,020. One of the two is about a different paper.
  • supportsA title is present.

Provenance

Everything above was read from one stored OpenAlex payload, fetched 2026-09-04T03:58:47+00:00.

sha256 a08467ae9504f237…