Skip to content

Entry

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Appears in 5 awesome lists

bidirectional transformer pretraining; foundation for most encoder-based NLP work since 2018. Read online with section navigation and the ACL source attached.

Open arxiv.org

Found in these lists

Awesome AI Papers

Section: Historical Papers

SlowScore 51

awesome-nlp

Section: Pretraining and Adaptation · bidirectional transformer pretraining; foundation for most encoder-based NLP work since 2018. Read online with section navigation and the ACL source attached.

FreshScore 90

Awesome Question Answering

Section: Recent Language Models · , Jacob Devlin, et al., NAACL 2019, 2018.

StaleScore 48

Awesome Transformer & Transfer Learning in NLP

Section: Papers · by Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova.

StaleScore 52

awesome-chatgpt

Section: The technical principle of ChatGPT · This paper ushered in the era of pre-training in NLP. BERT came out of nowhere.

StaleScore 52

ChatGPT

Announcement of ChatGPT, a conversational model trained to answer follow-up questions, admit mistakes, challenge incorrect premises, and reject inappropriate requests. OpenAI blog, November 30, 2022.

In 8 listsDetails

Qwen

(Alibaba, 2024-2025) - strong multilingual coverage, especially Chinese; often top open model on multilingual benchmarks.

In 3 lists

T5

and FLAN-T5 - text-to-text framing for NLP tasks; strong instruction-tuned encoder-decoder baselines.

In 3 lists

AlphaFold: Protein Structure Prediction

Nature, 2021. [All Versions]. This paper provides the first computational method that can regularly predict protein structures with atomic accuracy even in cases in which no similar structure is known. This approach is a canonical application of observation- and explanation- based method for…

In 3 lists

Llama 3 / 3.1 / 3.3

(Meta, 2024-2025) - widely adopted open-weight family; default base for fine-tuning across NLP tasks.

In 2 lists

RoBERTa

robustly optimized BERT pretraining; common encoder baseline.

In 2 lists

OLMo 2

(AI2, 2025) - fully open: weights, training data, code; reproducibility benchmark.

In 2 lists

What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization?

encoder vs decoder vs encoder-decoder for NLP transfer.