Theme
BERT and the Encoder Family
BERT (Bidirectional Encoder Representations from Transformers) is a bidirectional representation model based on the Transformer encoder (encoder-only) with "masked language modeling (MLM)" as its pretraining task, proposed by Google in 2018. It ignited the "pretraining + fine-tuning" paradigm alongside GPT, but diverged from GPT in architecture: GPT went "decoder + generation," while BERT went "encoder + understanding." This page breaks down BERT's principles, family evolution, and its position in the LLM era; for underlying architecture mechanisms, see Transformer Architecture Explained; for a complete comparison with the GPT path, see The GPT Series: From GPT-1 to GPT-4o.
I. What Is BERT: A One-Sentence Definition
BERT = bidirectional encoder + masked language modeling (MLM) pretraining + downstream fine-tuning. It pretrains by "randomly masking words in sentences and having the model guess them from the bidirectional context on both sides," yielding deep bidirectional text representations; then fine-tunes with a lightweight output head on each downstream task (classification, extraction, QA, similarity).
Pretraining objective (MLM):
Input: China's [MASK] is the capital, with a population exceeding [MASK]0 million.
Predict: [MASK]₁ → "capital" [MASK]₂ → "1" (based on bidirectional context)
Downstream fine-tuning (e.g., sentiment classification):
[CLS] This movie is great [SEP] → output layer → positiveThe BERT paper, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, was posted to arXiv in October 2018 (and won NAACL best long paper in 2019). It set new SOTA on 11 NLP benchmarks, appeared almost simultaneously with GPT-1, and together they established the "pretraining + fine-tuning" era.
II. Principle Breakdown: Three Key BERT Designs
1. Bidirectional: The Core Difference from GPT
GPT uses causal masking — each position can only "see the left" (unidirectional autoregressive); BERT has no causal masking, and each position sees the full context (bidirectional). Intuitively, understanding tasks (judging sentiment, extracting entities) heavily depend on the following context — "this movie isn't that good" requires seeing the "not" before knowing it's negative. Bidirectionality is BERT's fundamental advantage over same-scale GPT on understanding tasks.
2. MLM: Masked Language Modeling
The mask strategy's details (from the original paper): randomly select 15% of tokens for prediction, of which 80% are replaced with [MASK], 10% are replaced with random words, and 10% are kept unchanged. The motivation: during both pretraining and fine-tuning, the model never sees [MASK], forcing it to learn contextual semantics rather than "seeing [MASK] and guessing a word."
The 80/10/10 ratio is deliberately nuanced: if 100% were replaced with [MASK], the model would only learn "seeing [MASK] means guess a word," but during fine-tuning there are no [MASK] tokens — the pretraining and fine-tuning input distributions would be inconsistent. If 100% were kept unchanged, the model would have almost no prediction pressure. The balanced 80/10/10 lets the model both depend on contextual semantics and remain robust to "randomly replaced words." This design was later extended or revised by dynamic masking and word-level masking, but the core idea — "align the pretraining objective with the downstream distribution" — remains the central consideration in pretraining objective design (see Pretraining: Data and Objectives).
3. NSP: Next Sentence Prediction
To help the model understand "inter-sentence relationships" (important for QA and entailment tasks), BERT simultaneously trains a binary classification task: judging whether sentence B is the next sentence of sentence A (50% real / 50% random). Subsequent research found NSP contributed little (removing it in RoBERTa actually made it stronger), but it reflected the era's exploration of "inter-sentence structure."
4. Input Construction and Specifications
| Item | Specification |
|---|---|
| Input format | [CLS] sentenceA [SEP] sentenceB [SEP] + segment embedding + position embedding |
| Output head | [CLS] position vector for classification; remaining position vectors for token-level tasks |
| BERT-base | 110M parameters (12 layers / 768-dim / 12 heads) |
| BERT-large | 340M parameters (24 layers / 1024-dim / 16 heads) |
| Pretraining data | BookCorpus + English Wikipedia (~3.3 billion words) |
| Maximum input length | 512 tokens (position embedding limit) |
5. Results
BERT-large set new SOTA on 11 benchmarks at release, including GLUE (79.6 → 82.1), SQuAD 1.1 (F1 93.2, surpassing human level), and SQuAD 2.0. Its publicly available weights + open-source implementation made it realistic for "anyone to fine-tune a world-class NLP model in a few hours," directly catalyzing the explosion of the Hugging Face ecosystem.
III. BERT vs GPT: A Comparison of Two Paths
This is the most important architectural comparison of 2018–2023, worth revisiting:
The value of bidirectional attention varies significantly across tasks: for any task where "the answer depends on the full context" (entailment judgment, coreference resolution, QA), bidirectional has a clear advantage; for tasks where "order matters" (generation, translation, temporal reasoning), unidirectional/autoregressive is more natural. This is why encoder models remain strong baselines for natural language entailment and extraction tasks, while generation tasks have been entirely taken over by decoders. When choosing an architecture, don't ask "which is better" — ask "does my task rely more on global context or sequential generation?"
| Dimension | BERT | GPT |
|---|---|---|
| Architecture | Encoder (encoder-only) | Decoder (decoder-only) |
| Attention direction | Bidirectional (no causal mask) | Unidirectional (causal mask, sees only the left) |
| Pretraining objective | Masked language modeling (MLM) | Autoregressive next-token prediction |
| Suitable tasks | Understanding: classification, extraction, ranking, retrieval | Generation: continuation, translation, dialogue, code |
| Downstream adaptation | Add output head and fine-tune | Prompts / in-context learning / fine-tuning |
| Scale trajectory | Stalled at the hundred-million level (at most ~tens of billions) | Continuously grew to hundreds of billions and trillions |
| Era outcome | "Tool component" in the LLM era | "Protagonist" of the LLM era |
Why did GPT ultimately win?
Both objectives are fundamentally complementary: MLM makes representations more "understanding," while autoregression makes models more "writing-capable." But killer apps like dialogue, programming, and Agents all require generation; and once generation capability is paired with scale and alignment (see Alignment: RLHF and DPO), it naturally covers most understanding tasks. BERT's understanding advantage became insignificant as the absolute capability gap widened.
IV. The Encoder Family: BERT Improvements and Compression
A wave of improvements emerged within a year of BERT's release, targeting four directions: more data, longer training, fewer parameters, faster inference.
| Model | Time | Parameters | Key Changes | Results |
|---|---|---|---|---|
| BERT-large | 2018.10 | 340M | Baseline | GLUE 82.1 |
| RoBERTa | 2019.7 | 355M | Removed NSP, dynamic masking, 10× data (160 GB), 4× longer training | GLUE 88.5 |
| XLNet | 2019.6 | 340M | Permutation language modeling (autoregressive objective + bidirectional context) | Multiple SOTAs at the time |
| ALBERT | 2019.9 | ~12M (after sharing) | Cross-layer parameter sharing + factorized embeddings + SOP replacing NSP | Approached RoBERTa with far fewer parameters |
| DistilBERT | 2019.10 | 66M | Knowledge distillation (teacher: BERT-base) | Retained ~97% capability, ~60% faster inference |
| ELECTRA | 2020.3 | 330M | Replaced token detection (discriminative pretraining) | Matched RoBERTa with ¼ compute |
1. RoBERTa: A Victory for Data and Training Details
RoBERTa's findings are highly illuminating: BERT had significant room left to "train harder" — by increasing data from 16 GB to 160 GB, using dynamic masking, removing NSP, and training for 4× as many steps, GLUE went from 82.1 to 88.5. It proved that in the pretraining era, "data volume + training duration" matters as much as architectural innovation (this is isomorphic with the conclusion of Scaling Laws).
2. ALBERT and DistilBERT: Saving Parameters and Speeding Up
- ALBERT: Three moves to reduce parameters — cross-layer parameter sharing, embedding matrix factorization, and replacing NSP with sentence order prediction (SOP). Parameters dropped from the billions to the tens of millions, with only a slight capability loss.
- DistilBERT: A textbook example of knowledge distillation — training a 66M-parameter student using a teacher BERT's soft labels, retaining 97% of GLUE capability while being ~60% faster at inference. "Big models teaching small models" later became the mainstream approach for Building a Large Model from Scratch and distilling small models.
3. XLNet and ELECTRA: Two "Philosophical Variants"
- XLNet: Wanted both "autoregressive + bidirectional" — used permutations to let an autoregressive model see the full context during training, but didn't achieve a fundamental win.
- ELECTRA: Changed pretraining from "generative" to "discriminative" — having a generator mask tokens and a discriminator judge whether each position was replaced, achieving high computational efficiency and performance close to RoBERTa.
V. T5: Unifying Every Task as "Text-to-Text"
1. Core Idea
T5 (Text-to-Text Transfer Transformer, October 2019, Google) chose a third path: encoder-decoder, and unified every task (translation, summarization, classification, QA) into "input a piece of text, output a piece of text":
Input: "Translate to English: I love China" → Output: "I love China"
Input: "Sentiment: This movie is amazing" → Output: "positive"2. Specifications and Contributions
| Item | Specification |
|---|---|
| Largest version | T5-11B (encoder-decoder, ~11 billion parameters) |
| Training data | C4 (Colossal Clean Crawled Corpus, ~750 GB, cleaned Common Crawl) |
| Key innovation | Unified text-to-text framework, task prefix, C4 data cleaning methodology |
| Results | Multiple SOTAs on GLUE / SuperGLUE / SQuAD at release |
T5's other methodological legacy is the "task prefix + unified input/output" design, which made "using one model to handle hundreds of tasks" possible. This idea is in the same lineage as later instruction fine-tuning (see Fine-tuning: SFT and Parameter-Efficient Fine-tuning) — in fact, many instruction fine-tuning datasets treat task descriptions as T5-style prefixes. Understanding T5 helps explain why today's models can switch skills based on a single instruction: the task itself is part of the input, and the model's "generality" comes from generalizing over "task descriptions."
T5's contribution is methodological: it systematically validated the idea of "unified task representation" — the same pretrained model switches tasks through different "text prefixes," even without changing the output head. This directly inspired two things: first, later instruction fine-tuning (see Fine-tuning: SFT and PEFT) where "task descriptions go in the input"; and second, the generality idea of "one model for everything."
A mnemonic for the three architectures
Encoders "understand," decoders "write," encoder-decoders "transform." Classification, retrieval, representations → BERT family; generation, dialogue → GPT family; translation, summarization, multitask → T5 family. Understanding architecture choices lets you predict a model family's capability boundaries.
VI. BERT's Legacy and Limitations
1. Legacy: Four Active Directions in the LLM Era
- Text representation (Embedding) models: Bidirectional models like BERT remain the primary foundation for high-quality sentence/document embeddings — retrieval embedding models like E5, bge, GTE, and text-embedding series are mostly distilled/fine-tuned from encoder architectures.
- Retrieval encoders: RAG's dense retrieval relies on bi-encoder/single-tower encoders to map queries and documents into the same vector space (e.g., DPR), which is the foundation of the "index" stage in RAG: Retrieval-Augmented Generation.
- Cost-effective choices for understanding tasks: On tasks like classification, NER, spam filtering, and ranking, a few-hundred-MB encoder fine-tuned model has far lower cost and latency than calling a generative large model — and remains the mainstream industrial solution.
- Distillation and efficiency: Small encoder models (DistilBERT, etc.) and the "teacher-student" distillation paradigm have been repeatedly reused in the LLM era's lightweighting efforts.
A natural question: since decoders kept growing to trillions of parameters, why did encoders stall at hundreds of millions to tens of billions? Three reasons: first, understanding task capability (classification, extraction) saturates early, and the marginal return on scale is low; second, retrieval representations value "alignment" and "discriminability" over "capacity," and bi-encoder models that are too large are harder to train; third, deployment cost — a billion-parameter encoder can already serve in real-time on CPU, and larger encoders offer no cost advantage. Encoders being "small but fine" is an engineering choice, not a capability ceiling.
2. Limitations: Why They're Bad at Generation
- No autoregressive decoding capability: BERT doesn't produce "the next word"; generating requires an external decoder or converting to seq2seq, with worse results and a poorer architectural fit than a native decoder.
- Bidirectionality creates "exposure bias": During training the model sees the full sentence, but during generation it can only see what has been generated — the distributions are inconsistent.
- 512-token limit: Early encoders had very short context, weak long-document capability (later Longformer and others improved, but didn't change the landscape).
- Masked objective is inefficient: MLM only uses 15% of tokens for prediction — less information-efficient than autoregressive objectives (one reason the GPT path later overtook it).
3. How Encoders and Decoders Cooperate
In real systems, encoders and decoders don't compete — they collaborate. In a typical retrieval-augmented QA system (see RAG: Retrieval-Augmented Generation): the encoder compresses a library of millions of documents into a searchable vector index (low cost, low latency), while the decoder organizes a natural-language answer from the retrieved results (high capability, high cost). The encoder's recall quality directly determines the information boundary the decoder can access — if retrieval is wrong, even the strongest generation can't save it. Similar division of labor appears in evaluation pipelines: using encoders for vector similarity scoring, and decoders for LLM-as-a-judge semantic review — two complementary approaches. Figuring out "who finds, who speaks" gets half the system design right.
Encoders also play many "hidden roles": large-scale data deduplication (finding near-duplicate text via embeddings), classification routing (judging question type before assigning to different models), and safety filtering (sensitive content detection) — encoders are far faster than generative large models in all of these. This means even a team that fully uses closed-source large models for its frontend will almost certainly be running a fleet of encoder models behind the scenes. Encoders won't disappear — they'll just recede into "background infrastructure."
4. The Absence of Hallucination and Alignment Issues
Encoder models don't "speak," so they don't have "generation-side" problems like hallucination: causes and mitigation, alignment, or safety; but when they serve as a component of an LLM application (retrieval, classification, scoring), errors propagate upward, and they still need to be incorporated into the evaluation and benchmark system.
VII. Position in the LLM Era: From "Protagonist" to "Tool Component"
From 2018–2021, BERT was the de facto protagonist of NLP; after ChatGPT in 2022, generative large models took over the mainstream scenarios of dialogue, Agents, and creative writing, and encoders receded from "protagonist" to "tool component":
| Scenario | Answer in 2020 | Answer in 2025 |
|---|---|---|
| Sentiment classification / extraction | Fine-tuned BERT | Small encoder fine-tuning or generative model |
| QA / dialogue | Fine-tuned BERT / T5 | Generative large model |
| Semantic retrieval vectors | BERT embeddings | Encoder embeddings (still mainstream) |
| Zero-shot multitask | Unlikely | Generative large model + prompts |
One-sentence summary: BERT defined "pretraining + fine-tuning" and "deep bidirectional representations," and both remain deeply embedded in the underlying layers of LLM applications today (retrieval, embeddings, distillation, understanding subtasks) — it's just that the spotlight has shifted from BERT to the decoder family.
The research focus for encoders over the next few years has three directions: first, multilingual and long-document embeddings, enabling retrieval coverage across more languages and longer documents; second, integration with multimodal approaches, incorporating images, tables, and code into the same vector space (related ideas in Multimodal LLMs); and third, online learning and continuous updating, so representation models can keep up with knowledge base changes without frequent retraining. These directions point to one assessment: the encoder's role as an "efficient understanding and retrieval component" will persist long-term in the LLM ecosystem.
A direct recommendation for practitioners: treat encoders as engineering components rather than research hotspots. Pick a mature base (bge / E5 / GTE-type), fine-tune or distill on your data, pair it with an evaluation set, and you'll reliably capture the dividend of "low-latency, low-cost understanding"; invest the saved energy in retrieval strategy and generative model application layers, where the ROI is usually higher. Methods for evaluating encoders are in Evaluation and Benchmarks and Evaluations in Practice.
Don't use the BERT family for generation
A common mistake is "fine-tuning BERT to write articles or have dialogues." Encoders have no autoregressive generation mechanism — forcing output is both slow and poor quality. Use decoder models for generation tasks; the correct use of BERT is for representation, retrieval, classification, and distillation.
VIII. Modern Applications of Encoders: Retrieval, Classification, and Distillation in Practice
After understanding the principles and family, let's look at how encoders are used concretely today. This section provides four mainstream deployment paths and engineering parameters.
1. Retrieval Encoders: Bi-Encoders and Cross-Encoders
Two types of encoders serve different roles in retrieval:
| Type | Approach | Speed | Accuracy | Use Case |
|---|---|---|---|---|
| Bi-encoder | Query and doc encoded separately, compared for similarity | Fast (vector-indexable) | Medium | Recall: coarse screening of millions of candidates |
| Cross-encoder | Query and doc concatenated and scored jointly | Slow | High | Reranking: re-sorting top-k results |
The standard retrieval pipeline for RAG is "bi-encoder recall + cross-encoder reranking" (see RAG: Retrieval-Augmented Generation). Representative open-source embedding models include bge, E5, GTE, and text-embedding series; their general evaluation benchmark is MTEB (covering retrieval, classification, clustering, semantic similarity, and more; see Datasets and Benchmarks Archive).
# Bi-encoder recall flow (pseudocode)
query_vec = embed_query(question)
doc_vecs = vector_db.search(query_vec, top_k=100) # vector search for coarse screening
rerank = cross_encoder(question, doc) for doc in doc_vecs # fine ranking
final = sort(rerank)[:5] # take top 5 snippets2. Fine-tuning for Classification and Extraction
Classification (sentiment, intent, spam filtering) and extraction (NER, keyword) are classic fine-tuning scenarios for encoders:
- Classification: take the
[CLS]vector → connect a linear layer → softmax; - Extraction: output a label for each token position (BIO sequence labeling);
- Data volume: hundreds to tens of thousands of labeled examples suffice, far fewer than instruction data for generative models;
- Fine-tuning tips: low learning rate, small batch, early stopping, to prevent catastrophic forgetting (methods in Fine-tuning: SFT and Parameter-Efficient Fine-tuning).
3. Distillation and Compression
The encoder family is the most mature application area for knowledge distillation: training small models on soft labels from large models (BERT-large, or even GPT-type teachers) while retaining 95%+ capability and reducing inference latency by an order of magnitude. Combined with quantization (INT8/INT4), small encoder models can serve in real-time on CPU. Deployment tips in Deployment and Servicing.
4. Quick Reference for Deployment Selection
| Task | Recommended Approach | Size Reference |
|---|---|---|
| Semantic retrieval embeddings | bge / E5 / GTE (encoder distillation) | 0.1–1B parameters |
| QA reranking | Cross-encoder (BGE-Reranker, etc.) | Tens to hundreds of MB |
| Sentiment classification / NER | Fine-tuned BERT/RoBERTa small model | 100–400 MB |
| Large-scale clustering / deduplication | Embeddings + clustering algorithms | Any encoder |
5. When Not to Use Encoders
| Scenario | Why Not an Encoder | Correct Choice |
|---|---|---|
| Dialogue / creative writing / summarization | No generation capability | Decoder large model |
| Open-domain QA | Need to organize full answers | Decoder + RAG |
| Need zero-shot generalization | Encoder fine-tuning only works after training | Decoder + prompt |
| Complex multi-step reasoning | Encoders don't do reasoning | Decoder + CoT/Agent |
The contemporary positioning of encoders
Encoders are "high-speed understanding components," decoders are "general-purpose brains." For any job of "converting text into vectors or labels," encoders are fast and cheap; for any job of "speaking and doing things," hand it to a decoder. They collaborate extensively in RAG, distillation, and evaluation pipelines.
IX. Encoder Selection and Practical Q&A
1. Task-Driven Selection Table
| Task | Primary Choice | Alternative | Why |
|---|---|---|---|
| Sentiment / intent classification | Fine-tuned BERT/RoBERTa | Generative model + prompt | Classification is fast, stable, and cheap |
| NER / relation extraction | Fine-tuned encoder | Generative model + structured output | Sequence labeling is naturally suited to encoders |
| Semantic retrieval | bge / E5 / GTE | Closed-source embedding API | Encoder embeddings have high quality |
| Semantic similarity | Sentence-BERT-type | Cross-encoder | Bi-encoder is indexable |
| Text deduplication / clustering | Embeddings + clustering | — | Scales well |
| QA / dialogue / creative writing | Decoder large model | — | Requires generation capability |
2. Fine-tuning Data and Hyperparameters for Encoders
| Hyperparameter | Recommended Value | Notes |
|---|---|---|
| Learning rate | 2e-5 ~ 5e-5 | An order of magnitude lower than training from scratch |
| Batch size | 16–64 | Too small is unstable; too large causes OOM |
| Training epochs | 2–4 | Encoders overfit easily; use early stopping |
| Data volume | Hundreds to tens of thousands | Quality >> quantity |
| Adversarial validation | Yes | Prevents learning dataset noise |
Data preparation and annotation guidelines are covered in the post-training data section of Pretraining: Data and Objectives; evaluation comparison in Evaluation and Benchmarks.
3. Evaluation Differences: Encoders vs Generative Models
| Dimension | Encoder | Generative Model |
|---|---|---|
| Output form | Vectors / labels | Freeform text |
| Evaluation method | Accuracy, F1, Recall@k | Automatic metrics + human preference + LLM judge |
| Reproducibility | High (deterministic scoring) | Medium (sampling variance) |
| Regression testing | Cheap | Expensive (regenerate each round) |
4. Will Encoders Be Replaced by Large Models?
Not in the short term, for four reasons:
- Cost magnitude difference: encoder inference is one to two orders of magnitude cheaper than generative models;
- Latency-sensitive scenarios: search, recommendation, and risk control require millisecond-level responses;
- Interpretable and controllable: classification results are auditable, suitable for compliance scenarios;
- Distillation still needs teacher representations: hidden representations from generative models are often distilled into small encoders for retrieval.
Long-term, the two will merge more deeply: generative models provide the capability ceiling, while encoders handle high-frequency, low-latency tasks — RAG, distillation, and evaluation pipelines are all stages where they collaborate (see RAG: Retrieval-Augmented Generation).
5. FAQ Quick Answers
| Question | Quick Answer |
|---|---|
| Can encoders generate? | No; use a decoder for generation |
| Which embedding for Chinese retrieval? | bge / GTE / E5 Chinese versions |
| How much data for fine-tuning? | Hundreds of examples for classification, fewer for extraction |
| Will fine-tuning cause forgetting? | Yes; use low learning rate + mix in original corpus |
| How much VRAM for deployment? | A 400M model in FP16 needs ~0.8 GB; less after quantization |
| How do encoders and LLMs divide labor? | Encoders do understanding/retrieval; LLMs do generation/decision-making |
One more common misconception about "encoder vs decoder": many assume decoder large models have completely subsumed encoder capability. In practice, for scenarios requiring "fixed-structure output + high throughput" (ranking, scoring, clustering, large-scale classification), a decoder's generative output is both slow and hard to format — an encoder remains the more practical choice. This isn't a capability issue; it's a form-fit issue — the two model types each have strengths, and selection should depend on the task form rather than "which is stronger."
One sentence to remember encoder selection
Use "whether you need to generate text" as the dividing line: no generation → encoder (fast, cheap, stable); generation needed → decoder. Most systems need both, connected via RAG and distillation.
XI. Deep Dive into Encoder Evaluation Benchmarks
1. GLUE: The "Original Physical Exam" for Understanding Capability
GLUE consists of 9 subtasks covering sentence-level and single-sentence understanding:
| Task | Type | Assesses |
|---|---|---|
| CoLA | Grammatical acceptability | Linguistic capability |
| SST-2 | Sentiment binary classification | Sentiment understanding |
| MRPC / QQP | Sentence similarity/repetition | Semantic matching |
| STS-B | Semantic similarity scoring | Continuous semantics |
| MNLI / RTE | Entailment reasoning | Logical reasoning |
| QNLI | QA entailment | Reading and reasoning |
| WNLI | Coreference resolution | World knowledge |
BERT-large topped GLUE at 82.1; RoBERTa reached 88.5; subsequent models broke 90, "maxing out" GLUE, and researchers turned to harder tasks.
2. SuperGLUE: Difficulty Escalation
SuperGLUE was designed for models that had beaten GLUE: adding common-sense QA (COPA), multi-choice reading (MultiRC), and abstract reasoning (ReCoRD), emphasizing knowledge and reasoning more. T5, ELECTRA, and others topped it.
3. MTEB: The Modern Measure for Retrieval/Representation
As encoders' primary use shifted from "classification" to "retrieval and representation," MTEB (Massive Text Embedding Benchmark) became the new standard — covering 8 task categories (retrieval, classification, clustering, reranking, STS, summarization) and dozens of datasets, evaluating embedding quality uniformly.
| Category | Representative Datasets | Notes |
|---|---|---|
| Retrieval | BEIR series | Dense retrieval quality |
| Classification | AmazonPolarity, etc. | Linear classification capability |
| Clustering | ArXiv, etc. | Unsupervised representation |
| Reranking | Various corpora | Ranking quality |
4. Evaluation Contamination and Correct Usage
- Contamination: model training data may include evaluation sets, inflating scores;
- Correct approach: build a custom golden set + cross-validate on public benchmarks, and record model and data versions (methods in Evaluation and Benchmarks);
- Trend: encoder evaluation is shifting toward "retrieval-scenarioized" (enterprise knowledge bases, multilingual, long documents).
Also a reminder: historical scores on GLUE/SuperGLUE shouldn't be used directly for comparing today's models — first, tasks may have been covered by training data (contamination), and second, old leaderboards can't reflect long-context and multilingual capabilities. Modern selection should rely on new benchmarks like MTEB plus custom sets.
XII. Classic Resources and Tools
| Resource | Use |
|---|---|
| HF Transformers | Load/fine-tune BERT family and embedding models |
| HF sentence-transformers | Sentence embeddings, bi-encoder training and inference |
| MTEB leaderboard | Horizontal comparison of embedding models |
| GLUE / SuperGLUE | Understanding task benchmarks |
| Hugging Face Trainer / PEFT | Fine-tuning and LoRA |
| bge / E5 / GTE repos | Chinese and multilingual embedding bases |
Modern advice for encoder evaluation
For small classification tasks, check GLUE history + custom set; for retrieval tasks, check MTEB + test on your own data directly. Leaderboards are just the starting point — running your own data through is the real validation.
One more commonly cited ratio about distillation: the typical encoder distillation result is "about 40% of parameters retaining about 90% of capability" — DistilBERT, with 66M parameters (about 60% of BERT-base), retaining 97% of GLUE capability is one of the better cases. Distillation doesn't just save VRAM; it makes models easier to deploy on edge devices (phones, IoT), and when combined with quantization ideas from Deployment and Servicing, it's the primary path for "pushing down" large model capability to low-cost hardware.
XIII. Further Reading
- Language Modeling: The Next-Token Prediction Paradigm — comparing MLM and autoregressive objectives
- Transformer Architecture Explained — the attention differences between encoders and decoders
- The GPT Series: From GPT-1 to GPT-4o — the decoder family on a parallel but different path
- RAG: Retrieval-Augmented Generation — the role of encoder embeddings in retrieval
- Fine-tuning: SFT and Parameter-Efficient Fine-tuning — comparing fine-tuning methods for encoders and decoders
- Model Compendium — the encoder family's position in the model landscape
References
- Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2018) — BERT original paper (arXiv)
- Liu et al. RoBERTa: A Robustly Optimized BERT Pretraining Approach (2019) — RoBERTa paper (arXiv)
- Lan et al. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations (2019) — ALBERT paper (arXiv)
- Sanh et al. DistilBERT, a distilled version of BERT (2019) — DistilBERT paper (arXiv)
- Raffel et al. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5, 2019) — T5 paper (arXiv)
- Yang et al. XLNet: Generalized Autoregressive Pretraining for Language Understanding (2019) — XLNet paper (arXiv)
- Clark et al. ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators (2020) — ELECTRA paper (arXiv)
- Reimers & Gurevych. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks (2019) — Representative work on sentence embeddings and retrieval applications (arXiv)