Skip to content

BERT and the Encoder Family

At a glance BERT redefined the pretraining paradigm of NLP with "masked language modeling + bidirectional encoding." This article breaks down BERT's MLM/NSP mechanisms, its divergence from GPT, the RoBERTa/ALBERT/DistilBERT family evolution, T5's text-to-text framework, and BERT's legacy in the LLM era.

This page contains time-sensitive content, current as of 2025-06; job descriptions, rankings, product features, and other information may have changed. Please verify with original sources before citing.

BERT and the Encoder Family ​

BERT (Bidirectional Encoder Representations from Transformers) is a bidirectional representation model based on the Transformer encoder (encoder-only) with "masked language modeling (MLM)" as its pretraining task, proposed by Google in 2018. It ignited the "pretraining + fine-tuning" paradigm alongside GPT, but diverged from GPT in architecture: GPT went "decoder + generation," while BERT went "encoder + understanding." This page breaks down BERT's principles, family evolution, and its position in the LLM era; for underlying architecture mechanisms, see Transformer Architecture Explained; for a complete comparison with the GPT path, see The GPT Series: From GPT-1 to GPT-4o.

I. What Is BERT: A One-Sentence Definition ​

BERT = bidirectional encoder + masked language modeling (MLM) pretraining + downstream fine-tuning. It pretrains by "randomly masking words in sentences and having the model guess them from the bidirectional context on both sides," yielding deep bidirectional text representations; then fine-tunes with a lightweight output head on each downstream task (classification, extraction, QA, similarity).

Pretraining objective (MLM):
Input:  China's [MASK] is the capital, with a population exceeding [MASK]0 million.
Predict: [MASK]₁ → "capital"    [MASK]₂ → "1" (based on bidirectional context)

Downstream fine-tuning (e.g., sentiment classification):
[CLS] This movie is great [SEP] → output layer → positive

The BERT paper, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, was posted to arXiv in October 2018 (and won NAACL best long paper in 2019). It set new SOTA on 11 NLP benchmarks, appeared almost simultaneously with GPT-1, and together they established the "pretraining + fine-tuning" era.

II. Principle Breakdown: Three Key BERT Designs ​

1. Bidirectional: The Core Difference from GPT ​

GPT uses causal masking — each position can only "see the left" (unidirectional autoregressive); BERT has no causal masking, and each position sees the full context (bidirectional). Intuitively, understanding tasks (judging sentiment, extracting entities) heavily depend on the following context — "this movie isn't that good" requires seeing the "not" before knowing it's negative. Bidirectionality is BERT's fundamental advantage over same-scale GPT on understanding tasks.

2. MLM: Masked Language Modeling ​

The mask strategy's details (from the original paper): randomly select 15% of tokens for prediction, of which 80% are replaced with [MASK], 10% are replaced with random words, and 10% are kept unchanged. The motivation: during both pretraining and fine-tuning, the model never sees [MASK], forcing it to learn contextual semantics rather than "seeing [MASK] and guessing a word."

The 80/10/10 ratio is deliberately nuanced: if 100% were replaced with [MASK], the model would only learn "seeing [MASK] means guess a word," but during fine-tuning there are no [MASK] tokens — the pretraining and fine-tuning input distributions would be inconsistent. If 100% were kept unchanged, the model would have almost no prediction pressure. The balanced 80/10/10 lets the model both depend on contextual semantics and remain robust to "randomly replaced words." This design was later extended or revised by dynamic masking and word-level masking, but the core idea — "align the pretraining objective with the downstream distribution" — remains the central consideration in pretraining objective design (see Pretraining: Data and Objectives).

3. NSP: Next Sentence Prediction ​

To help the model understand "inter-sentence relationships" (important for QA and entailment tasks), BERT simultaneously trains a binary classification task: judging whether sentence B is the next sentence of sentence A (50% real / 50% random). Subsequent research found NSP contributed little (removing it in RoBERTa actually made it stronger), but it reflected the era's exploration of "inter-sentence structure."

4. Input Construction and Specifications ​

ItemSpecification
Input format[CLS] sentenceA [SEP] sentenceB [SEP] + segment embedding + position embedding
Output head[CLS] position vector for classification; remaining position vectors for token-level tasks
BERT-base110M parameters (12 layers / 768-dim / 12 heads)
BERT-large340M parameters (24 layers / 1024-dim / 16 heads)
Pretraining dataBookCorpus + English Wikipedia (~3.3 billion words)
Maximum input length512 tokens (position embedding limit)

5. Results ​

BERT-large set new SOTA on 11 benchmarks at release, including GLUE (79.6 → 82.1), SQuAD 1.1 (F1 93.2, surpassing human level), and SQuAD 2.0. Its publicly available weights + open-source implementation made it realistic for "anyone to fine-tune a world-class NLP model in a few hours," directly catalyzing the explosion of the Hugging Face ecosystem.

III. BERT vs GPT: A Comparison of Two Paths ​

This is the most important architectural comparison of 2018–2023, worth revisiting:

The value of bidirectional attention varies significantly across tasks: for any task where "the answer depends on the full context" (entailment judgment, coreference resolution, QA), bidirectional has a clear advantage; for tasks where "order matters" (generation, translation, temporal reasoning), unidirectional/autoregressive is more natural. This is why encoder models remain strong baselines for natural language entailment and extraction tasks, while generation tasks have been entirely taken over by decoders. When choosing an architecture, don't ask "which is better" — ask "does my task rely more on global context or sequential generation?"

DimensionBERTGPT
ArchitectureEncoder (encoder-only)Decoder (decoder-only)
Attention directionBidirectional (no causal mask)Unidirectional (causal mask, sees only the left)
Pretraining objectiveMasked language modeling (MLM)Autoregressive next-token prediction
Suitable tasksUnderstanding: classification, extraction, ranking, retrievalGeneration: continuation, translation, dialogue, code
Downstream adaptationAdd output head and fine-tunePrompts / in-context learning / fine-tuning
Scale trajectoryStalled at the hundred-million level (at most ~tens of billions)Continuously grew to hundreds of billions and trillions
Era outcome"Tool component" in the LLM era"Protagonist" of the LLM era

Why did GPT ultimately win?

Both objectives are fundamentally complementary: MLM makes representations more "understanding," while autoregression makes models more "writing-capable." But killer apps like dialogue, programming, and Agents all require generation; and once generation capability is paired with scale and alignment (see Alignment: RLHF and DPO), it naturally covers most understanding tasks. BERT's understanding advantage became insignificant as the absolute capability gap widened.

IV. The Encoder Family: BERT Improvements and Compression ​

A wave of improvements emerged within a year of BERT's release, targeting four directions: more data, longer training, fewer parameters, faster inference.

ModelTimeParametersKey ChangesResults
BERT-large2018.10340MBaselineGLUE 82.1
RoBERTa2019.7355MRemoved NSP, dynamic masking, 10× data (160 GB), 4× longer trainingGLUE 88.5
XLNet2019.6340MPermutation language modeling (autoregressive objective + bidirectional context)Multiple SOTAs at the time
ALBERT2019.9~12M (after sharing)Cross-layer parameter sharing + factorized embeddings + SOP replacing NSPApproached RoBERTa with far fewer parameters
DistilBERT2019.1066MKnowledge distillation (teacher: BERT-base)Retained ~97% capability, ~60% faster inference
ELECTRA2020.3330MReplaced token detection (discriminative pretraining)Matched RoBERTa with ¼ compute

1. RoBERTa: A Victory for Data and Training Details ​

RoBERTa's findings are highly illuminating: BERT had significant room left to "train harder" — by increasing data from 16 GB to 160 GB, using dynamic masking, removing NSP, and training for 4× as many steps, GLUE went from 82.1 to 88.5. It proved that in the pretraining era, "data volume + training duration" matters as much as architectural innovation (this is isomorphic with the conclusion of Scaling Laws).

2. ALBERT and DistilBERT: Saving Parameters and Speeding Up ​

  • ALBERT: Three moves to reduce parameters — cross-layer parameter sharing, embedding matrix factorization, and replacing NSP with sentence order prediction (SOP). Parameters dropped from the billions to the tens of millions, with only a slight capability loss.
  • DistilBERT: A textbook example of knowledge distillation — training a 66M-parameter student using a teacher BERT's soft labels, retaining 97% of GLUE capability while being ~60% faster at inference. "Big models teaching small models" later became the mainstream approach for Building a Large Model from Scratch and distilling small models.

3. XLNet and ELECTRA: Two "Philosophical Variants" ​

  • XLNet: Wanted both "autoregressive + bidirectional" — used permutations to let an autoregressive model see the full context during training, but didn't achieve a fundamental win.
  • ELECTRA: Changed pretraining from "generative" to "discriminative" — having a generator mask tokens and a discriminator judge whether each position was replaced, achieving high computational efficiency and performance close to RoBERTa.

V. T5: Unifying Every Task as "Text-to-Text" ​

1. Core Idea ​

T5 (Text-to-Text Transfer Transformer, October 2019, Google) chose a third path: encoder-decoder, and unified every task (translation, summarization, classification, QA) into "input a piece of text, output a piece of text":

Input: "Translate to English: I love China"  →  Output: "I love China"
Input: "Sentiment: This movie is amazing"  →  Output: "positive"

2. Specifications and Contributions ​

ItemSpecification
Largest versionT5-11B (encoder-decoder, ~11 billion parameters)
Training dataC4 (Colossal Clean Crawled Corpus, ~750 GB, cleaned Common Crawl)
Key innovationUnified text-to-text framework, task prefix, C4 data cleaning methodology
ResultsMultiple SOTAs on GLUE / SuperGLUE / SQuAD at release

T5's other methodological legacy is the "task prefix + unified input/output" design, which made "using one model to handle hundreds of tasks" possible. This idea is in the same lineage as later instruction fine-tuning (see Fine-tuning: SFT and Parameter-Efficient Fine-tuning) — in fact, many instruction fine-tuning datasets treat task descriptions as T5-style prefixes. Understanding T5 helps explain why today's models can switch skills based on a single instruction: the task itself is part of the input, and the model's "generality" comes from generalizing over "task descriptions."

T5's contribution is methodological: it systematically validated the idea of "unified task representation" — the same pretrained model switches tasks through different "text prefixes," even without changing the output head. This directly inspired two things: first, later instruction fine-tuning (see Fine-tuning: SFT and PEFT) where "task descriptions go in the input"; and second, the generality idea of "one model for everything."

A mnemonic for the three architectures

Encoders "understand," decoders "write," encoder-decoders "transform." Classification, retrieval, representations → BERT family; generation, dialogue → GPT family; translation, summarization, multitask → T5 family. Understanding architecture choices lets you predict a model family's capability boundaries.

VI. BERT's Legacy and Limitations ​

1. Legacy: Four Active Directions in the LLM Era ​

  1. Text representation (Embedding) models: Bidirectional models like BERT remain the primary foundation for high-quality sentence/document embeddings — retrieval embedding models like E5, bge, GTE, and text-embedding series are mostly distilled/fine-tuned from encoder architectures.
  2. Retrieval encoders: RAG's dense retrieval relies on bi-encoder/single-tower encoders to map queries and documents into the same vector space (e.g., DPR), which is the foundation of the "index" stage in RAG: Retrieval-Augmented Generation.
  3. Cost-effective choices for understanding tasks: On tasks like classification, NER, spam filtering, and ranking, a few-hundred-MB encoder fine-tuned model has far lower cost and latency than calling a generative large model — and remains the mainstream industrial solution.
  4. Distillation and efficiency: Small encoder models (DistilBERT, etc.) and the "teacher-student" distillation paradigm have been repeatedly reused in the LLM era's lightweighting efforts.

A natural question: since decoders kept growing to trillions of parameters, why did encoders stall at hundreds of millions to tens of billions? Three reasons: first, understanding task capability (classification, extraction) saturates early, and the marginal return on scale is low; second, retrieval representations value "alignment" and "discriminability" over "capacity," and bi-encoder models that are too large are harder to train; third, deployment cost — a billion-parameter encoder can already serve in real-time on CPU, and larger encoders offer no cost advantage. Encoders being "small but fine" is an engineering choice, not a capability ceiling.

2. Limitations: Why They're Bad at Generation ​

  • No autoregressive decoding capability: BERT doesn't produce "the next word"; generating requires an external decoder or converting to seq2seq, with worse results and a poorer architectural fit than a native decoder.
  • Bidirectionality creates "exposure bias": During training the model sees the full sentence, but during generation it can only see what has been generated — the distributions are inconsistent.
  • 512-token limit: Early encoders had very short context, weak long-document capability (later Longformer and others improved, but didn't change the landscape).
  • Masked objective is inefficient: MLM only uses 15% of tokens for prediction — less information-efficient than autoregressive objectives (one reason the GPT path later overtook it).

3. How Encoders and Decoders Cooperate ​

In real systems, encoders and decoders don't compete — they collaborate. In a typical retrieval-augmented QA system (see RAG: Retrieval-Augmented Generation): the encoder compresses a library of millions of documents into a searchable vector index (low cost, low latency), while the decoder organizes a natural-language answer from the retrieved results (high capability, high cost). The encoder's recall quality directly determines the information boundary the decoder can access — if retrieval is wrong, even the strongest generation can't save it. Similar division of labor appears in evaluation pipelines: using encoders for vector similarity scoring, and decoders for LLM-as-a-judge semantic review — two complementary approaches. Figuring out "who finds, who speaks" gets half the system design right.

Encoders also play many "hidden roles": large-scale data deduplication (finding near-duplicate text via embeddings), classification routing (judging question type before assigning to different models), and safety filtering (sensitive content detection) — encoders are far faster than generative large models in all of these. This means even a team that fully uses closed-source large models for its frontend will almost certainly be running a fleet of encoder models behind the scenes. Encoders won't disappear — they'll just recede into "background infrastructure."

4. The Absence of Hallucination and Alignment Issues ​

Encoder models don't "speak," so they don't have "generation-side" problems like hallucination: causes and mitigation, alignment, or safety; but when they serve as a component of an LLM application (retrieval, classification, scoring), errors propagate upward, and they still need to be incorporated into the evaluation and benchmark system.

VII. Position in the LLM Era: From "Protagonist" to "Tool Component" ​

From 2018–2021, BERT was the de facto protagonist of NLP; after ChatGPT in 2022, generative large models took over the mainstream scenarios of dialogue, Agents, and creative writing, and encoders receded from "protagonist" to "tool component":

ScenarioAnswer in 2020Answer in 2025
Sentiment classification / extractionFine-tuned BERTSmall encoder fine-tuning or generative model
QA / dialogueFine-tuned BERT / T5Generative large model
Semantic retrieval vectorsBERT embeddingsEncoder embeddings (still mainstream)
Zero-shot multitaskUnlikelyGenerative large model + prompts

One-sentence summary: BERT defined "pretraining + fine-tuning" and "deep bidirectional representations," and both remain deeply embedded in the underlying layers of LLM applications today (retrieval, embeddings, distillation, understanding subtasks) — it's just that the spotlight has shifted from BERT to the decoder family.

The research focus for encoders over the next few years has three directions: first, multilingual and long-document embeddings, enabling retrieval coverage across more languages and longer documents; second, integration with multimodal approaches, incorporating images, tables, and code into the same vector space (related ideas in Multimodal LLMs); and third, online learning and continuous updating, so representation models can keep up with knowledge base changes without frequent retraining. These directions point to one assessment: the encoder's role as an "efficient understanding and retrieval component" will persist long-term in the LLM ecosystem.

A direct recommendation for practitioners: treat encoders as engineering components rather than research hotspots. Pick a mature base (bge / E5 / GTE-type), fine-tune or distill on your data, pair it with an evaluation set, and you'll reliably capture the dividend of "low-latency, low-cost understanding"; invest the saved energy in retrieval strategy and generative model application layers, where the ROI is usually higher. Methods for evaluating encoders are in Evaluation and Benchmarks and Evaluations in Practice.

Don't use the BERT family for generation

A common mistake is "fine-tuning BERT to write articles or have dialogues." Encoders have no autoregressive generation mechanism — forcing output is both slow and poor quality. Use decoder models for generation tasks; the correct use of BERT is for representation, retrieval, classification, and distillation.

VIII. Modern Applications of Encoders: Retrieval, Classification, and Distillation in Practice ​

After understanding the principles and family, let's look at how encoders are used concretely today. This section provides four mainstream deployment paths and engineering parameters.

1. Retrieval Encoders: Bi-Encoders and Cross-Encoders ​

Two types of encoders serve different roles in retrieval:

TypeApproachSpeedAccuracyUse Case
Bi-encoderQuery and doc encoded separately, compared for similarityFast (vector-indexable)MediumRecall: coarse screening of millions of candidates
Cross-encoderQuery and doc concatenated and scored jointlySlowHighReranking: re-sorting top-k results

The standard retrieval pipeline for RAG is "bi-encoder recall + cross-encoder reranking" (see RAG: Retrieval-Augmented Generation). Representative open-source embedding models include bge, E5, GTE, and text-embedding series; their general evaluation benchmark is MTEB (covering retrieval, classification, clustering, semantic similarity, and more; see Datasets and Benchmarks Archive).

# Bi-encoder recall flow (pseudocode)
query_vec = embed_query(question)
doc_vecs  = vector_db.search(query_vec, top_k=100)   # vector search for coarse screening
rerank    = cross_encoder(question, doc) for doc in doc_vecs  # fine ranking
final     = sort(rerank)[:5]                          # take top 5 snippets

2. Fine-tuning for Classification and Extraction ​

Classification (sentiment, intent, spam filtering) and extraction (NER, keyword) are classic fine-tuning scenarios for encoders:

  • Classification: take the [CLS] vector → connect a linear layer → softmax;
  • Extraction: output a label for each token position (BIO sequence labeling);
  • Data volume: hundreds to tens of thousands of labeled examples suffice, far fewer than instruction data for generative models;
  • Fine-tuning tips: low learning rate, small batch, early stopping, to prevent catastrophic forgetting (methods in Fine-tuning: SFT and Parameter-Efficient Fine-tuning).

3. Distillation and Compression ​

The encoder family is the most mature application area for knowledge distillation: training small models on soft labels from large models (BERT-large, or even GPT-type teachers) while retaining 95%+ capability and reducing inference latency by an order of magnitude. Combined with quantization (INT8/INT4), small encoder models can serve in real-time on CPU. Deployment tips in Deployment and Servicing.

4. Quick Reference for Deployment Selection ​

TaskRecommended ApproachSize Reference
Semantic retrieval embeddingsbge / E5 / GTE (encoder distillation)0.1–1B parameters
QA rerankingCross-encoder (BGE-Reranker, etc.)Tens to hundreds of MB
Sentiment classification / NERFine-tuned BERT/RoBERTa small model100–400 MB
Large-scale clustering / deduplicationEmbeddings + clustering algorithmsAny encoder

5. When Not to Use Encoders ​

ScenarioWhy Not an EncoderCorrect Choice
Dialogue / creative writing / summarizationNo generation capabilityDecoder large model
Open-domain QANeed to organize full answersDecoder + RAG
Need zero-shot generalizationEncoder fine-tuning only works after trainingDecoder + prompt
Complex multi-step reasoningEncoders don't do reasoningDecoder + CoT/Agent

The contemporary positioning of encoders

Encoders are "high-speed understanding components," decoders are "general-purpose brains." For any job of "converting text into vectors or labels," encoders are fast and cheap; for any job of "speaking and doing things," hand it to a decoder. They collaborate extensively in RAG, distillation, and evaluation pipelines.

IX. Encoder Selection and Practical Q&A ​

1. Task-Driven Selection Table ​

TaskPrimary ChoiceAlternativeWhy
Sentiment / intent classificationFine-tuned BERT/RoBERTaGenerative model + promptClassification is fast, stable, and cheap
NER / relation extractionFine-tuned encoderGenerative model + structured outputSequence labeling is naturally suited to encoders
Semantic retrievalbge / E5 / GTEClosed-source embedding APIEncoder embeddings have high quality
Semantic similaritySentence-BERT-typeCross-encoderBi-encoder is indexable
Text deduplication / clusteringEmbeddings + clustering—Scales well
QA / dialogue / creative writingDecoder large model—Requires generation capability

2. Fine-tuning Data and Hyperparameters for Encoders ​

HyperparameterRecommended ValueNotes
Learning rate2e-5 ~ 5e-5An order of magnitude lower than training from scratch
Batch size16–64Too small is unstable; too large causes OOM
Training epochs2–4Encoders overfit easily; use early stopping
Data volumeHundreds to tens of thousandsQuality >> quantity
Adversarial validationYesPrevents learning dataset noise

Data preparation and annotation guidelines are covered in the post-training data section of Pretraining: Data and Objectives; evaluation comparison in Evaluation and Benchmarks.

3. Evaluation Differences: Encoders vs Generative Models ​

DimensionEncoderGenerative Model
Output formVectors / labelsFreeform text
Evaluation methodAccuracy, F1, Recall@kAutomatic metrics + human preference + LLM judge
ReproducibilityHigh (deterministic scoring)Medium (sampling variance)
Regression testingCheapExpensive (regenerate each round)

4. Will Encoders Be Replaced by Large Models? ​

Not in the short term, for four reasons:

  1. Cost magnitude difference: encoder inference is one to two orders of magnitude cheaper than generative models;
  2. Latency-sensitive scenarios: search, recommendation, and risk control require millisecond-level responses;
  3. Interpretable and controllable: classification results are auditable, suitable for compliance scenarios;
  4. Distillation still needs teacher representations: hidden representations from generative models are often distilled into small encoders for retrieval.

Long-term, the two will merge more deeply: generative models provide the capability ceiling, while encoders handle high-frequency, low-latency tasks — RAG, distillation, and evaluation pipelines are all stages where they collaborate (see RAG: Retrieval-Augmented Generation).

5. FAQ Quick Answers ​

QuestionQuick Answer
Can encoders generate?No; use a decoder for generation
Which embedding for Chinese retrieval?bge / GTE / E5 Chinese versions
How much data for fine-tuning?Hundreds of examples for classification, fewer for extraction
Will fine-tuning cause forgetting?Yes; use low learning rate + mix in original corpus
How much VRAM for deployment?A 400M model in FP16 needs ~0.8 GB; less after quantization
How do encoders and LLMs divide labor?Encoders do understanding/retrieval; LLMs do generation/decision-making

One more common misconception about "encoder vs decoder": many assume decoder large models have completely subsumed encoder capability. In practice, for scenarios requiring "fixed-structure output + high throughput" (ranking, scoring, clustering, large-scale classification), a decoder's generative output is both slow and hard to format — an encoder remains the more practical choice. This isn't a capability issue; it's a form-fit issue — the two model types each have strengths, and selection should depend on the task form rather than "which is stronger."

One sentence to remember encoder selection

Use "whether you need to generate text" as the dividing line: no generation → encoder (fast, cheap, stable); generation needed → decoder. Most systems need both, connected via RAG and distillation.

XI. Deep Dive into Encoder Evaluation Benchmarks ​

1. GLUE: The "Original Physical Exam" for Understanding Capability ​

GLUE consists of 9 subtasks covering sentence-level and single-sentence understanding:

TaskTypeAssesses
CoLAGrammatical acceptabilityLinguistic capability
SST-2Sentiment binary classificationSentiment understanding
MRPC / QQPSentence similarity/repetitionSemantic matching
STS-BSemantic similarity scoringContinuous semantics
MNLI / RTEEntailment reasoningLogical reasoning
QNLIQA entailmentReading and reasoning
WNLICoreference resolutionWorld knowledge

BERT-large topped GLUE at 82.1; RoBERTa reached 88.5; subsequent models broke 90, "maxing out" GLUE, and researchers turned to harder tasks.

2. SuperGLUE: Difficulty Escalation ​

SuperGLUE was designed for models that had beaten GLUE: adding common-sense QA (COPA), multi-choice reading (MultiRC), and abstract reasoning (ReCoRD), emphasizing knowledge and reasoning more. T5, ELECTRA, and others topped it.

3. MTEB: The Modern Measure for Retrieval/Representation ​

As encoders' primary use shifted from "classification" to "retrieval and representation," MTEB (Massive Text Embedding Benchmark) became the new standard — covering 8 task categories (retrieval, classification, clustering, reranking, STS, summarization) and dozens of datasets, evaluating embedding quality uniformly.

CategoryRepresentative DatasetsNotes
RetrievalBEIR seriesDense retrieval quality
ClassificationAmazonPolarity, etc.Linear classification capability
ClusteringArXiv, etc.Unsupervised representation
RerankingVarious corporaRanking quality

4. Evaluation Contamination and Correct Usage ​

  • Contamination: model training data may include evaluation sets, inflating scores;
  • Correct approach: build a custom golden set + cross-validate on public benchmarks, and record model and data versions (methods in Evaluation and Benchmarks);
  • Trend: encoder evaluation is shifting toward "retrieval-scenarioized" (enterprise knowledge bases, multilingual, long documents).

Also a reminder: historical scores on GLUE/SuperGLUE shouldn't be used directly for comparing today's models — first, tasks may have been covered by training data (contamination), and second, old leaderboards can't reflect long-context and multilingual capabilities. Modern selection should rely on new benchmarks like MTEB plus custom sets.

XII. Classic Resources and Tools ​

ResourceUse
HF TransformersLoad/fine-tune BERT family and embedding models
HF sentence-transformersSentence embeddings, bi-encoder training and inference
MTEB leaderboardHorizontal comparison of embedding models
GLUE / SuperGLUEUnderstanding task benchmarks
Hugging Face Trainer / PEFTFine-tuning and LoRA
bge / E5 / GTE reposChinese and multilingual embedding bases

Modern advice for encoder evaluation

For small classification tasks, check GLUE history + custom set; for retrieval tasks, check MTEB + test on your own data directly. Leaderboards are just the starting point — running your own data through is the real validation.

One more commonly cited ratio about distillation: the typical encoder distillation result is "about 40% of parameters retaining about 90% of capability" — DistilBERT, with 66M parameters (about 60% of BERT-base), retaining 97% of GLUE capability is one of the better cases. Distillation doesn't just save VRAM; it makes models easier to deploy on edge devices (phones, IoT), and when combined with quantization ideas from Deployment and Servicing, it's the primary path for "pushing down" large model capability to low-cost hardware.

XIII. Further Reading ​

References ​