Skip to content

Loss Functions and Output Layers

Quick overview The loss function is the source of gradient signals and the direct definition of "what the model optimizes for." This article explains why output layers and losses must match, covers classification (cross-entropy, BCE, label smoothing, focal), regression (MSE, MAE, Huber), ranking and contrastive losses, and provides guidance on multi-task balancing and debugging.

Loss Functions and Output Layers ​

One-line definition: A loss function quantifies "how far the model's prediction is from the ground truth" and is the source of all gradient signals. Backpropagation passes ∂L/∂W layer by layer back to the parameters (see Backpropagation and Automatic Differentiation); what the loss function looks like determines the quality and direction of these signals — "the loss defines the objective, optimization is responsible for reaching it."

I. Why the Output Layer Must Match the Loss Function ​

This is the most important principle in this entire article: the design of the output layer (whether to use activation, which activation to use) and the loss function are a matched pair — they cannot be selected independently.

  • Multi-class classification: output layer produces logits (unnormalized scores), paired with Softmax + cross-entropy.
  • Binary classification: output layer has one logit, paired with Sigmoid + binary cross-entropy.
  • Regression: output layer typically has no activation (linear output), paired with MSE/MAE and similar losses.

The mathematical reason is that cross-entropy requires input in "probability distribution" form, and Softmax/Sigmoid normalize logits into probabilities; the combination also matters for numerical stability. In PyTorch, nn.CrossEntropyLoss already applies Softmax to logits internally before computing cross-entropy, so do not manually add nn.Softmax — that causes double normalization and loss of numerical precision (see the note on "the output layer returns logits" in Neural Networks Fundamentals).

Typical Symptoms of Mismatched Output Layer and Loss

Manually adding Softmax to the output layer and then passing it to CrossEntropyLoss results in loss that is too high at the beginning of training, slow convergence, and degraded accuracy — because the gradient is distorted after softmax(logits). For troubleshooting thinking, see Debugging and Diagnostics.

II. Classification Losses: CrossEntropy, BCE, Label Smoothing, Focal ​

Cross-Entropy (CrossEntropy) ​

The standard loss for multi-class classification. For a single sample where the true class is y (from a one-hot perspective) and the model's predicted distribution is p:

L = −log p_y

It is equivalent to "negative log-likelihood": it encourages the predicted probability of the correct class to approach 1. Intuition: correct-class probability of 0.9 → loss ≈ 0.105; probability of 0.1 → loss ≈ 2.3. The penalty increases sharply as the model gets it more wrong.

Binary Cross-Entropy (BCE) ​

For binary classification (or multi-label where "each label is independently yes/no"):

L = −[y·log σ(z) + (1−y)·log(1−σ(z))]

where σ is Sigmoid. For multi-label classification, sum across each label dimension, with the output layer having "one Sigmoid per logit."

Label Smoothing ​

Hard labels 0/1 under cross-entropy encourage logits to grow without bound, leading to overconfidence and overfitting. Label smoothing changes the target to y' = (1−ε)·y + ε/K (K classes, ε≈0.1), preventing the model from pushing probabilities all the way to 1. This is a standard technique in ImageNet classification, machine translation, and even GPT-series training, working in concert with other methods in Overfitting and Regularization.

Focal Loss ​

When classes are extremely imbalanced, regular cross-entropy is dominated by "easy-to-classify majority samples." Focal Loss multiplies cross-entropy by a modulating factor (1−p_t)^γ:

L = −(1−p_t)^γ · log p_t

The closer p_t is to 1 (the easier the classification), the smaller the weight — directing the model's attention to "hard, minority-class" samples. It was proposed by Lin et al. in 2017, initially for dense object detection (RetinaNet). For a more systematic treatment of class imbalance, see Deep Learning Evaluation and Experiments.

III. Regression Losses: MSE, MAE, Huber ​

LossFormulaCharacteristicsWhen to Use
MSE(ŷ−y)²Differentiable everywhere, penalizes large errors quadratically; extremely sensitive to outliersErrors follow a Gaussian distribution, no outliers
MAE|ŷ−y|Linear penalty for large errors, robust to outliers; but non-differentiable at zeroContains outliers, robustness is priority
HuberMSE for small errors, linear for largeCombines the advantages of both, with a hyperparameter δ controlling the inflection pointDefault engineering compromise

Why is MSE sensitive to outliers? Because the gradient is 2(ŷ−y), a single sample deviating 10× produces a gradient 10× larger than a normal sample, dominating the entire update step. MAE's gradient is bounded (±1). Choosing a regression loss is essentially answering "how much should outliers influence the model."

Intuition

MSE assumes errors follow a Gaussian distribution; MAE implicitly assumes a Laplacian distribution — they correspond to "mean regression" and "median regression" respectively. If your data has heavy tails, MAE or Huber is more honest.

IV. Ranking Losses: Pairwise and Listwise ​

In ranking problems (recommendation, retrieval), we often care not about exact scores but about the relative order. Typical approaches:

  • Pointwise: Treat ranking as regression/classification, score each item independently. Simple, but doesn't directly optimize for ranking.
  • Pairwise: Construct a loss for pairs where "the positive example should rank above the negative," such as the pairwise constraints in RankNet and LambdaRank.
  • Listwise: Directly optimize the ranking quality of an entire list, such as ListNet, which directly optimizes ranking metrics like NDCG (using differentiable surrogates).

In recommendation systems, BPR (Bayesian Personalized Ranking) is the most commonly used pairwise loss: for a given user, it enforces that the score of "interacted items" exceeds that of "non-interacted items." See Deep Learning Recommender Systems.

V. Contrastive Losses: Triplet, InfoNCE, and Metric Learning ​

The goal of contrastive losses is not to fit labels, but to pull same-class samples together and push different-class samples apart in representation space — this directly serves Representation Learning and Pretraining.

Triplet Loss (FaceNet, 2015): a triplet of (anchor a, positive p, negative n):

L = max( d(a,p) − d(a,n) + margin, 0 )

Intuition: the "distance to the positive" should be at least margin smaller than the "distance to the negative."

InfoNCE (2018, Oord et al.): frames contrastive learning as a classification task of "picking the positive sample out of N candidates":

L = −log [ exp(sim(a,p)/τ) / Σⱼ exp(sim(a,qⱼ)/τ) ]

τ is temperature, controlling the sharpness of the distribution. Self-supervised methods like SimCLR and MoCo all use it; it is the standard loss for modern contrastive learning. It turns "similarity" into "multi-class probability," which is inherently the same family as Softmax cross-entropy.

VI. Multi-Task Loss Balancing: A Brief Introduction to GradNorm ​

Multi-task models minimize a weighted sum L = Σₖ wₖ·Lₖ. The problem: gradients from different tasks can vary widely in magnitude, and if wₖ isn't tuned well, the tasks with smaller losses get "drowned out." Common approaches:

  • Manually tune wₖ or normalize by magnitude.
  • Uncertainty weighting: learn the variance of each task as a parameter (Kendall et al., 2018).
  • GradNorm (Chen et al., 2018): dynamically adjusts wₖ during training so that the gradient norms of all tasks converge back to a common "target norm," preventing any single task from dominating.

There is no silver bullet for multi-task balancing, but an engineering rule of thumb: first get the magnitudes of all task losses close, then worry about weights. Multi-task architecture design (shared backbone + task heads) is covered in Anatomy of Deep Learning Architectures.

VII. Loss Function Design Checklist and Debugging ​

First Rule of the Road

If the loss doesn't decrease, check the data and code first — don't switch the loss function. 80% of "loss is stuck" issues are data problems: misaligned labels, NaN values, severe class imbalance, or data not being shuffled. Start by checking Data and Data Engineering and Debugging and Diagnostics.

When designing or selecting a loss function, go through this checklist:

  1. Is the task type matched? Classification/regression/ranking/contrastive? Are the output layer and loss paired (Section I)?
  2. Data distribution? Class imbalance? Many outliers? → Consider Focal/weighted/MAE.
  3. What metric are you truly trying to optimize? You evaluate using business metrics like accuracy/NDCG/FID, but the training loss is a differentiable surrogate for it. There may be a gap between the surrogate and the metric — this is called "optimization objective misalignment." See Deep Learning Evaluation and Experiments for evaluation guidelines.
  4. Numerical stability of the loss? Prefer framework-built-in fused implementations (e.g., CrossEntropyLoss) to avoid hand-written log-sum-exp overflow.
  5. Watch the loss magnitude at the start of training: the initial expected value for cross-entropy ≈ log(number of classes). If it's far off, either the initialization or the data has a problem.

VIII. Tradeoffs ​

Tradeoffs

Fitting the objective vs. optimization difficulty: MSE is smooth and easy to optimize but gets dragged by outliers; MAE is robust but non-differentiable at zero with slow convergence; Huber is a compromise but adds a hyperparameter. There is no "optimal loss," only "a loss that matches the data distribution."

Task metrics vs. differentiable surrogates: NDCG, AUC, FID, and other evaluation metrics are mostly non-differentiable and must be replaced with surrogate losses (ranking losses, contrastive losses, generative objectives). The closer the surrogate is to the metric, the more "honest" the training is — but it is typically harder to optimize.

Single loss vs. multi-loss weighting: Multiple losses (e.g., GAN's adversarial + feature matching, multi-task) inject more supervisory signals but bring balancing problems and tuning overhead. The simpler the loss, the easier it is to debug.

A loss function is a declaration of "what the model wants to become." It, the output layer, and the evaluation metrics must form a self-consistent loop — the output layer provides shape, the loss provides gradient direction, and the evaluation metric provides truth. For a complete methodology on evaluation, see the evaluation section. For a glossary of common losses, see Glossary.

Further Reading ​

References ​