Skip to content

Deep Learning Recommender Systems

Quick overview Recommender systems are the most profitable battlefield for deep learning in internet industry. This article covers the four-layer architecture of recall/ranking/re-ranking, from matrix factorization to neural collaborative filtering, two-tower recall, SASRec sequential recommendation, DCN/Wide&Deep CTR ranking, PinSage graph-based recommendation, and MMoE multi-objective learning, with practical perspectives on cold start, evaluation, and deployment.

Deep Learning Recommender Systems ​

In a sentence: Deep learning recommender systems model "user-item interactions" as probability prediction problems in high-dimensional feature space — quickly recalling from billions of candidates, then precisely ranking among thousands, ultimately turning "guessing what you like" into a real-time, personalized, measurable industrial pipeline — it is one of the scenarios with the most direct deep learning value (data and evaluation infrastructure span Data and Data Engineering and Deep Learning Evaluation and Experimentation).

1. The Overall Architecture of Recommender Systems ​

Industrial recommender systems are "funnels": data volume shrinks and models grow more refined as you go down.

LayerCandidate ScaleGoalCharacteristics
RecallBillions → ThousandsFast coarse filtering, maximize recallTwo-tower, vector retrieval, rules (trending/regional)
Pre-rankingThousands → HundredsIntermediate filteringLightweight model, balancing speed and accuracy
RankingHundreds → DozensPrecise scoring and sortingCTR/CVR models, heavy feature engineering
Re-rankingDozens → Final listDiversity, rules, business constraintsShuffling, MMR, business policies

This two-stage "recall + ranking" structure (with pre-ranking added as a performance optimization) almost universally spans all major tech companies. Understanding the full picture is more important than chasing any individual model — this is the "systems thinking" embodied in DL Design Principles.

2. From Collaborative Filtering to Matrix Factorization to Neural ​

  • Collaborative Filtering (CF): The core assumption is "similar people like similar things" (UserCF/ItemCF). The advantage is that it needs no content features; the drawback is sparsity and cold start;
  • Matrix Factorization (MF): Decomposes the "user × item" interaction matrix into two low-dimensional latent factor matrices $R \approx U V^\top$, predicting scores via latent vector dot product. SVD and Funk-SVD (2006) are classics. MF is the first success of "representation learning" in recommendation (the vectorization philosophy is covered in Representation Learning and Pretraining);
  • Neural Collaborative Filtering (NCF, 2017): Uses MLPs to replace dot products, modeling "nonlinear interactions between users and items." But what truly dominated industry combined MF's vectorization with large-scale feature engineering via two-tower / CTR models.

3. Two-Tower Models: The Recall Standard ​

Two-Tower (representative: DSSM, 2013) encodes users and items each into a single vector, approximating match quality via dot product:

User features → User tower → u_vec ┐
                                    ├→ dot product → score
Item features → Item tower → i_vec ┘

Key design: The item tower's output can be precomputed offline at full scale and indexed as vectors, so the online system only computes user vectors and performs ANN (approximate nearest neighbor) search, retrieving hundreds of millions of items in milliseconds. It sacrifices user-item cross features (the dot product is shallow) in exchange for scalability. Improvement directions include using attention to aggregate multiple behaviors per user (Multi-Interest) and graph-neighbor-enhanced representations.

4. Sequential Recommendation: SASRec ​

User clicks/purchases form ordered sequences that hide clues about "what the user wants next." Sequential recommendation encodes the "last N interactions" into a representation:

  • SASRec (2018) treats interaction sequences like sentences, using self-attention to model relationships between any two items (faster than RNNs, deeper than Markov chains) — it is the most direct application of Transformer Architecture to recommendation;
  • Industrial variants: use "behavior sequence + target item" for next-item prediction, often fused with CTR ranking;
  • Details on attention and sequence modeling are in Attention Mechanisms.

5. CTR Ranking: Wide&Deep and DCN ​

The core of ranking is CTR prediction: given massive features (user profiles, item attributes, context), predict click probability. Features are mostly high-cardinality sparse categorical values (user IDs, categories), first embedded, then fed to the network. Two classic architectures:

  • Wide & Deep (2016, Google): The Wide side preserves linear cross features (memory, e.g., "male ∧ football" strong rules), while the Deep side does embedding + MLP (generalization). The design philosophy of "memory and generalization coexisting" influenced virtually all subsequent ranking models;
  • DCN (Deep & Cross Network, 2017): Uses a cross layer to explicitly perform high-order feature cross, capturing "interactions between features" rather than relying on MLPs to learn them implicitly. Subsequent DCN-V2 replaces the cross layer with low-rank decomposition to reduce parameters.

Ranking feature engineering best practices: bin continuous features, match embedding dimensionality to cardinality, ensure feature timeliness — see Training Recipes and Hyperparameter Tuning.

6. Graph-Based Recommendation: PinSage ​

Items and users, items and items, are naturally graphs (who viewed what, who co-watched with whom). PinSage (2018, Pinterest) uses graph neural networks (GNNs) to learn "neighbor-aggregated" representations for items: multi-layer neighbor sampling + aggregation encodes graph structure into vectors. Its industrial breakthrough was scalability — defining neighbors via random walks, batch training offline, building indexes for online serving. It is a landmark success of Graph Neural Networks at real-world massive scale.

7. Multi-Objective Learning: MMoE ​

Real businesses never have just one objective: click-through rate (CTR), conversion rate (CVR), watch time, completion rate, positive feedback rate, merchant satisfaction… often conflicting (high clicks but poor conversions). MMoE (Multi-gate Mixture-of-Experts, 2018) uses multiple expert networks + gate-weighted combinations, letting each objective share some capacity while maintaining its own focus. It is a representative work of "multi-task learning" (more general multi-task frameworks see the fusion philosophy in Multimodal Models).

8. Cold Start and Bias Correction ​

  • Cold start: New users / new items have no interaction data. Common approaches: content-feature fallback (category, brand), exploratory strategies (diversity injection, exploration-exploitation), cross-domain transfer (leveraging behavior from other product lines), meta-learning;
  • Position bias: Items ranked higher naturally get more clicks; naively learning CTR amplifies the position effect — requires adding position features and zeroing them at inference time, or using debiasing sample weighting methods;
  • Popularity bias: Popular items get more clicks, making the model prone to a "Matthew effect" — only promoting top items and starving the long tail.

9. Evaluation and Online Experimentation ​

  • Offline metrics: AUC (ranking quality), GAUC (AUC weighted by user-group, more representative of real distributions), NDCG/Recall@K (recall side). A 0.1% offline AUC lift is big news in industry, but offline improvement ≠ online revenue;
  • Online experimentation: A/B tests, gradual rollouts, stratified experiments (Overlapping Experimentation). Online metrics (watch time, retention, business KPIs) are the ultimate arbiters;
  • Feedback loops: Feeding online data back into training sets (real-time features, negative sample construction), but watch for data leakage — mistaking "items that were recommended" for "items the user actively chose" is the most common pitfall. See Evaluation in Practice for evaluation methodology.

The Biggest Pitfall in Recommender Systems

In real systems, data, architecture, and the experiment platform typically contribute more to results than the model itself. A ranking model with 1% higher AUC is often less impactful than a successful recall enhancement or faster online latency. Get the data pipeline and evaluation loop right first, then talk about models.

10. New Developments in the LLM Era ​

  • LLM-assisted recommendation: Use LLMs to generate user profile summaries, item semantic tags, cold-start recommendation rationales, or even power conversational recommendation directly (see Large Language Models (LLM));
  • Multimodal features: Cover images, titles, and video pre-features are new information sources; multimodal embeddings are covered in the Multimodal Models article;
  • RL-based ranking: Treat ranking as sequential decision-making (Deep Q-Learning to "play with watch time"), echoing the reward modeling of Deep Reinforcement Learning Applications.

Further Reading ​

References ​