Collaborative Knowledge Base Embedding for Recommender Systems

Fuzheng ZhangNicholas Jing YuanDefu LianXing XieWei-Ying Ma

article2016KDD1,834 citations

Proposes the Collaborative Knowledge Base Embedding framework to tackle interaction sparsity by jointly training collaborative filtering models with structural, textual, and visual representations extracted from knowledge bases via deep learning and TransR.

Listen

Online recommendation systems heavily rely on user interaction history, but they often struggle when data is sparse or when recommending newly introduced items. While auxiliary information from knowledge bases can alleviate these issues, existing solutions typically focus on single network structures and depend on labor-intensive manual feature engineering, ignoring rich unstructured signals such as text descriptions and visual imagery.

The article develops and evaluates Collaborative Knowledge Base Embedding, a unified recommendation framework that automatically extracts semantic features from structured relational data, text summaries, and visual images, jointly learning them alongside user interaction patterns.

The authors designed three specialized embedding modules: a Bayesian translation network model to capture heterogeneous entity relationships, stacked denoising auto-encoders for text summaries, and stacked convolutional auto-encoders to extract visual representations from item images like movie posters and book covers. These representations are unified into an item latent vector and trained end-to-end with collaborative filtering using implicit feedback ranking. The framework was evaluated across two large real-world datasets: the benchmark MovieLens-1M dataset (5,883 users, 3,230 movies) and IntentBooks, a web-scale dataset derived from Bing search logs (92,564 users, 18,475 books, mapped to Microsoft's Satori knowledge base).

Key findings show that the unified framework significantly and consistently outperforms competitive baseline recommendation methods across ranking accuracy and retrieval metrics. Structural relationship data provided the largest single performance boost among auxiliary inputs, followed by text summaries, while visual features provided a smaller yet statistically meaningful improvement. Automated deep embedding techniques demonstrated clear superiority over manual feature engineering and conventional topic modeling. Furthermore, joint end-to-end training outperformed two-stage pipelines where embeddings were learned separately from user feedback.

These results demonstrate that enterprises can enhance recommendation quality and user discovery without burdensome manual feature engineering by integrating diverse knowledge base data directly into model training. The approach reduces reliance on dense user interaction histories and provides a flexible blueprint for leveraging structured and unstructured data across broader information retrieval and search domains.

Organizations operating content-rich platforms should evaluate integrating multimodal knowledge bases into their core recommendation pipelines. Initial deployments can prioritize relational and textual features for the greatest immediate return, incorporating visual models as infrastructure allows. Further research should explore expanding this architecture to broader domain contexts and assessing computational scalability in real-time production environments.

Confidence in these findings is strong given the rigorous multi-run evaluations across two distinct domains. However, practical applicability depends on the availability of a well-maintained knowledge base and reliable entity-matching pipelines, which showed an observed mapping error rate of approximately 8% in the study.

Cover for Collaborative Knowledge Base Embedding for Recommender Systems

Abstract

Among different recommendation techniques, collaborative filtering usually suffer from limited performance due to the sparsity of user-item interactions. To address the issues, auxiliary information is usually used to boost the performance. Due to the rapid collection of information on the web, the knowledge base provides heterogeneous information including both structured and unstructured data with different semantics, which can be consumed by various applications. In this paper, we investigate how to leverage the heterogeneous information in a knowledge base to improve the quality of recommender systems. First, by exploiting the knowledge base, we design three components to extract items' semantic representations from structural content, textual content and visual content, respectively. To be specific, we adopt a heterogeneous network embedding method, termed as TransR, to extract items' structural representations by considering the heterogeneity of both nodes and relationships. We apply stacked denoising auto-encoders and stacked convolutional auto-encoders, which are two types of deep learning based embedding techniques, to extract items' textual representations and visual representations, respectively. Finally, we propose our final integrated framework, which is termed as Collaborative Knowledge Base Embedding (CKE), to jointly learn the latent representations in collaborative filtering as well as items' semantic representations from the knowledge base. To evaluate the performance of each embedding component as well as the whole system, we conduct extensive experiments with two real-world datasets from different scenarios. The results reveal that our approaches outperform several widely adopted state-of-the-art recommendation methods.

Table of Contents

  • 1. INTRODUCTION
  • 2. PRELIMINARY
  • 2.1 User Implicit Feedback
  • 2.2 Knowledge Base
  • 2.3 Problem Formulation
  • 3. OVERVIEW
  • 4. KNOWLEDGE BASE EMBEDDING
  • 4.1 Structural Embedding
  • 4.2 Textual Embedding
  • 4.3 Visual Embedding
  • 5. COLLABORATIVE JOINT LEARNING
  • 6. EXPERIMENTS
  • 6.1 Datasets Description
  • 6.2 Evaluation Schema
  • 6.3 Study of Structural Knowledge Usage
  • 6.4 Study of Textual Knowledge Usage
  • 6.5 Study of Visual Knowledge Usage
  • 6.6 Study of the Whole Framework
  • 7. RELATED WORK
  • 7.1 Knowledge Base for Recommendation
  • 7.2 Structural Knowledge Embedding
  • 7.3 Deep Learning for Recommendation
  • 8. CONCLUSION
  • 9. REFERENCES

Knowls

  1. Knowl 1 — Collaborative Knowledge Base Embedding Item Representation and Joint Ranking Model

    model/method

    Collaborative Knowledge Base Embedding (CKE) integrates implicit collaborative filtering with structural, textual, and visual representations extracted from a knowledge base. Given an implicit feedback matrix R∈Rm×n\mathbf{R} \in \mathbb{R}^{m \times n} where Rij=1R_{ij} = 1 indicates an observed interaction between user ii and item jj and Rij=0R_{ij} = 0 denotes unobserved interaction, CKE models user preferences via latent user vectors ui∈Rdim\mathbf{u}_i \in \mathbb{R}^{\text{dim}} and composite item latent vectors ej∈Rdim\mathbf{e}_j \in \mathbb{R}^{\text{dim}}.

    The composite item representation ej\mathbf{e}_j integrates an offset vector and three semantic embeddings:

    ej=ηj+vj+XLt2,j∗+ZLv2,j∗\mathbf{e}_j = \boldsymbol{\eta}_j + \mathbf{v}_j + \mathbf{X}_{\frac{L_t}{2}, j*} + \mathbf{Z}_{\frac{L_v}{2}, j*}

    where ηj∼N(0,λI−1I)\boldsymbol{\eta}_j \sim \mathcal{N}(0, \lambda_I^{-1} \mathbf{I}) is an item-specific offset latent vector, vj\mathbf{v}_j is the entity structural embedding from Bayesian TransR, XLt2,j∗\mathbf{X}_{\frac{L_t}{2}, j*} is the central hidden layer textual embedding from an LtL_t-layer Bayesian Stacked Denoising Auto-encoder (SDAE), and ZLv2,j∗\mathbf{Z}_{\frac{L_v}{2}, j*} is the central hidden layer visual embedding from an LvL_v-layer Bayesian Stacked Convolutional Auto-encoder (SCAE).

    For a user ii and a pair of items (j,j′)(j, j') where Rij=1R_{ij} = 1 and Rij′=0R_{ij'} = 0, the pairwise preference probability is parameterized as:

    p(j>j′;i∣θ)=σ(uiTej−uiTej′)p(j > j'; i \mid \theta) = \sigma\left(\mathbf{u}_i^T \mathbf{e}_j - \mathbf{u}_i^T \mathbf{e}_{j'}\right)

    where σ(x)=11+e−x\sigma(x) = \frac{1}{1 + e^{-x}} is the logistic sigmoid function, and θ\theta denotes the model parameters.

  2. Knowl 2 — Complete CKE Joint MAP Objective Function

    equation

    The parameters u,e,r,M,W,b,Q,c\mathbf{u}, \mathbf{e}, \mathbf{r}, \mathbf{M}, \mathbf{W}, \mathbf{b}, \mathbf{Q}, \mathbf{c} of the Collaborative Knowledge Base Embedding (CKE) framework are learned by maximizing the joint log-posterior objective function L\mathcal{L}:

    \begin{aligned} \mathcal{L} &= \sum_{(i, j, j') \in \mathcal{D}} \ln \sigma\left(\mathbf{u}_i^T \mathbf{e}_j - \mathbf{u}_i^T \mathbf{e}_{j'}\right) \\ &\quad - \frac{\lambda_X}{2} \sum_l \|\sigma(\mathbf{X}_{l-1} \mathbf{W}_l + \mathbf{b}_l) - \mathbf{X}_l\|_2^2 \\ &\quad + \sum_{(v_h, r, v_t, v_{t'}) \in \mathcal{S}} \ln \sigma\left(\|\mathbf{v}_h \mathbf{M}_r + \mathbf{r} - \mathbf{v}_{t'} \mathbf{M}_r\|_2^2 - \|\mathbf{v}_h \mathbf{M}_r + \mathbf{r} - \mathbf{v}_t \mathbf{M}_r\|_2^2\right) \\ &\quad - \frac{\lambda_Z}{2} \sum_{l \notin \{\frac{L_v}{2}, \frac{L_v}{2}+1\}} \|\sigma(\mathbf{Z}_{l-1} * \mathbf{Q}_l + \mathbf{c}_l) - \mathbf{Z}_l\|_2^2 \\ &\quad - \frac{\lambda_Z}{2} \sum_{l \in \{\frac{L_v}{2}, \frac{L_v}{2}+1\}} \|\sigma(\mathbf{Z}_{l-1} \mathbf{Q}_l + \mathbf{c}_l) - \mathbf{Z}_l\|_2^2 \\ &\quad - \frac{\lambda_U}{2} \sum_i \|\mathbf{u}_i\|_2^2 - \frac{\lambda_v}{2} \sum_v \|\mathbf{v}\|_2^2 - \frac{\lambda_r}{2} \sum_r \|\mathbf{r}\|_2^2 - \frac{\lambda_M}{2} \sum_r \|\mathbf{M}_r\|_F^2 \\ &\quad - \frac{1}{2} \sum_l \left(\lambda_W \|\mathbf{W}_l\|_F^2 + \lambda_b \|\mathbf{b}_l\|_2^2\right) - \frac{1}{2} \sum_l \left(\lambda_Q \|\mathbf{Q}_l\|_F^2 + \lambda_c \|\mathbf{c}_l\|_2^2\right) \\ &\quad - \frac{\lambda_I}{2} \sum_j \|\mathbf{e}_j - \mathbf{v}_j - \mathbf{X}_{\frac{L_t}{2}, j*} - \mathbf{Z}_{\frac{L_v}{2}, j*\|_2^2 \end{aligned}

    where D={(i,j,j′)∣Rij=1∧Rij′=0}\mathcal{D} = \{(i, j, j') \mid R_{ij} = 1 \land R_{ij'} = 0\} denotes pairwise interaction triples, S\mathcal{S} is the set of quadruples (vh,r,vt,vt′)(v_h, r, v_t, v_{t'}) consisting of a valid relation triple (vh,r,vt)(v_h, r, v_t) and a corrupted negative triple (vh,r,vt′)(v_h, r, v_{t'}), Xl\mathbf{X}_l and Zl\mathbf{Z}_l are layer activations for textual SDAE and visual SCAE, ∗* denotes convolution, and λU,λI,λv,λr,λM,λX,λZ,λW,λb,λQ,λc\lambda_U, \lambda_I, \lambda_v, \lambda_r, \lambda_M, \lambda_X, \lambda_Z, \lambda_W, \lambda_b, \lambda_Q, \lambda_c are Gaussian precision hyperparameters.

  3. Knowl 3 — Bayesian TransR for Structural Knowledge Base Embedding

    model/method

    Structural knowledge is represented as a heterogeneous network G=(V,E)G = (\mathcal{V}, \mathcal{E}) where V\mathcal{V} is the entity set and E\mathcal{E} is the set of typed relation edges. Bayesian TransR models entities and relations in distinct semantic spaces bridged by relation-specific projection matrices. For a triple (vh,r,vt)(v_h, r, v_t), entities are embedded as vectors vh,vt∈Rk\mathbf{v}_h, \mathbf{v}_t \in \mathbb{R}^k, the relation as r∈Rd\mathbf{r} \in \mathbb{R}^d, and the relation projection matrix as Mr∈Rk×d\mathbf{M}_r \in \mathbb{R}^{k \times d}.

    The projected entity representations in relation space are defined as:

    vhr=vhMr,vtr=vtMr\mathbf{v}_h^r = \mathbf{v}_h \mathbf{M}_r, \quad \mathbf{v}_t^r = \mathbf{v}_t \mathbf{M}_r

    The dissimilarity score function for the triple is:

    fr(vh,vt)=∥vhr+r−vtr∥22f_r(v_h, v_t) = \|\mathbf{v}_h^r + \mathbf{r} - \mathbf{v}_t^r\|_2^2

    The Bayesian generative process specifies zero-mean spherical Gaussian priors:

    v∼N(0,λv−1I),r∼N(0,λr−1I),Mr∼N(0,λM−1I)\mathbf{v} \sim \mathcal{N}(0, \lambda_v^{-1} \mathbf{I}), \quad \mathbf{r} \sim \mathcal{N}(0, \lambda_r^{-1} \mathbf{I}), \quad \mathbf{M}_r \sim \mathcal{N}(0, \lambda_M^{-1} \mathbf{I})

    For each quadruple (vh,r,vt,vt′)∈S(v_h, r, v_t, v_{t'}) \in \mathcal{S}, where (vh,r,vt)(v_h, r, v_t) is a valid knowledge base triple and (vh,r,vt′)(v_h, r, v_{t'}) is a corrupted negative triple (formed by replacing vtv_t with an entity of the same type), the pairwise ranking probability is:

    p((vh,r,vt,vt′))=σ(fr(vh,vt′)−fr(vh,vt))p((v_h, r, v_t, v_{t'})) = \sigma\left(f_r(v_h, v_{t'}) - f_r(v_h, v_t)\right)

    For an item entity jj, its embedding vector vj\mathbf{v}_j learned via this process serves as its structural representation in the recommendation framework.

  4. Knowl 4 — Bayesian Stacked Denoising Auto-Encoder for Textual Embedding

    model/method

    Textual knowledge for items (e.g., plots or book descriptions) is embedded using an LtL_t-layer Bayesian Stacked Denoising Auto-encoder (Bayesian SDAE). The clean textual input matrix is denoted by XLt\mathbf{X}_{L_t}, where the jj-th row XLt,j∗\mathbf{X}_{L_t, j*} is the bag-of-words representation for item entity jj. A corrupted input matrix X0\mathbf{X}_0 is constructed by randomly masking entries of XLt\mathbf{X}_{L_t} to zero with noise masking level ϵ\epsilon.

    The encoder comprises the first Lt2\frac{L_t}{2} layers mapping X0\mathbf{X}_0 to a latent representation XLt2\mathbf{X}_{\frac{L_t}{2}}, and the decoder comprises the remaining Lt2\frac{L_t}{2} layers reconstructing the clean input XLt\mathbf{X}_{L_t}.

    The generative process for each layer l∈{1,…,Lt}l \in \{1, \dots, L_t\} is:

    1. Draw weight matrix Wl∼N(0,λW−1I)\mathbf{W}_l \sim \mathcal{N}(0, \lambda_W^{-1} \mathbf{I}).
    2. Draw bias vector bl∼N(0,λb−1I)\mathbf{b}_l \sim \mathcal{N}(0, \lambda_b^{-1} \mathbf{I}).
    3. Draw layer output activations Xl∼N(σ(Xl−1Wl+bl),λX−1I)\mathbf{X}_l \sim \mathcal{N}\left(\sigma(\mathbf{X}_{l-1} \mathbf{W}_l + \mathbf{b}_l), \lambda_X^{-1} \mathbf{I}\right), where σ\sigma is the element-wise sigmoid activation function.

    The vector XLt2,j∗\mathbf{X}_{\frac{L_t}{2}, j*} at the bottleneck layer serves as the textual semantic embedding for item jj.

  5. Knowl 5 — Bayesian Stacked Convolutional Auto-Encoder for Visual Embedding

    model/method

    Visual knowledge of items (e.g., movie poster images or book covers) is embedded using an LvL_v-layer Bayesian Stacked Convolutional Auto-encoder (Bayesian SCAE). The clean input ZLv\mathbf{Z}_{L_v} is a 4-dimensional tensor whose jj-th slice ZLv,j∗\mathbf{Z}_{L_v, j*} is a 3×64×643 \times 64 \times 64 RGB pixel tensor for item jj. The corrupted input Z0\mathbf{Z}_0 is generated by adding zero-mean Gaussian noise with standard deviation σ\sigma to masked entries of ZLv\mathbf{Z}_{L_v}.

    The architecture uses convolutional layers for l∉{Lv2,Lv2+1}l \notin \{\frac{L_v}{2}, \frac{L_v}{2}+1\} and fully connected layers for the central bottleneck transition l∈{Lv2,Lv2+1}l \in \{\frac{L_v}{2}, \frac{L_v}{2}+1\}.

    The generative process for each layer l∈{1,…,Lv}l \in \{1, \dots, L_v\} is:

    1. Draw weight parameter tensor/matrix Ql∼N(0,λQ−1I)\mathbf{Q}_l \sim \mathcal{N}(0, \lambda_Q^{-1} \mathbf{I}).
    2. Draw bias parameter cl∼N(0,λc−1I)\mathbf{c}_l \sim \mathcal{N}(0, \lambda_c^{-1} \mathbf{I}).
    3. Draw layer output activations:
      • If layer ll is fully connected: Zl∼N(σ(Zl−1Ql+cl),λZ−1I)\mathbf{Z}_l \sim \mathcal{N}\left(\sigma(\mathbf{Z}_{l-1} \mathbf{Q}_l + \mathbf{c}_l), \lambda_Z^{-1} \mathbf{I}\right).
      • If layer ll is convolutional/deconvolutional: Zl∼N(σ(Zl−1∗Ql+cl),λZ−1I)\mathbf{Z}_l \sim \mathcal{N}\left(\sigma(\mathbf{Z}_{l-1} * \mathbf{Q}_l + \mathbf{c}_l), \lambda_Z^{-1} \mathbf{I}\right), where ∗* is the convolution operator.

    The row vector ZLv2,j∗\mathbf{Z}_{\frac{L_v}{2}, j*} from the central representation matrix ZLv2\mathbf{Z}_{\frac{L_v}{2}} serves as the visual embedding of item entity jj.

  6. Knowl 6 — Stochastic Gradient Descent Training Procedure for CKE

    algorithm

    CKE parameters are jointly optimized by maximizing the log-likelihood objective function using Stochastic Gradient Descent (SGD).

    Input: User implicit feedback matrix R∈Rm×nR \in \mathbb{R}^{m \times n}, structural knowledge graph G=(V,E)G=(\mathcal{V}, \mathcal{E}), textual data XLtX_{L_t}, visual data ZLvZ_{L_v}, hyperparameters λU,λI,λv,λr,λM,λX,λZ,λW,λb,λQ,λc\lambda_U, \lambda_I, \lambda_v, \lambda_r, \lambda_M, \lambda_X, \lambda_Z, \lambda_W, \lambda_b, \lambda_Q, \lambda_c, learning rate γ\gamma
    Output: User latent vectors {ui}\{\mathbf{u}_i\}, item latent vectors {ej}\{\mathbf{e}_j\}
    Initialize parameters u,e,v,r,M,W,b,Q,c\mathbf{u}, \mathbf{e}, \mathbf{v}, \mathbf{r}, \mathbf{M}, \mathbf{W}, \mathbf{b}, \mathbf{Q}, \mathbf{c} randomly from their Gaussian priors
    Construct corrupted textual matrix X0X_0 and corrupted visual tensor Z0Z_0
    repeat
        Sample a positive user-item interaction (i,j)(i, j) where Rij=1R_{ij} = 1
        Sample an unobserved item j′j' where Rij′=0R_{ij'} = 0
        Identify subset Sj,j′⊂S\mathcal{S}_{j, j'} \subset \mathcal{S} containing quadruples involving item entities jj or j′j'
        
        Compute gradient of joint objective L\mathcal{L} with respect to ui,ej,ej′\mathbf{u}_i, \mathbf{e}_j, \mathbf{e}_{j'}
        Compute gradient of L\mathcal{L} with respect to structural parameters v,r,Mr\mathbf{v}, \mathbf{r}, \mathbf{M}_r using quadruples in Sj,j′\mathcal{S}_{j, j'}
        Compute gradient of L\mathcal{L} with respect to textual SDAE parameters Wl,bl\mathbf{W}_l, \mathbf{b}_l and activations XlX_l
        Compute gradient of L\mathcal{L} with respect to visual SCAE parameters Ql,cl\mathbf{Q}_l, \mathbf{c}_l and activations ZlZ_l
        
        Update all parameters using SGD: θ←θ+γ∇θL\theta \leftarrow \theta + \gamma \nabla_\theta \mathcal{L}
    until convergence
    return {ui}\{\mathbf{u}_i\} and {ej}\{\mathbf{e}_j\}

    Recommendations for a user ii are generated by sorting candidate items jj in descending order of the predicted score:

    y^ij=uiTej\hat{y}_{ij} = \mathbf{u}_i^T \mathbf{e}_j

  7. Knowl 7 — Experimental Datasets and Knowledge Base Statistics

    data/table

    The experiments evaluate recommendation performance on two datasets linked to Microsoft's Satori knowledge base:

    1. MovieLens-1M: Implicit feedback extracted by taking rating 5 as positive interactions and filtering users with fewer than 3 positive ratings. Movies are matched to Satori knowledge entities with a verified match precision of 92% (134 movies could not be mapped). Structural subgraphs contain 1-step relations including genre, director, writer, actors, language, country, production date, rating, nominated awards, and received awards.
    2. IntentBooks: User book interests derived from Bing search query logs (Sep 2014 to Jun 2015) using similarity computation and classification (precision: 91.5%), filtering users with fewer than 5 book interactions. Structural subgraphs include 1-step entities: genre, author, publish date, belonged series, language, and rating.
    Statistic MovieLens-1M IntentBooks
    #user 5,883 92,564
    #item 3,230 18,475
    #interactions 226,101 897,871
    #sk nodes 84,011 26,337
    #sk edges 169,368 57,408
    #sk edge types 10 6
    #tk items 2,752 17,331
    #vk items 2,958 16,719

    In the table, #sk nodes, #sk edges, and #sk edge types indicate structural knowledge node, edge, and relation counts; #tk items and #vk items indicate the number of items possessing textual summaries and visual images, respectively.

  8. Knowl 8 — Hyperparameter Settings for CKE on Benchmark Datasets

    data/table

    Optimal hyperparameters determined via validation sets (70% train / 30% test split randomly partitioned 5 times) for collaborative filtering (cf), structural knowledge embedding (sk), textual knowledge embedding (tk), and visual knowledge embedding (vk) in CKE:

    Component MovieLens-1M IntentBooks
    cf dim=150,λU=λI=0.0025\text{dim}=150, \lambda_U=\lambda_I=0.0025 dim=100,λU=λI=0.005\text{dim}=100, \lambda_U=\lambda_I=0.005
    sk λv=λr=0.001,λM=0.01\lambda_v=\lambda_r=0.001, \lambda_M=0.01 λv=λr=0.001,λM=0.1\lambda_v=\lambda_r=0.001, \lambda_M=0.1
    tk λW=λb=0.01,λX=0.0001,\lambda_W=\lambda_b=0.01, \lambda_X=0.0001, λW=λb=0.01,λX=0.001,\lambda_W=\lambda_b=0.01, \lambda_X=0.001,
    ϵ=0.2,Lt=4,Nl=300\epsilon=0.2, L_t=4, N_l=300 ϵ=0.1,Lt=6,Nl=200\epsilon=0.1, L_t=6, N_l=200
    vk λQ=λc=0.01,λZ=0.0001,σ=3,\lambda_Q=\lambda_c=0.01, \lambda_Z=0.0001, \sigma=3, λQ=λc=0.01,λZ=0.001,σ=2,\lambda_Q=\lambda_c=0.01, \lambda_Z=0.001, \sigma=2,
    Lv=6,Nf=20,Sf=(5,5)L_v=6, N_f=20, S_f=(5,5) Lv=8,Nf=20,Sf=(5,5)L_v=8, N_f=20, S_f=(5,5)

    Here, dim\text{dim} is the latent vector dimension, ϵ\epsilon is the input masking ratio, σ\sigma is the visual Gaussian noise standard deviation, LtL_t and LvL_v are total network layer depths, NlN_l is the hidden unit count in intermediate textual SDAE layers, NfN_f is the number of convolutional filter maps, and SfS_f is the convolutional kernel filter size.

  9. Knowl 9 — Ablation and Empirical Performance of Individual Knowledge Base Modalities

    empirical result

    Evaluating individual knowledge embedding components under MAP@K and Recall@K across MovieLens-1M and IntentBooks demonstrates the following:

    1. Structural Knowledge (CKE(S)): CKE(S) outperforms all structural baselines (BPRMF+TransE, PER, LIBFM(S), PRP, and BPRMF). Standard BPRMF performs worst, showing the substantial value of structural data. CKE(S) outperforms BPRMF+TransE, confirming that TransR's distinct relation projection spaces capture entity and relation heterogeneity better than TransE's single-space translation. Network embedding approaches (CKE(S) and BPRMF+TransE) outperform manual meta-path feature engineering (PER) and feature factorization (LIBFM(S)).
    2. Textual Knowledge (CKE(T)): CKE(T) outperforms LIBFM(T), CMF(T), and pure collaborative filtering (BPRMF). CKE(T) also consistently surpasses Collaborative Topic Regression (CTR), demonstrating that deep representation learning via SDAE extracts semantic representations superior to topic modeling.
    3. Visual Knowledge (CKE(V)): CKE(V) outperforms LIBFM(V), CMF(V), and BPRMF+SDAE(V). The clear performance margin of CKE(V) over BPRMF+SDAE(V) demonstrates that convolutional layers (SCAE) preserve spatial locality and image neighborhood relations substantially better than fully connected denoising auto-encoders.
    4. Relative Modality Contribution: Across both datasets, structural knowledge delivers the largest individual performance improvement, followed by textual knowledge, while visual knowledge delivers a smaller but statistically significant improvement.
  10. Knowl 10 — Performance Superiority of Full CKE Integration and Collaborative Joint Learning

    empirical result

    Empirical comparison of the complete model, CKE(STV), against partial multimodal variants and competing baselines yields the following key results:

    1. Multimodal Synergy: CKE(STV) achieves the highest MAP@K and Recall@K scores across both MovieLens-1M and IntentBooks, consistently outperforming two-modality ablations: CKE(ST), CKE(SV), and CKE(TV). This confirms that structural, textual, and visual knowledge sources provide complementary semantic information.
    2. Joint Learning vs. Decoupled Learning: CKE(STV) significantly outperforms BPRMF+STV (which trains structural TransR, textual SDAE, and visual SCAE independently in advance and feeds fixed embeddings into collaborative ranking). Jointly optimizing knowledge embeddings and implicit feedback ranking aligns feature extraction directly with the recommendation objective.
    3. Deep Knowledge Embedding vs. Feature Engineering: CKE(STV) substantially outperforms LIBFM(STV) (which combines raw structural, textual, and visual features in a factorization machine), confirming the superiority of automated deep network embedding over raw feature combinations.

Coverage note — None was omitted; all contributed models, auto-encoder architectures, objective functions, optimization algorithms, dataset statistics, hyperparameters, and experimental ablation results are covered.

References

  1. 1.ahmet uyar, f. m. a. Evaluating search features of google knowledge graph and bing satori. Online Information Review (2015).
  2. 2.Bobadilla, J., Ortega, F., Hernando, A., and Gutiérrez, A. Recommender systems survey. Knowledge-Based Systems 46 (2013), 109–132.
  3. 3.Bordes, A., Usunier, N., Garcia-Duran, A., Weston, J., and Yakhnenko, O. Translating embeddings for modeling multi-relational data. In Advances in Neural Information Processing Systems (2013), 2787–2795.
  4. 4.Cambria, E., Schuller, B., Liu, B., Wang, H., and Havasi, C. Knowledge-based approaches to concept-level sentiment analysis. IEEE Intelligent Systems, 2 (2013), 12–14.
  5. 5.Cheekula, S. K., Kapanipathi, P., Doran, D., Jain, P., and Sheth, A. Entity recommendations using hierarchical knowledge bases.
  6. 6.Chui, M., Manyika, J., and Kuiken, S. What executives should know about open data. McKinsey Quarterly, January 2014 (2014).
  7. 7.Ciresan, D. C., Meier, U., Masci, J., Maria Gambardella, L., and Schmidhuber, J. Flexible, high performance convolutional neural networks for image classification. In IJCAI Proceedings-International Joint Conference on Artificial Intelligence, vol. 22 (2011), 1237.
  8. 8.Elkahky, A. M., Song, Y., and He, X. A multi-view deep learning approach for cross domain user modeling in recommendation systems. In Proceedings of the 24th International Conference on World Wide Web, International World Wide Web Conferences Steering Committee (2015), 278–288.
  9. 9.Freitas, A., Oliveira, J. G., OâA˘ ZRiain, S., Curry, E., and Da Silva, J. ´ C. P. Querying linked data using semantic relatedness: a vocabulary independent approach. In Natural Language Processing and Information Systems. Springer, 2011, 40–51.
  10. 10.Gantner, Z., Rendle, S., Freudenthaler, C., and Schmidt-Thieme, L. Mymedialite: A free recommender system library. In Proceedings of the fifth ACM conference on Recommender systems, ACM (2011), 305–308.
  11. 11.Grad-Gyenge, L., Filzmoser, P., and Werthner, H. Recommendations on a knowledge graph.
  12. 12.Huang, P.-S., He, X., Gao, J., Deng, L., Acero, A., and Heck, L. Learning deep structured semantic models for web search using clickthrough data. In Proceedings of the 22nd ACM international conference on Conference on information & knowledge management, ACM (2013), 2333–2338.
  13. 13.Ji, H., and Grishman, R. Knowledge base population: Successful approaches and challenges. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1, Association for Computational Linguistics (2011), 1148–1158.
  14. 14.Lian, D., Zhao, C., Xie, X., Sun, G., Chen, E., and Rui, Y. Geomf: joint geographical modeling and matrix factorization for point-of-interest recommendation. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, ACM (2014), 831–840.
  15. 15.Lin, Y., Liu, Z., Sun, M., Liu, Y., and Zhu, X. Learning entity and relation embeddings for knowledge graph completion. In Proceedings of AAAI (2015).
  16. 16.Masci, J., Meier, U., Cire¸san, D., and Schmidhuber, J. Stacked convolutional auto-encoders for hierarchical feature extraction. In Artificial Neural Networks and Machine Learning–ICANN 2011. Springer, 2011, 52–59.
  17. 17.Nguyen, P., Tomeo, P., Di Noia, T., and Di Sciascio, E. An evaluation of simrank and personalized pagerank to build a recommender system for the web of data. In Proceedings of the 24th International Conference on World Wide Web Companion, International World Wide Web Conferences Steering Committee (2015), 1477–1482.
  18. 18.Passant, A. dbrecâA˘Tmusic recommendations using dbpedia. In ˇ The Semantic Web–ISWC 2010. Springer, 2010, 209–224.
  19. 19.Purushotham, S., Liu, Y., and Kuo, C.-C. J. Collaborative topic regression with social matrix factorization for recommendation systems. arXiv preprint arXiv:1206.4684 (2012).
  20. 20.Ren, J. S., and Xu, L. On vectorization of deep convolutional neural networks for vision tasks. In Twenty-Ninth AAAI Conference on Artificial Intelligence (2015).
  21. 21.Rendle, S. Factorization machines with libfm. ACM Transactions on Intelligent Systems and Technology (TIST) 3, 3 (2012), 57.
  22. 22.Rendle, S., Freudenthaler, C., Gantner, Z., and Schmidt-Thieme, L. Bpr: Bayesian personalized ranking from implicit feedback. In Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence, AUAI Press (2009), 452–461.
  23. 23.Schafer, J. B., Frankowski, D., Herlocker, J., and Sen, S. Collaborative filtering recommender systems. In The adaptive web. Springer, 2007, 291–324.
  24. 24.Scheffler, T., Schirru, R., and Lehmann, P. Matching points of interest from different social networking sites. In KI 2012: Advances in Artificial Intelligence. Springer, 2012, 245–248.
  25. 25.Tiddi, I., d’Aquin, M., and Motta, E. Using linked data traversal to label academic communities. In Proceedings of the 24th International Conference on World Wide Web Companion, International World Wide Web Conferences Steering Committee (2015), 1029–1034.
  26. 26.Van den Oord, A., Dieleman, S., and Schrauwen, B. Deep content-based music recommendation. In Advances in Neural Information Processing Systems (2013), 2643–2651.
  27. 27.Vincent, P., Larochelle, H., Lajoie, I., Bengio, Y., and Manzagol, P.-A. Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. The Journal of Machine Learning Research 11 (2010), 3371–3408.
  28. 28.Wang, C., and Blei, D. M. Collaborative topic modeling for recommending scientific articles. In Proceedings of the 17th ACM SIGKDD international conference on Knowledge discovery and data mining, ACM (2011), 448–456.
  29. 29.Wang, H., Wang, N., and Yeung, D.-Y. Collaborative deep learning for recommender systems. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’15, ACM (New York, NY, USA, 2015), 1235–1244.
  30. 30.Yu, X., Ren, X., Sun, Y., Gu, Q., Sturt, B., Khandelwal, U., Norick, B., and Han, J. Personalized entity recommendation: A heterogeneous information network approach. In Proceedings of the 7th ACM international conference on Web search and data mining, ACM (2014), 283–292.
  31. 31.Zhou, K., and Zha, H. Learning binary codes for collaborative filtering. In Proceedings of the 18th ACM SIGKDD international conference on Knowledge discovery and data mining, ACM (2012), 498–506.

Citation

MLA
Zhang, F., et al. “Collaborative Knowledge Base Embedding for Recommender Systems”. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016, pp. 353–62, https://doi.org/10.1145/2939672.2939673.
APA
Zhang, F., Yuan, N. J., Lian, D., Xie, X., & Ma, W.-Y. (2016). Collaborative Knowledge Base Embedding for Recommender Systems. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 353–362. https://doi.org/10.1145/2939672.2939673
Chicago
Zhang, F., N. J. Yuan, D. Lian, X. Xie, and W.-Y. Ma. 2016. “Collaborative Knowledge Base Embedding for Recommender Systems”. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 353–62. https://doi.org/10.1145/2939672.2939673.
Harvard
Zhang, F. et al. (2016) “Collaborative Knowledge Base Embedding for Recommender Systems”, Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, pp. 353–362. Available at: https://doi.org/10.1145/2939672.2939673.
Vancouver
1. Zhang F, Yuan NJ, Lian D, Xie X, Ma W-Y (2016) Collaborative Knowledge Base Embedding for Recommender Systems. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, pp 353–362

BibTeX

@inproceedings{Zhang_2016, series={KDD ’16}, title={Collaborative Knowledge Base Embedding for Recommender Systems}, url={http://dx.doi.org/10.1145/2939672.2939673}, DOI={10.1145/2939672.2939673}, booktitle={Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining}, publisher={ACM}, author={Zhang, Fuzheng and Yuan, Nicholas Jing and Lian, Defu and Xie, Xing and Ma, Wei-Ying}, year={2016}, month=Aug, pages={353–362}, collection={KDD ’16} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF