Neural Factorization Machines for Sparse Predictive Analytics

Xiangnan HeTat-Seng Chua

article2017SIGIR1,438 citations

Proposes Neural Factorization Machines, a model that unifies linear factorization machines with non-linear neural networks to capture higher-order feature interactions in sparse data, delivering superior predictive accuracy with simpler training than deeper alternatives.

Listen

Online applications such as personalized recommendation and targeted advertising rely heavily on sparse categorical data, including user identities, demographics, and context. Standard machine learning models struggle to capture the complex, non-linear feature interactions hidden within this sparse data without expensive, manual feature engineering. While traditional Factorization Machines efficiently capture pairwise relationships linearly, they lack expressiveness for complex patterns. Conversely, existing deep learning architectures attempt to learn these interactions by stacking many deep layers, but they often suffer from severe training instability and overfitting.

The article designs and evaluates a new model, the Neural Factorization Machine, to demonstrate that sparse predictive performance improves significantly by uniting the linear pairwise modeling of Factorization Machines with non-linear neural network layers.

The researchers developed a specialized Bilinear Interaction pooling layer that explicitly captures pairwise feature interactions in linear time without adding parameters, followed by non-linear hidden layers to learn higher-order patterns. They evaluated the model across two public regression benchmarks: the Frappe mobile app context dataset (comprising over 288,000 instances and 5,382 features) and the MovieLens tag recommendation dataset (comprising over 2 million instances and 90,445 features). The evaluation measured prediction error against standard Factorization Machines, higher-order Factorization Machines, and state-of-the-art deep learning architectures including Google’s Wide&Deep and Microsoft’s DeepCross.

The evaluation yielded several key findings. First, the Neural Factorization Machine achieved the lowest prediction error on both benchmarks, outperforming standard Factorization Machines by approximately 7.3% relative improvement with just a single non-linear hidden layer. Second, the proposed model outperformed deeper baselines while requiring far fewer parameters; for example, it attained lower error than the 10-layer DeepCross model, which struggled with severe overfitting. Third, the model demonstrated superior training stability and robustness, achieving peak accuracy from random parameter initialization without requiring the complex pre-training steps needed by other deep baselines. Finally, incorporating standard neural network regularization techniques directly on the interaction layer—specifically dropout and batch normalization—proved more effective at preventing overfitting and speeding up training convergence than traditional penalty methods.

These findings indicate that predictive modeling teams do not need excessively deep, computationally expensive architectures to achieve high accuracy on sparse web data. Instead, engineering a more informative low-level interaction layer reduces model complexity, lowering training costs, maintenance overhead, and deployment risks compared to heavily parameterized deep neural networks.

Organizations handling sparse prediction tasks should consider adopting this hybrid architecture over complex deep networks, testing single-layer non-linear configurations first before attempting deeper designs. Practitioners should also employ dropout on the interaction layer to optimize generalization. Before widespread commercial rollout, technical teams should conduct pilot studies in domain-specific workflows, such as search ranking and ad click-through rate prediction, and explore compression or hashing techniques for massive production scales. Confidence in these results is high across the evaluated recommendation tasks, though performance should still be validated when adapting the approach to dense data or complex sequential behaviors.

  • Paper: Factorization Machines, Steffen Rendle (2010). This seminal paper introduces Factorization Machines, providing the theoretical and mathematical foundation of second-order feature interactions that Neural Factorization Machines directly extend with non-linear neural components.
  • Paper: Entity Embeddings of Categorical Variables, Cheng Guo et al. (2016). This work establishes the entity embedding technique for mapping sparse categorical predictors into dense continuous vector spaces, a core prerequisite for neural models on tabular data like NFM.
Cover for Neural Factorization Machines for Sparse Predictive Analytics

Abstract

Many predictive tasks of web applications need to model categorical variables, such as user IDs and demographics like genders and occupations. To apply standard machine learning techniques, these categorical predictors are always converted to a set of binary features via one-hot encoding, making the resultant feature vector highly sparse. To learn from such sparse data effectively, it is crucial to account for the interactions between features.

Factorization Machines (FMs) are a popular solution for efficiently using the second-order feature interactions. However, FM models feature interactions in a linear way, which can be insufficient for capturing the non-linear and complex inherent structure of real-world data. While deep neural networks have recently been applied to learn non-linear feature interactions in industry, such as the Wide&Deep by Google and DeepCross by Microsoft, the deep structure meanwhile makes them difficult to train.

In this paper, we propose a novel model Neural Factorization Machine (NFM) for prediction under sparse settings. NFM seamlessly combines the linearity of FM in modelling second-order feature interactions and the non-linearity of neural network in modelling higher-order feature interactions. Conceptually, NFM is more expressive than FM since FM can be seen as a special case of NFM without hidden layers. Empirical results on two regression tasks show that with one hidden layer only, NFM significantly outperforms FM with a 7.3% relative improvement. Compared to the recent deep learning methods Wide&Deep and DeepCross, our NFM uses a shallower structure but offers better performance, being much easier to train and tune in practice.

Table of Contents

  • 1 Introduction
  • 2 Modelling Feature Interactions
  • 2.1 Factorization Machines
  • 2.1.1 Expressiveness Limitation of FM
  • 2.2 Deep Neural Networks
  • 2.2.1 Optimization Difficulties of DNN
  • 3 Neural Factorization Machines
  • 3.1 The NFM Model
  • 3.1.1 NFM Generalizes FM
  • 3.1.2 Relation to Wide&Deep and DeepCross
  • 3.1.3 Time Complexity Analysis
  • 3.2 Learning
  • 3.2.1 Dropout
  • 3.2.2 Batch Normalization
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.1.1 Datasets
  • 4.1.2 Evaluation Protocols
  • 4.1.3 Baselines
  • 4.1.4 Parameter Settings.
  • 4.2 Study of Bi-Interaction Pooling (RQ1)
  • 4.2.1 Dropout Improves Generalization
  • 4.2.2 Batch Normalization Speeds up Training
  • 4.3 Impact of Hidden Layers (RQ2)
  • 4.3.1 Pre-training Speeds up Training
  • 4.4 Performance Comparison (RQ3)
  • 5 Conclusion and Future Work
  • References

Knowls

  1. Knowl 1 — Neural Factorization Machine Architecture

    model/method

    The Neural Factorization Machine (NFM) is a predictive neural network architecture designed for sparse input vectors x∈Rn\mathbf{x} \in \mathbb{R}^n, where non-zero entries represent active discrete or categorical features. NFM computes a scalar prediction y^NFM(x)\hat{y}_{\text{NFM}}(\mathbf{x}) by combining a linear regression component with a non-linear feature interaction neural network f(x)f(\mathbf{x}):

    y^NFM(x)=w0+∑i=1nwixi+f(x)\hat{y}_{\text{NFM}}(\mathbf{x}) = w_0 + \sum_{i=1}^n w_i x_i + f(\mathbf{x})

    where w0∈Rw_0 \in \mathbb{R} is the global bias and wi∈Rw_i \in \mathbb{R} is the first-order weight for feature ii.

    The non-linear interaction network f(x)f(\mathbf{x}) consists of four sequential stages:

    1. Embedding Layer: Maps active features to dense embedding vectors Vx={xivi∣xi≠0}\mathcal{V}_{\mathbf{x}} = \{x_i \mathbf{v}_i \mid x_i \neq 0\}, where vi∈Rk\mathbf{v}_i \in \mathbb{R}^k is the latent embedding vector for feature ii scaled by its input value xix_i.
    2. Bi-Interaction Pooling Layer: Converts the set Vx\mathcal{V}_{\mathbf{x}} into a single kk-dimensional vector fBI(Vx)f_{\text{BI}}(\mathcal{V}_{\mathbf{x}}) capturing pairwise feature interactions.
    3. Hidden Layers: A multi-layer feedforward network with LL hidden layers that learns high-order non-linear feature interactions: z1=σ1(W1fBI(Vx)+b1)\mathbf{z}_1 = \sigma_1(\mathbf{W}_1 f_{\text{BI}}(\mathcal{V}_{\mathbf{x}}) + \mathbf{b}_1) zl=σl(Wlzl−1+bl)for l=2,…,L\mathbf{z}_l = \sigma_l(\mathbf{W}_l \mathbf{z}_{l-1} + \mathbf{b}_l) \quad \text{for } l = 2, \dots, L where Wl∈Rdl×dl−1\mathbf{W}_l \in \mathbb{R}^{d_l \times d_{l-1}} (with d0=kd_0 = k), bl∈Rdl\mathbf{b}_l \in \mathbb{R}^{d_l}, and σl\sigma_l is an activation function (such as ReLU, sigmoid, or tanh).
    4. Prediction Layer: Computes the interaction score using weight vector h∈RdL\mathbf{h} \in \mathbb{R}^{d_L}: f(x)=hTzLf(\mathbf{x}) = \mathbf{h}^T \mathbf{z}_L
  2. Knowl 2 — Bilinear Interaction Pooling Operation and Linear Time Formulation

    equation

    The Bilinear Interaction (Bi-Interaction) pooling operation aggregates a set of scaled embedding vectors Vx={xivi}i=1n\mathcal{V}_{\mathbf{x}} = \{x_i \mathbf{v}_i\}_{i=1}^n (where vi∈Rk\mathbf{v}_i \in \mathbb{R}^k and xi∈Rx_i \in \mathbb{R}) into a single kk-dimensional vector encoding all second-order interactions via element-wise vector products:

    fBI(Vx)=∑i=1n∑j=i+1n(xivi)⊙(xjvj)f_{\text{BI}}(\mathcal{V}_{\mathbf{x}}) = \sum_{i=1}^n \sum_{j=i+1}^n (x_i \mathbf{v}_i) \odot (x_j \mathbf{v}_j)

    where ⊙\odot represents the Hadamard (element-wise) product, defined as (u⊙v)f=ufvf(\mathbf{u} \odot \mathbf{v})_f = u_f v_f.

    Without introducing additional parameters, this operation can be evaluated in O(kNx)O(k N_x) time (where NxN_x is the number of non-zero entries in x\mathbf{x}) by rewriting it as:

    fBI(Vx)=12[(∑i=1nxivi)2−∑i=1n(xivi)2]f_{\text{BI}}(\mathcal{V}_{\mathbf{x}}) = \frac{1}{2} \left[ \left( \sum_{i=1}^n x_i \mathbf{v}_i \right)^2 - \sum_{i=1}^n (x_i \mathbf{v}_i)^2 \right]

    where v2≡v⊙v\mathbf{v}^2 \equiv \mathbf{v} \odot \mathbf{v}.

    For backpropagation in gradient-based optimization, the derivative of the Bi-Interaction pooling layer with respect to an embedding vector vi\mathbf{v}_i is:

    ∂fBI(Vx)∂vi=(∑j=1nxjvj)xi−xi2vi=∑j=1,j≠inxixjvj\frac{\partial f_{\text{BI}}(\mathcal{V}_{\mathbf{x}})}{\partial \mathbf{v}_i} = \left( \sum_{j=1}^n x_j \mathbf{v}_j \right) x_i - x_i^2 \mathbf{v}_i = \sum_{j=1, j \neq i}^n x_i x_j \mathbf{v}_j

  3. Knowl 3 — Subsumption and Equivalence of Factorization Machines in NFM

    theoretical result

    The standard second-order Factorization Machine (FM) is mathematically equivalent to a Neural Factorization Machine with zero hidden layers (L=0L = 0).

    When L=0L = 0, the NFM architecture (denoted NFM-0) directly projects the output of the Bi-Interaction pooling layer to the prediction score using weight vector h∈Rk\mathbf{h} \in \mathbb{R}^k:

    y^NFM-0(x)=w0+∑i=1nwixi+hT(∑i=1n∑j=i+1nxivi⊙xjvj)=w0+∑i=1nwixi+∑i=1n∑j=i+1n(∑f=1khfvifvjf)xixj\hat{y}_{\text{NFM-0}}(\mathbf{x}) = w_0 + \sum_{i=1}^n w_i x_i + \mathbf{h}^T \left( \sum_{i=1}^n \sum_{j=i+1}^n x_i \mathbf{v}_i \odot x_j \mathbf{v}_j \right) = w_0 + \sum_{i=1}^n w_i x_i + \sum_{i=1}^n \sum_{j=i+1}^n \left( \sum_{f=1}^k h_f v_{if} v_{jf} \right) x_i x_j

    Fixing h=(1,1,…,1)T\mathbf{h} = (1, 1, \dots, 1)^T exactly recovers the standard Factorization Machine:

    y^FM(x)=w0+∑i=1nwixi+∑i=1n∑j=i+1n⟨vi,vj⟩xixj\hat{y}_{\text{FM}}(\mathbf{x}) = w_0 + \sum_{i=1}^n w_i x_i + \sum_{i=1}^n \sum_{j=i+1}^n \langle \mathbf{v}_i, \mathbf{v}_j \rangle x_i x_j

    Allowing h\mathbf{h} to be trainable does not increase the expressiveness of NFM-0 over standard FM, as any scaling from h\mathbf{h} can be absorbed into the latent factor vectors vi\mathbf{v}_i.

  4. Knowl 4 — Dropout and Batch Normalization Regularization for NFM

    model/method

    To regularize Neural Factorization Machines (NFM) against overfitting and stabilize gradient training under sparse input conditions, dropout and batch normalization (BN) are applied across layers:

    1. Bi-Interaction Dropout: During training, a dropout ratio ρ∈[0,1)\rho \in [0, 1) is applied directly to the kk-dimensional pooled representation fBI(Vx)f_{\text{BI}}(\mathcal{V}_{\mathbf{x}}), randomly setting latent components to zero. This regularizes the feature embedding space and prevents co-adaptation among feature embeddings.
    2. Hidden Layer Dropout: Dropout is applied to the output of each fully connected hidden layer zl\mathbf{z}_l to prevent overfitting on higher-order non-linear feature combinations.
    3. Batch Normalization: To counter internal covariate shift caused by updating embedding parameters, batch normalization is applied to the output of the Bi-Interaction pooling layer and to the inputs of each subsequent hidden layer: BN(z)=γ⊙(z−μBσB)+β\text{BN}(\mathbf{z}) = \boldsymbol{\gamma} \odot \left( \frac{\mathbf{z} - \boldsymbol{\mu}_{\mathcal{B}}}{\boldsymbol{\sigma}_{\mathcal{B}}} \right) + \boldsymbol{\beta} where μB\boldsymbol{\mu}_{\mathcal{B}} and σB\boldsymbol{\sigma}_{\mathcal{B}} are the mini-batch mean and standard deviation vectors, and γ,β\boldsymbol{\gamma}, \boldsymbol{\beta} are learnable scale and shift parameter vectors.
  5. Knowl 5 — Computational Complexity of Neural Factorization Machines

    theoretical result

    Given a sparse input vector x∈Rn\mathbf{x} \in \mathbb{R}^n with NxN_x non-zero entries, embedding dimension kk, and LL hidden layers of dimensions d1,d2,…,dLd_1, d_2, \dots, d_L (with d0=kd_0 = k), the overall computational time complexity for evaluating the Neural Factorization Machine is:

    O(kNx+∑l=1Ldl−1dl)O\left(k N_x + \sum_{l=1}^L d_{l-1} d_l\right)

    • The linear component ∑i=1nwixi\sum_{i=1}^n w_i x_i requires O(Nx)O(N_x) operations.
    • The Bi-Interaction pooling layer requires O(kNx)O(k N_x) operations using its quadratic-free reformulation.
    • The ll-th fully connected layer requires O(dl−1dl)O(d_{l-1} d_l) operations for matrix-vector multiplication.
    • The linear prediction layer requires O(dL)O(d_L) operations for the dot product hTzL\mathbf{h}^T \mathbf{z}_L.

    This makes the asymptotic evaluation complexity of NFM equivalent to deep models like Wide&Deep and DeepCross, while avoiding quadratic interaction computation.

  6. Knowl 6 — Benchmark Performance Comparison of NFM Against Linear and Deep Baselines

    empirical result

    On two sparse benchmark datasets—Frappe (context-aware app recommendation with 288,609 instances and 5,382 features) and MovieLens (personalized tag recommendation with 2,006,859 instances and 90,445 features), framed as regression with targets {1,−1}\{1, -1\} and evaluated by Root Mean Square Error (RMSE)—NFM with 1 hidden layer outperforms Factorization Machines (LibFM), 3rd-order Factorization Machines (HOFM), Wide&Deep, and DeepCross across embedding sizes k∈{128,256}k \in \{128, 256\}.

    Frappe MovieLens
    Method k=128k=128 RMSE k=256k=256 RMSE k=128k=128 RMSE k=256k=256 RMSE
    LibFM 0.3437 0.3385 0.4793 0.4735
    HOFM 0.3405 0.3331 0.4752 0.4636
    WideDeep (no pre-train) 0.3621 0.3661 0.5323 0.5313
    WideDeep (FM pre-train) 0.3311 0.3246 0.4595 0.4512
    DeepCross (no pre-train) 0.4025 0.4071 0.5885 0.5907
    DeepCross (FM pre-train) 0.3388 0.3548 0.5084 0.5130
    NFM (1 hidden layer) 0.3127 0.3095 0.4557 0.4443

    NFM achieves lower test RMSE than all baselines on both datasets with statistical significance (p<0.01p < 0.01 on Frappe, p<0.05p < 0.05 on MovieLens compared to the best baseline). NFM requires significantly fewer parameters than Wide&Deep and DeepCross (e.g., 0.71M parameters for NFM vs. 2.66M for Wide&Deep and 4.47M for DeepCross at k=128k=128 on Frappe).

  7. Knowl 7 — Impact of Network Depth and Non-linear Activation on NFM

    empirical result

    Validation error analysis on Frappe and MovieLens demonstrates that NFM achieves its best performance with a single hidden layer (L=1L=1).

    Model Variant Frappe (RMSE) MovieLens (RMSE)
    NFM-0 (L=0L=0, linear FM equivalent) 0.3562 0.4901
    NFM-1 (L=1L=1 hidden layer) 0.3133 0.4646
    NFM-2 (L=2L=2 hidden layers) 0.3193 0.4681
    NFM-3 (L=3L=3 hidden layers) 0.3219 0.4752
    NFM-4 (L=4L=4 hidden layers) 0.3202 0.4703

    Key observations:

    1. Adding one non-linear hidden layer improves relative performance over NFM-0 by 12.0% on Frappe and 5.2% on MovieLens.
    2. When an identity activation function is used (making the hidden layer a linear transformation), performance does not improve over NFM-0, demonstrating that non-linear transformations are required to capture higher-order feature interactions.
    3. Stacking additional layers (L≥2L \ge 2) degrades RMSE. Because the Bi-Interaction pooling layer explicitly captures pairwise feature interactions at the lowest level, a shallow non-linear layer is sufficient to learn high-order interactions without deep stacking.
  8. Knowl 8 — Sensitivity of NFM to Parameter Initialization and Pre-training

    empirical result

    Empirical analysis of parameter pre-training reveals different behaviors between NFM and deep concatenation-based models:

    1. Pre-training for Concatenation Models: Models that concatenate raw embeddings (Wide&Deep and DeepCross) experience severe optimization difficulty when trained from scratch (e.g., DeepCross without pre-training gets 0.4025 RMSE on Frappe at k=128k=128, but drops to 0.3388 when initialized with FM embeddings).
    2. Pre-training for NFM: Initializing NFM embeddings with FM-learned weights speeds up convergence (matching 40 epochs of training from scratch within 5 epochs), but does not improve final converged RMSE. NFM trained from random initialization with batch normalization converges to equivalent or slightly better final accuracy than with FM pre-training.
  9. Knowl 9 — Effect of Bi-Interaction Dropout versus L2 Regularization

    empirical result

    On the linear NFM-0 model (which is mathematically equivalent to standard Factorization Machines), applying dropout to the Bi-Interaction pooling layer yields superior generalization compared to conventional L2L_2 weight regularization on feature embeddings:

    • On Frappe, tuning the dropout ratio on the Bi-Interaction layer to ρ=0.3\rho = 0.3 yields a validation RMSE of 0.3562, outperforming the best tuned L2L_2 regularization score of 0.3799.
    • Dropout operates by preventing complex co-adaptations between embedding dimensions and acts as an implicit ensemble over sub-networks, whereas L2L_2 regularization only scales down parameter magnitudes uniformly.
    • Introducing Batch Normalization (BN) to the Bi-Interaction output speeds up convergence (e.g., on Frappe, epoch 20 with BN achieves a lower training error than epoch 60 without BN).

Coverage note — None was omitted; all primary methodological contributions, mathematical formulations, equivalence proofs, regularizations, complexity analyses, and empirical findings are fully covered.

References

  1. 1.L. Baltrunas, K. Church, A. Karatzoglou, and N. Oliver. Frappe: Understanding the usage and perception of mobile app recommendations in-the-wild. CoRR, abs/1505.03014, 2015.
  2. 2.I. Bayer, X. He, B. Kanagal, and S. Rendle. A generic coordinate descent framework for learning from implicit feedback. In WWW, 2017.
  3. 3.M. Blondel, A. Fujino, N. Ueda, and M. Ishihata. Higher-order factorization machines. In NIPS, 2016.
  4. 4.M. Blondel, M. Ishihata, A. Fujino, and N. Ueda. Polynomial networks and factorization machines: New insights and efficient training algorithms. In ICML, 2016.
  5. 5.D. Cao, X. He, L. Nie, X. Wei, X. Hu, S. Wu, and T.-S. Chua. Cross-platform app recommendation by jointly modeling ratings and texts. ACM TOIS, 2017.
  6. 6.J. Chen, B. Sun, H. Li, H. Lu, and X.-S. Hua. Deep ctr prediction in display advertising. In MM, 2016.
  7. 7.J. Chen, H. Zhang, X. He, L. Nie, W. Liu, and T.-S. Chua. Attentive collaborative filtering: Multimedia recommendation with feature- and item-level attention. In SIGIR, 2017.
  8. 8.T. Chen, X. He, and M.-Y. Kan. Context-aware image tweets modelling and recommendation. In MM, 2016.
  9. 9.H.-T. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir, R. Anil, Z. Haque, L. Hong, V. Jain, X. Liu, and H. Shah. Wide & deep learning for recommender systems. In DLRS, 2016.
  10. 10.J. Duchi, E. Hazan, and Y. Singer. Adaptive subgradient methods for online learning and stochastic optimization. Journal of Machine Learning Research, 2011.
  11. 11.D. Erhan, Y. Bengio, A. Courville, P.-A. Manzagol, P. Vincent, and S. Bengio. Why does unsupervised pre-training help deep learning? Journal of Machine Learning Research, 2010.
  12. 12.M. Genzel and G. Kutyniok. A mathematical framework for feature selection from real-world data with non-linear observations. arXiv preprint arXiv:1608.08852, 2016.
  13. 13.F. M. Harper and J. A. Konstan. The movielens datasets: History and context. ACM Transactions on Interactive Intelligent Systems, 2015.
  14. 14.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016.
  15. 15.X. He, M. Gao, M.-Y. Kan, Y. Liu, and K. Sugiyama. Predicting the popularity of web 2.0 items based on user comments. In SIGIR, 2014.
  16. 16.X. He, L. Liao, H. Zhang, L. Nie, X. Hu, and T.-S. Chua. Neural collaborative filtering. In WWW, 2017.
  17. 17.X. He, H. Zhang, M.-Y. Kan, and T.-S. Chua. Fast matrix factorization for online recommendation with implicit feedback. In SIGIR, 2016.
  18. 18.L. Hong, A. S. Doumith, and B. D. Davison. Co-factorization machines: Modeling user interests and predicting individual decisions in twitter. In WSDM, 2013.
  19. 19.R. Hong, Y. Yang, M. Wang, and X.-S. Hua. Learning visual semantic relationships for efficient visual retrieval. IEEE Transactions on Big Data, 2015.
  20. 20.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015.
  21. 21.Y. Juan, Y. Zhuang, W.-S. Chin, and C.-J. Lin. Field-aware factorization machines for ctr prediction. In RecSys, 2016.
  22. 22.Y. Koren. Factorization meets the neighborhood: A multifaceted collaborative filtering model. In KDD, 2008.
  23. 23.A. Novikov, M. Trofimov, and I. Oseledets. Exponential machines. In ICLR Workshop, 2017.
  24. 24.R. J. Oentaryo, E.-P. Lim, J.-W. Low, D. Lo, and M. Finegold. Predicting response in mobile advertising with hierarchical importance-aware factorization machine. In WSDM, 2014.
  25. 25.F. Petroni, L. Del Corro, and R. Gemulla. Core: Context-aware open relation extraction with factorization machines. In EMNLP, 2015.
  26. 26.R. Qiang, F. Liang, and J. Yang. Exploiting ranking factorization machines for microblog retrieval. In CIKM, 2013.
  27. 27.S. Rendle. Factorization machines. In ICDM, 2010.
  28. 28.S. Rendle. Factorization machines with libfm. ACM Transactions on Intelligent Systems and Technology, 2012.
  29. 29.S. Rendle, C. Freudenthaler, Z. Gantner, and L. Schmidt-Thieme. Bpr: Bayesian personalized ranking from implicit feedback. In UAI, 2009.
  30. 30.S. Rendle, Z. Gantner, C. Freudenthaler, and L. Schmidt-Thieme. Fast context-aware recommendations with factorization machines. In SIGIR, 2011.
  31. 31.Y. Shan, T. R. Hoens, J. Jiao, H. Wang, D. Yu, and J. Mao. Deep crossing: Web-scale modeling without manually crafted combinatorial features. In KDD, 2016.
  32. 32.F. Shen, Y. Mu, Y. Yang, W. Liu, L. Liu, J. Song, and H. T. Shen. Classification by retrieval: Binarizing data and classifier. In SIGIR, 2017.
  33. 33.N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 2014.
  34. 34.M. Wang, W. Fu, S. Hao, D. Tao, and X. Wu. Scalable semi-supervised learning by efficient anchor graph regularization. IEEE Transaction on Knowledge and Data Engineering, 2016.
  35. 35.M. Wang, X. Liu, and X. Wu. Visual classification by l1-hypergraph modeling. IEEE Transaction on Knowledge and Data Engineering, 2015.
  36. 36.P. Wang, J. Guo, Y. Lan, J. Xu, S. Wan, and X. Cheng. Learning hierarchical representation model for nextbasket recommendation. In SIGIR, 2015.
  37. 37.X. Wang, X. He, L. Nie, and T.-S. Chua. Item silk road: Recommending items from information domains to social users. In SIGIR, 2017.
  38. 38.J. Xiao, H. Ye, X. He, H. Zhang, F. Wu, and T.-S. Chua. Attentional factorization machines: Learning the weight of feature interactions via attention networks. In IJCAI, 2017.
  39. 39.C. Xiong, J. Callan, and T.-Y. Liu. Learning to attend and to rank with word-entity duets. In SIGIR, 2017.
  40. 40.C. Zhang, G. Zhou, Q. Yuan, H. Zhuang, Y. Zheng, L. Kaplan, S. Wang, and J. Han. Geoburst: Real-time local event detection in geo-tagged tweet streams. In SIGIR, 2016.
  41. 41.H. Zhang, F. Shen, W. Liu, X. He, H. Luan, and T.-S. Chua. Discrete collaborative filtering. In SIGIR, 2016.
  42. 42.H. Zhang, M. Wang, R. Hong, and T.-S. Chua. Play and rewind: Optimizing binary representations of videos by self-supervised temporal hashing. In MM, 2016.
  43. 43.H. Zhang, Z.-J. Zha, Y. Yang, S. Yan, Y. Gao, and T.-S. Chua. Attribute-augmented semantic hierarchy: Towards bridging semantic gap and intention gap in image retrieval. In MM, 2013.
  44. 44.W. Zhang, T. Du, and J. Wang. Deep learning over multi-field categorical data. In ECIR, 2016.

Citation

MLA
He, X., and T.-S. Chua. “Neural Factorization Machines for Sparse Predictive Analytics”. arXiv, 2017, http://arxiv.org/abs/1708.05027v1.
APA
He, X., & Chua, T.-S. (2017). Neural Factorization Machines for Sparse Predictive Analytics. arXiv. http://arxiv.org/abs/1708.05027v1
Chicago
He, X., and T.-S. Chua. 2017. “Neural Factorization Machines for Sparse Predictive Analytics”. arXiv. http://arxiv.org/abs/1708.05027v1.
Harvard
He, X. and Chua, T.-S. (2017) “Neural Factorization Machines for Sparse Predictive Analytics”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1708.05027v1.
Vancouver
1. He X, Chua T-S (2017) Neural Factorization Machines for Sparse Predictive Analytics. arXiv

BibTeX

@article{he2017neural,
  title = {Neural Factorization Machines for Sparse Predictive Analytics},
  author = {He, Xiangnan and Chua, Tat-Seng},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1708.05027v1},
  eprint = {1708.05027}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF