Memory Networks

Jason WestonSumit ChopraAntoine Bordes

article2014ICLR1,884 citations

Introduces Memory Networks, a learning architecture that integrates explicit read-write memory with neural inference to enable multi-step reasoning and dynamic knowledge retrieval for question answering.

Listen

Standard machine learning models often struggle with tasks that require tracking long narratives, complex reasoning, and memorization because their internal memory is small and compressed into continuous vectors. The article introduces and evaluates "Memory Networks," a class of models designed to address this problem by combining machine learning inference with an explicit, compartmentalized long-term memory that can be read from and written to dynamically.

The article evaluates this framework across two core benchmarks: a large-scale factoid question-answering dataset containing 14 million knowledge-base statements, and a synthetic environment designed to test multi-step logical deduction, entity tracking, temporal reasoning, and adaptation to previously unseen vocabulary. The evaluated model processes text inputs, manages memory storage, scores relevant supporting facts iteratively, and generates single-word or multi-word textual responses.

The findings show that Memory Networks substantially outperform conventional recurrent neural networks. On complex multi-step reasoning tasks in the simulated world, the Memory Network achieved 99.9% to 100% accuracy, whereas standard recurrent models and long short-term memory networks scored below 30%. On the 14-million-statement dataset, the model achieved an F1 score of 0.82, surpassing existing baseline methods. Furthermore, implementing an embedding cluster hash reduced the candidate retrieval search space by approximately 80-fold with negligible accuracy loss. The system also successfully answered reasoning questions involving completely unfamiliar words by leveraging surrounding context.

These results demonstrate that separating long-term storage from the inference mechanism enables artificial intelligence systems to perform multi-hop reasoning over large knowledge repositories without losing critical context. This capability directly reduces computational bottlenecks and improves reliability in complex query-answering applications.

Future efforts should explore more sophisticated memory management mechanisms, broader evaluation on open-domain comprehension tasks, and extensions into weakly supervised training environments where intermediate supporting facts are not explicitly annotated. Readers should note that current multi-step validation relies heavily on structured simulations and fully supervised supporting labels; performance on unstructured, noisy, real-world narrative text may vary until further benchmarks are established.

arXiv: 1410.3916facebookarchive/MemNN
  • Paper: Long Short-Term Memory, Sepp Hochreiter et al. (1997). Introduces the Long Short-Term Memory architecture, establishing the foundational gated recurrent mechanisms whose limitations in explicit long-term storage motivated the design of external Memory Networks.
  • Paper: Natural Language Processing (almost) from Scratch, Ronan Collobert et al. (2011). Pioneers unified deep neural architectures and distributed vector representations for text understanding that directly inform the continuous embedding techniques used in Memory Networks.
  • Paper: Reasoning With Neural Tensor Networks for Knowledge Base Completion, Richard Socher et al. (2013). Demonstrates continuous relational representations for knowledge base completion, providing foundational techniques for reasoning over structured facts that Memory Networks adapt into dynamic long-term memory.
  • Paper: Semantic Parsing on Freebase from Question-Answer Pairs, Jonathan Berant et al. (2013). Establishes question-answering benchmarks and semantic querying over large knowledge bases, defining the core problem formulation targeted by Memory Networks.
  • Paper: A Convolutional Neural Network for Modelling Sentences, Nal Kalchbrenner et al. (2014). Develops convolutional neural models for mapping sentences into continuous semantic spaces, which underpins the sentence-level scoring components in Memory Networks.
  • Paper: Linguistic Regularities in Continuous Space Word Representations, Tomas Mikolov et al. (2013). Demonstrates that vector offsets in continuous embedding spaces capture semantic and syntactic regularities, enabling the continuous scoring functions used to query memory slots.
  • Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). Extends the original Memory Networks framework to an end-to-end differentiable architecture with continuous multi-hop attention, eliminating the need for strong supervision of supporting facts.
  • Paper: Relational recurrent neural networks, Adam Santoro et al. (2018). Generalizes memory-augmented architectures by introducing relational memory cores that allow stored memory slots to interact directly through self-attention.
  • Paper: A simple neural network module for relational reasoning, Adam Santoro et al. (2017). Develops plug-and-play Relation Networks that build on the multi-hop reasoning tasks formulated in Memory Networks to perform pairwise relational reasoning across modalities.
  • Paper: Gated Graph Sequence Neural Networks, Yujia Li et al. (2015). Builds upon the reasoning tasks evaluated in Memory Networks by applying gated graph neural networks to multi-step relational and sequential reasoning problems.
  • Paper: Teaching Machines to Read and Comprehend, Karl Moritz Hermann et al. (2015). Expands the synthetic question-answering and multi-sentence reasoning paradigms of Memory Networks into large-scale natural-language reading comprehension benchmarks.
  • Paper: REALM: Retrieval-Augmented Language Model Pre-Training, Kelvin Guu et al. (2020). Scales the concept of neural external memory retrieval to massive text corpora by integrating a learned retriever into language model pre-training.
  • Paper: Matching Networks for One Shot Learning, Oriol Vinyals et al. (2016). Adapts external associative memory and differentiable attention mechanisms from memory-augmented models to solve few-shot metric learning.
Cover for Memory Networks

Abstract

We describe a new class of learning models called memory networks. Memory networks reason with inference components combined with a long-term memory component; they learn how to use these jointly. The long-term memory can be read and written to, with the goal of using it for prediction. We investigate these models in the context of question answering (QA) where the long-term memory effectively acts as a (dynamic) knowledge base, and the output is a textual response. We evaluate them on a large-scale QA task, and a smaller, but more complex, toy task generated from a simulated world. In the latter, we show the reasoning power of such models by chaining multiple supporting sentences to answer questions that require understanding the intension of verbs.

Table of Contents

  • 1 Introduction
  • 2 Memory Networks
  • 3 A MemNN Implementation For Text
  • 3.1 Basic model
  • 3.2 Word Sequences as Input
  • 3.3 Efficient Memory via Hashing
  • 3.4 Modeling Write Time
  • 3.5 Modeling Previously Unseen Words
  • 3.6 Exact Matches and Unseen Words
  • 4 Related Work
  • 5 Experiments
  • 5.1 Large-scale QA
  • 5.2 Simulated World QA
  • 5.2.1 QA with Previously Unseen Words
  • 5.3 Combining simulated data and large-scale QA
  • 6 Conclusions and Future Work
  • References
  • A Simulation Data Generation
  • B Word Sequence Training
  • C Write Time Feature Training
  • D Word-sequence Learning Curve Experiments
  • E Sentence-level Experiments
  • F Multi-word Answer Setting Experiments

Knowls

  1. Knowl 1 — General Memory Network Architecture

    model/method

    A memory network is a machine learning framework that couples an explicit, addressable long-term memory array with inference components. The architecture consists of a memory m=(m1,m2,…,mN)\mathbf{m} = (m_1, m_2, \dots, m_N), where each mim_i is an indexed memory object (such as a string, dense vector, or sparse vector), and four primary operational components:

    1. II (Input Feature Map): Maps incoming raw input xx (such as an incoming character, word, sentence, image, or audio signal) to an internal feature representation I(x)I(x). It can perform preprocessing such as parsing, entity extraction, or embedding mapping.
    2. GG (Generalization): Updates existing memory slots given new input features: mi=G(mi,I(x),m)m_i = G(m_i, I(x), \mathbf{m}) for all ii. In its simplest form, GG writes I(x)I(x) to the next vacant slot mH(x)=I(x)m_{H(x)} = I(x) via a slot selection function HH, leaving older slots unchanged. More complex variants can compress, rewrite, organize (e.g., by entity or topic), or execute forgetting strategies on full memories.
    3. OO (Output Feature Map): Reads from memory given the input representation and current memory state to produce an internal output feature representation: o=O(I(x),m)o = O(I(x), \mathbf{m}). This module typically performs retrieval and multi-step reasoning over stored facts.
    4. RR (Response Decoder): Converts the output features oo into the final target response format r=R(o)r = R(o), such as generating natural language text via a language model, choosing a single word, or producing an action.

    During execution, the model processes an input xx through the sequential pipeline I→G→O→RI \rightarrow G \rightarrow O \rightarrow R. The memory m\mathbf{m} is dynamically updated during both training and testing phases, whereas the model parameters parameterizing I,G,O,RI, G, O, R are frozen during evaluation.

  2. Knowl 2 — Text Memory Neural Network with Multi-Hop Iterative Retrieval

    model/method

    In a Memory Neural Network (MemNN) applied to question answering on textual inputs, statements are placed sequentially into memory slots m1,m2,…,mNm_1, m_2, \dots, m_N. When presented with a question xx, the model extracts supporting evidence across kk reasoning hops using an iterative matching process within the output module OO, and subsequently decodes an answer via the response module RR.

    For a 2-hop configuration (k=2k = 2):

    1. First-Hop Retrieval: The module searches all memory slots to find the single most relevant statement mo1m_{o_1} matching the query xx:

    o1=O1(x,m)=arg⁡max⁡i=1,…,NsO(x,mi)o_1 = O_1(x, \mathbf{m}) = \arg\max_{i=1,\dots,N} s_O(x, m_i)

    where sOs_O is a scoring function computing the compatibility between text representations.

    1. Second-Hop Retrieval: Conditioned on both the original query xx and the first supporting memory mo1m_{o_1}, the module searches memory for a second supporting statement mo2m_{o_2}:

    o2=O2(x,m)=arg⁡max⁡i=1,…,NsO([x,mo1],mi)o_2 = O_2(x, \mathbf{m}) = \arg\max_{i=1,\dots,N} s_O([x, m_{o_1}], m_i)

    where [x,mo1][x, m_{o_1}] denotes the combined representation of the query and first retrieved fact.

    1. Response Formulation: The final output vector o=[x,mo1,mo2]o = [x, m_{o_1}, m_{o_2}] is passed to RR. For single-word answers, the model scores every word ww in the vocabulary WW using an answer scoring function sRs_R:

    r=arg⁡max⁡w∈WsR([x,mo1,mo2],w)r = \arg\max_{w \in W} s_R([x, m_{o_1}, m_{o_2}], w)

    Alternatively, RR can be parameterized as a recurrent neural network (RNN or LSTM) conditioned on the sequence [x,mo1,mo2][x, m_{o_1}, m_{o_2}] to generate multi-word natural language sentences.

  3. Knowl 3 — Bilinear Embedding Scoring Function and Feature Spaces for MemNN

    equation

    The scoring functions sOs_O (for memory retrieval) and sRs_R (for response selection) are parameterized as bilinear embedding models over bag-of-words feature mappings:

    s(x,y)=Φx(x)⊤U⊤UΦy(y)s(x, y) = \Phi_x(x)^\top U^\top U \Phi_y(y)

    where U∈Rn×DU \in \mathbb{R}^{n \times D} is a learned weight matrix mapping DD-dimensional sparse feature vectors into an nn-dimensional latent embedding space. Separate parameter matrices UOU_O and URU_R are used for sOs_O and sRs_R.

    To allow the model to distinguish between different roles of words, the feature dimensionality is partitioned into distinct sub-dictionaries (D=3∣W∣D = 3|W|, where ∣W∣|W| is vocabulary size):

    • One bag-of-words dictionary for candidate memories Φy(y)\Phi_y(y).
    • Two distinct bag-of-words dictionaries for Φx(x)\Phi_x(x) depending on whether the word originates from the initial query xx or from a previously retrieved supporting memory mo1m_{o_1}.

    To explicitly retain exact word matches despite the low embedding rank nn, the representation can be extended with matching indicator features (D=8∣W∣D = 8|W|) or scored additively:

    s(x,y)=Φx(x)⊤U⊤UΦy(y)+λΦx(x)⊤Φy(y)s(x, y) = \Phi_x(x)^\top U^\top U \Phi_y(y) + \lambda \Phi_x(x)^\top \Phi_y(y)

    where λ\lambda is a mixing hyperparameter and Φx(x)⊤Φy(y)\Phi_x(x)^\top \Phi_y(y) computes the direct bag-of-words dot product.

  4. Knowl 4 — Supervised Margin Ranking Objective for MemNNs

    equation

    Memory Neural Networks are trained in a fully supervised setting where target supporting sentence indices (mo1,mo2m_{o_1}, m_{o_2}) and ground-truth answer words rr are provided for each training question xx. The parameters UOU_O and URU_R are trained jointly using stochastic gradient descent (SGD) to minimize a sum of margin ranking hinge losses:

    L(x,r,mo1,mo2)=∑fˉ≠mo1max⁡(0,γ−sO(x,mo1)+sO(x,fˉ))+∑fˉ′≠mo2max⁡(0,γ−sO([x,mo1],mo2)+sO([x,mo1],fˉ′))+∑rˉ≠rmax⁡(0,γ−sR([x,mo1,mo2],r)+sR([x,mo1,mo2],rˉ))\mathcal{L}(x, r, m_{o_1}, m_{o_2}) = \sum_{\bar{f} \neq m_{o_1}} \max\left(0, \gamma - s_O(x, m_{o_1}) + s_O(x, \bar{f})\right) + \sum_{\bar{f}' \neq m_{o_2}} \max\left(0, \gamma - s_O([x, m_{o_1}], m_{o_2}) + s_O([x, m_{o_1}], \bar{f}')\right) + \sum_{\bar{r} \neq r} \max\left(0, \gamma - s_R([x, m_{o_1}, m_{o_2}], r) + s_R([x, m_{o_1}, m_{o_2}], \bar{r})\right)

    where γ>0\gamma > 0 denotes the margin parameter, and fˉ,fˉ′,rˉ\bar{f}, \bar{f}', \bar{r} denote incorrect candidate memories and answer labels respectively. At each step of stochastic gradient descent, negative candidate examples are sampled randomly rather than summed over the full dictionary and memory space.

  5. Knowl 5 — Relative Write-Time Modeling and Pairwise Tournament Retrieval

    model/method

    To handle temporal ordering in dynamic narratives (e.g., distinguishing where an actor is now versus earlier), the scoring function is adapted to score triples (x,y,y′)(x, y, y') indicating whether memory yy is preferred over memory y′y' given context xx:

    sOt(x,y,y′)=Φx(x)⊤UOt⊤UOt(Φy(y)−Φy(y′)+Φt(x,y,y′))s_{Ot}(x, y, y') = \Phi_x(x)^\top U_{Ot}^\top U_{Ot} \left( \Phi_y(y) - \Phi_y(y') + \Phi_t(x, y, y') \right)

    where Φt(x,y,y′)\Phi_t(x, y, y') provides three binary features indicating: (1) whether xx is older than yy, (2) whether xx is older than y′y', and (3) whether yy is older than y′y'. When selecting the second supporting memory, xx is replaced by the first supporting memory mo1m_{o_1} to encode relative ages between facts.

    Memory retrieval replaces the global argmax with a sequential tournament over all memory slots:

    function Ot(q, m)
        t = 1
        for i = 2 to N do
            if sOt(q, m_i, m_t) > 0 then
                t = i
            end if
        end for
        return t
    end function

    The corresponding hinge loss replaces the individual ranking terms with pairs of terms for both ordering permutations:

    Lt=∑fˉ≠mo1[max⁡(0,γ−sOt(x,mo1,fˉ))+max⁡(0,γ+sOt(x,fˉ,mo1))]\mathcal{L}_t = \sum_{\bar{f} \neq m_{o_1}} \left[ \max\left(0, \gamma - s_{Ot}(x, m_{o_1}, \bar{f})\right) + \max\left(0, \gamma + s_{Ot}(x, \bar{f}, m_{o_1})\right) \right]

  6. Knowl 6 — Handling Out-of-Vocabulary Words via Context Bags and Word Dropout

    model/method

    To handle unseen words (e.g., novel named entities) at test time without prior embeddings, each word ww is augmented with representations of its surrounding context. Specifically, the model maintains two bag-of-words distributions for each word: one for words occurring immediately in its left context and one for its right context.

    1. Feature Space Expansion: The feature representation dimension is expanded from D=3∣W∣D = 3|W| to D=5∣W∣D = 5|W| (adding ∣W∣|W| dimensions for left-context bags and ∣W∣|W| dimensions for right-context bags). When combined with exact-word matching indicator features, the feature dimensionality reaches D=8∣W∣D = 8|W|.
    2. Context Dropout Training: During training, a dropout rate of d%d\% is applied to word tokens: with probability dd, the model pretends an observed word has never been seen, omits its direct nn-dimensional word embedding, and represents the token solely using its surrounding context feature bags.

    This forces the model to generalize linguistic templates (such as actor-action-location structures) and enables correct multi-hop reasoning over novel entities at inference time.

  7. Knowl 7 — Sublinear Memory Retrieval via Word and Embedding Cluster Hashing

    model/method

    When scaling to very large memory stores (e.g., millions of facts), evaluating sO(x,mi)s_O(x, m_i) across all NN memory slots becomes computationally intractable. Memory Networks prune candidate slots using two hashing strategies:

    1. Word Hashing: The memory is partitioned into ∣W∣|W| buckets corresponding to vocabulary words. A memory mim_i is mapped into all buckets corresponding to the words it contains. At query time, only candidate memories sharing at least one word with the query I(x)I(x) are scored.
    2. Clustered Embedding Hashing: KK-means clustering is applied to the word embedding vectors (UO)i(U_O)_i to construct KK semantic clusters. Memory sentences and input queries are hashed into all cluster buckets corresponding to their constituent words. Because synonyms and semantically related terms cluster together in embedding space, memories that share meaning without exact lexical overlap are still retrieved for scoring.

    Setting KK enables control over the trade-off between retrieval speed and QA accuracy.

  8. Knowl 8 — Word Stream Input Segmentation for Continuous Text

    model/method

    When input text is supplied as an unsegmented stream of words rather than discrete pre-split sentences, a learned segmentation module seg(c)\text{seg}(c) predicts whether the accumulated sequence of words cc constitutes a complete statement to be written into a new memory slot:

    seg(c)=Wseg⊤USΦseg(c)\text{seg}(c) = W_{\text{seg}}^\top U_S \Phi_{\text{seg}}(c)

    where Φseg(c)\Phi_{\text{seg}}(c) maps the word sequence cc into a bag-of-words feature space, USU_S is an embedding projection matrix, and WsegW_{\text{seg}} is a linear classifier weight vector in embedding space. A sequence boundary is triggered whenever seg(c)>γ\text{seg}(c) > \gamma.

    The segmenter is trained on supervision derived from labeled supporting facts using a margin ranking hinge loss:

    Lseg=∑f∈Fmax⁡(0,γ−seg(f))+∑fˉ∈Fˉmax⁡(0,γ+seg(fˉ))\mathcal{L}_{\text{seg}} = \sum_{f \in \mathcal{F}} \max(0, \gamma - \text{seg}(f)) + \sum_{\bar{f} \in \bar{\mathcal{F}}} \max(0, \gamma + \text{seg}(\bar{f}))

    where F\mathcal{F} is the set of ground-truth complete supporting segments, and Fˉ\bar{\mathcal{F}} consists of incomplete or non-supporting word sequences.

  9. Knowl 9 — Empirical Performance on Simulated World QA Benchmark

    data/table

    Performance of Memory Networks compared to Recurrent Neural Networks (RNNs) and Long Short-Term Memory networks (LSTMs) was evaluated on a synthetic QA environment involving actors, objects, and locations across two temporal distance difficulties (Difficulty 1 and Difficulty 5):

    Difficulty 1 Difficulty 5
    Method actor w/o before actor actor+object actor actor+object
    RNN 100% 60.9% 27.9% 23.8% 17.8%
    LSTM 100% 64.8% 49.1% 35.2% 29.0%
    MemNN k=1k=1 97.8% 31.0% 24.0% 21.9% 18.5%
    MemNN k=1k=1 (+time) 99.9% 60.2% 42.5% 60.8% 44.4%
    MemNN k=2k=2 (+time) 100% 100% 100% 100% 99.9%

    Standard RNNs and LSTMs fail on harder multi-hop reasoning tasks (dropping to 17.8%17.8\% and 29.0%29.0\% on Difficulty 5 actor+object) due to memory compression degradation over long sequence distances. MemNN with single-hop retrieval (k=1k=1) fails because actor+object queries require chaining two facts (finding the actor who dropped the object, and finding where that actor went). Adding relative write-time features and 2-hop memory retrieval (k=2k=2) achieves nearly perfect accuracy (99.9%99.9\%–100%100\%) across all difficulties.

  10. Knowl 10 — Large-Scale Question Answering and Hashing Acceleration Results

    data/table

    The MemNN architecture was evaluated on a large-scale QA benchmark consisting of 14 million REVERB knowledge base triples extracted from ClueWeb09, trained using pseudo-labeled QA pairs and 35M WikiAnswers paraphrase pairs. Performance was measured by re-ranking candidate answers using F1F_1 score, along with retrieval candidate speedups:

    Method F1F_1
    Fader et al. (2013) 0.54
    Bordes et al. (2014b) 0.73
    MemNN (embedding only) 0.72
    MemNN (with BoW matching features) 0.82
    Method Embedding F1F_1 Embedding + BoW F1F_1 Candidates Scored (Speedup)
    MemNN (no hashing) 0.72 0.82 14M (1×1\times)
    MemNN (word hash) 0.63 0.68 13k (1000×1000\times)
    MemNN (cluster hash, K=1000K=1000) 0.71 0.80 177k (80×80\times)

    Adding exact bag-of-words matching features increased F1F_1 from 0.720.72 to 0.820.82. Cluster hashing (K=1000K=1000) achieved an ∼80×\sim 80\times reduction in evaluated candidates with minimal accuracy loss (0.82→0.800.82 \rightarrow 0.80), whereas lexical word hashing achieved 1000×1000\times speedup but degraded F1F_1 to 0.680.68 due to missing synonym matches.

  11. Knowl 11 — Natural Language Sentence Generation via Recurrent Decoder Conditioning

    empirical result

    When evaluated in a multi-word answer generation setting where answers must be complete natural language sentences (e.g., "Bill is in the kitchen"), the MemNN response module RR is replaced with an RNN or LSTM decoder conditioned on the retrieved memory features [x,mo1,mo2][x, m_{o_1}, m_{o_2}].

    On the Difficulty 5 actor+object task:

    • Conventional RNNs conditioned solely on the raw word stream achieved 13.97%13.97\% accuracy.
    • Conventional LSTMs conditioned solely on the raw word stream achieved 14.01%14.01\% accuracy.
    • MemNN with an RNN response module R([x,mo1,mo2])R([x, m_{o_1}, m_{o_2}]) achieved 68.83%68.83\% accuracy.
    • MemNN with an LSTM response module R([x,mo1,mo2])R([x, m_{o_1}, m_{o_2}]) achieved 90.98%90.98\% accuracy.

    Decoupling memory addressing and multi-hop fact retrieval from sequence generation allows the recurrent decoder to focus solely on surface realization, resolving long-term dependency bottlenecks.

Coverage note — None omitted; all core components, formulations (general framework, text MemNN, scoring functions, loss objectives, write-time features, word hashing, out-of-vocabulary handling, word stream segmentation), and primary empirical findings from the simulated and large-scale QA tasks are covered.

References

  1. 1.Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014.
  2. 2.Berant, Jonathan, Chou, Andrew, Frostig, Roy, and Liang, Percy. Semantic parsing on freebase from question-answer pairs. In EMNLP, pp. 1533–1544, 2013.
  3. 3.Berant, Jonathan, Srikumar, Vivek, Chen, Pei-Chun, Huang, Brad, Manning, Christopher D, Vander Linden, Abby, Harding, Brittany, and Clark, Peter. Modeling biological processes for reading comprehension. In Proc. EMNLP, 2014.
  4. 4.Bordes, Antoine, Usunier, Nicolas, Collobert, Ronan, and Weston, Jason. Towards understanding situated natural language. In AISTATS, 2010.
  5. 5.Bordes, Antoine, Chopra, Sumit, and Weston, Jason. Question answering with subgraph embeddings. In Proc. EMNLP, 2014a.
  6. 6.Bordes, Antoine, Weston, Jason, and Usunier, Nicolas. Open question answering with weakly supervised embedding models. ECML-PKDD, 2014b.
  7. 7.Das, Sreerupa, Giles, C Lee, and Sun, Guo-Zheng. Learning context-free grammars: Capabilities and limitations of a recurrent neural network with an external stack memory. In Proceedings of The Fourteenth Annual Conference of Cognitive Science Society. Indiana University, 1992.
  8. 8.Fader, Anthony, Zettlemoyer, Luke, and Etzioni, Oren. Paraphrase-driven learning for open question answering. In ACL, pp. 1608–1618, 2013.
  9. 9.Graves, Alex. Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850, 2013.
  10. 10.Graves, Alex, Wayne, Greg, and Danihelka, Ivo. Neural turing machines. arXiv preprint arXiv:1410.5401, 2014.
  11. 11.Haykin, Simon. Neural networks: A comprehensive foundation. 1994.
  12. 12.Hochreiter, Sepp and Schmidhuber, J¨urgen. Long short-term memory. Neural computation, 9(8): 1735–1780, 1997.
  13. 13.Iyyer, Mohit, Boyd-Graber, Jordan, Claudino, Leonardo, Socher, Richard, and III, Hal Daum´e. A neural network for factoid question answering over paragraphs. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 633–644, 2014.
  14. 14.Kolomiyets, Oleksandr and Moens, Marie-Francine. A survey on question answering technology from an information retrieval perspective. Information Sciences, 181(24):5412–5434, 2011.
  15. 15.Lin, Tsungnam, Horne, Bil G, Tiˇno, Peter, and Giles, C Lee. Learning long-term dependencies in narx recurrent neural networks. Neural Networks, IEEE Transactions on, 7(6):1329–1338, 1996.
  16. 16.Miikkulainen, Risto. {DISCERN}:{A} distributed artificial neural network model of script processing and memory. 1990.
  17. 17.Mikolov, Tomas, Karafi´at, Martin, Burget, Lukas, Cernock`y, Jan, and Khudanpur, Sanjeev. Recurrent neural network based language model. In Interspeech, pp. 1045–1048, 2010.
  18. 18.Richardson, Matthew, Burges, Christopher JC, and Renshaw, Erin. Mctest: A challenge dataset for the open-domain machine comprehension of text. In EMNLP, pp. 193–203, 2013.
  19. 19.Schmidhuber, J¨urgen. Learning to control fast-weight memories: An alternative to dynamic recurrent networks. Neural Computation, 4(1):131–139, 1992.
  20. 20.Schmidhuber, J¨urgen. A self-referentialweight matrix. In ICANN93, pp. 446–450. Springer, 1993.
  21. 21.Tolkien, John Ronald Reuel. The Fellowship of the Ring. George Allen & Unwin, 1954.
  22. 22.Weston, Jason, Bengio, Samy, and Usunier, Nicolas. Wsabie: Scaling up to large vocabulary image annotation. In Proceedings of the Twenty-Second international joint conference on Artificial Intelligence-Volume Volume Three, pp. 2764–2770. AAAI Press, 2011.
  23. 23.Yih, Wen-Tau, He, Xiaodong, and Meek, Christopher. Semantic parsing for single-relation question answering. In Proceedings of ACL. Association for Computational Linguistics, June 2014. URL http://research.microsoft.com/apps/pubs/default.aspx?id=214353.
  24. 24.Zaremba, Wojciech and Sutskever, Ilya. Learning to execute. arXiv preprint arXiv:1410.4615, 2014.

Citation

MLA
Weston, J., et al. “Memory Networks”. arXiv, 2014, https://doi.org/10.48550/arxiv.1410.3916.
APA
Weston, J., Chopra, S., & Bordes, A. (2014). Memory Networks. arXiv. https://doi.org/10.48550/arxiv.1410.3916
Chicago
Weston, J., S. Chopra, and A. Bordes. 2014. “Memory Networks”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.1410.3916.
Harvard
Weston, J., Chopra, S. and Bordes, A. (2014) “Memory Networks”. arXiv. Available at: https://doi.org/10.48550/arxiv.1410.3916.
Vancouver
1. Weston J, Chopra S, Bordes A (2014) Memory Networks. https://doi.org/10.48550/arxiv.1410.3916

BibTeX

@misc{https://doi.org/10.48550/arxiv.1410.3916,
  doi = {10.48550/ARXIV.1410.3916},
  url = {https://arxiv.org/abs/1410.3916},
  author = {Weston, Jason and Chopra, Sumit and Bordes, Antoine},
  keywords = {Artificial Intelligence (cs.AI), Computation and Language (cs.CL), Machine Learning (stat.ML), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Memory Networks},
  publisher = {arXiv},
  year = {2014},
  copyright = {Creative Commons Attribution Non Commercial Share Alike 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors