Bidirectional Attention Flow for Machine Comprehension

Minjoon SeoAniruddha KembhaviAli FarhadiHannaneh Hajishirzi

article2016ICLR2,146 citations

Introduces the Bi-Directional Attention Flow network to compute query-aware context representations without early summarization, setting a new performance standard for machine comprehension benchmarks like SQuAD.

Listen

Automated question answering and machine reading comprehension are critical capabilities for modern natural language processing systems, enabling software to extract accurate answers from unstructured text. Historically, neural models summarized long context passages into single fixed-size vectors or relied on complex step-by-step attention updates. These conventions often resulted in early information loss or allowed early prediction errors to cascade through the system.

The article introduces and evaluates the Bi-Directional Attention Flow (BIDAF) network, a hierarchical neural architecture designed to match questions and context paragraphs without premature summarization. The primary objective is to demonstrate that computing static, memory-less attention in two directions—from context to query and query to context—significantly improves question answering and cloze-style comprehension accuracy.

The research evaluated BIDAF across two standard benchmarks: the Stanford Question Answering Dataset (SQuAD), comprising over 100,000 question-context pairs derived from Wikipedia, and the massive CNN/DailyMail cloze-style test datasets, which contain hundreds of thousands of news articles with masked entity queries. The architecture computes representations across multiple granularities (character, word, and contextual levels) and uses a modular output layer that adapts to extract answer spans or fill in cloze-style missing entities.

The evaluation produced several key findings: First, an ensemble of the proposed BIDAF model achieved state-of-the-art results on the competitive SQuAD leaderboard, achieving an Exact Match score of 73.3% and an F1 score of 81.1%. Second, on the CNN and DailyMail benchmarks, single-model BIDAF outperformed all existing single models and even surpassed prior ensemble methods on the DailyMail test set with a 79.6% accuracy score. Third, ablation analyses confirmed that context-to-query attention is vital, as its removal caused an accuracy drop of over 10 points. Finally, static memory-less attention outperformed dynamic recurrent attention mechanisms by more than 3 points while remaining computationally simpler.

These results show that decoupling attention calculation from recurrent modeling layers improves overall system performance. This clear division of labor allows the attention layer to focus exclusively on query-context relationships, passing richer features to downstream layers without compounding historical attention errors. For organizations deploying search and automated document processing, this architectural pattern provides higher accuracy and faster, modular adaptability to different text-based tasks.

Based on these findings, technical teams developing reading comprehension and retrieval tools should adopt bi-directional, memory-less attention flow frameworks. The primary recommended next step is extending the architecture to multi-hop reasoning, enabling systems to iteratively link evidence across multiple sentences. A detailed error analysis indicates that half of the model's incorrect predictions stemmed from imprecise answer boundary selection, with another 28% caused by syntactic ambiguity. Consequently, future research and deployment pilots should focus on refining answer-boundary detection and evaluating performance on complex, multi-sentence reasoning before full operational rollout.

  • Paper: Teaching Machines to Read and Comprehend, Karl Moritz Hermann et al. (2015). This foundational paper establishes the Cloze-style CNN/Daily Mail reading comprehension benchmark and early neural attentive readers that BiDAF directly aims to evaluate against and improve upon.
  • Paper: Neural Machine Translation by Jointly Learning to Align and Translate, Dzmitry Bahdanau et al. (2015). It introduces the fundamental soft attention mechanism for sequence modeling that BiDAF adapts into a multi-stage bidirectional flow.
  • Paper: Effective Approaches to Attention-based Neural Machine Translation, Minh-Thang Luong et al. (2015). It systematizes global and local attention scoring functions across recurrent representations, providing the core mathematical framework for BiDAF's attention layers.
  • Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). It formalizes multi-hop, continuous attention reads over context representations without strong step supervision, motivating BiDAF's multi-stage context-query modeling.
  • Paper: A Convolutional Neural Network for Modelling Sentences, Nal Kalchbrenner et al. (2014). It introduces dynamic character- and word-level convolutions that inform the multi-granularity representation layers foundational to BiDAF's architecture.
  • Paper: Reading Wikipedia to Answer Open-Domain Questions, Danqi Chen et al. (2017). This work scales machine comprehension architectures like BiDAF from single paragraph reading to open-domain question answering over the entirety of Wikipedia using the DrQA framework.
  • Paper: A Structured Self-attentive Sentence Embedding, Zhouhan Lin et al. (2017). It extends non-recurrent multi-dimensional attention concepts by formulating structured self-attention matrix representations of sentence contexts.
  • Paper: Get To The Point: Summarization with Pointer-Generator Networks, Abigail See et al. (2017). It builds upon span-extraction and attention-tracking principles by combining pointer networks with coverage mechanisms to generate abstractive summaries.
  • Paper: Dense Passage Retrieval for Open-Domain Question Answering, Vladimir Karpukhin et al. (2020). It advances the open-domain question answering pipeline by replacing sparse paragraph search with dense passage retrieval to feed reader architectures.
  • Paper: Passage Re-ranking with BERT, Rodrigo Nogueira et al. (2019). It investigates full cross-attention transformer models to re-rank candidate passages for question answering, surpassing recurrent context-query interaction networks.
Cover for Bidirectional Attention Flow for Machine Comprehension

Abstract

Machine comprehension (MC), answering a query about a given context paragraph, requires modeling complex interactions between the context and the query. Recently, attention mechanisms have been successfully extended to MC. Typically these methods use attention to focus on a small portion of the context and summarize it with a fixed-size vector, couple attentions temporally, and/or often form a uni-directional attention. In this paper we introduce the Bi-Directional Attention Flow (BIDAF) network, a multi-stage hierarchical process that represents the context at different levels of granularity and uses bi-directional attention flow mechanism to obtain a query-aware context representation without early summarization. Our experimental evaluations show that our model achieves the state-of-the-art results in Stanford Question Answering Dataset (SQuAD) and CNN/DailyMail cloze test.

Table of Contents

  • 1 Introduction
  • 2 Model
  • 3 Related Work
  • 4 Question Answering Experiments
  • 5 Cloze Test Experiments
  • 6 Conclusion
  • References
  • A Error Analysis
  • B Variations of Similarity and Fusion Functions

Knowls

  1. Knowl 1 — Bi-Directional Attention Flow Network Architecture

    model/method

    The Bi-Directional Attention Flow (BiDAF) network is a multi-stage hierarchical neural architecture designed for machine comprehension and question answering over a context paragraph and a query. The model comprises six functional layers:

    1. Character Embedding Layer: Maps each word of the context and query into a vector space using character-level convolutional neural networks (CNNs).
    2. Word Embedding Layer: Maps each word to a distributed vector representation using pre-trained embeddings (e.g., GloVe).
    3. Contextual Embedding Layer: Refines token representations by encoding temporal interactions among surrounding words via bidirectional Long Short-Term Memory (BiLSTM) networks, applied independently to the context and query.
    4. Attention Flow Layer: Computes bidirectional cross-attention between context and query contextual vectors without collapsing or summarizing sequences into fixed-size vectors, outputting query-aware representations for every context position.
    5. Modeling Layer: Employs a multi-layer BiLSTM over the query-aware representations to capture dependencies among context words conditioned on the query.
    6. Output Layer: Task-specific layer that maps modeling states to output predictions, such as answer span boundary distributions for extractive reading comprehension.

    Unlike traditional attention mechanisms that summarize the context into a single vector or maintain dynamic recurrence across attention steps, BiDAF employs a memory-less attention mechanism where attention at each time step depends solely on the current context word and query representations.

  2. Knowl 2 — Bi-Directional Attention Flow Layer and Query-Aware Context Fusion

    model/method

    The Attention Flow layer links contextual representations of the context H∈R2d×TH \in \mathbb{R}^{2d \times T} (for a context sequence of length TT) and query U∈R2d×JU \in \mathbb{R}^{2d \times J} (for a query sequence of length JJ), where dd is the hidden dimension of the contextual LSTMs. It computes attention in two directions based on a shared similarity matrix S∈RT×JS \in \mathbb{R}^{T \times J}:

    Stj=α(H:t,U:j)=w(S)⊤[H:t;U:j;H:t∘U:j]S_{tj} = \alpha(H_{:t}, U_{:j}) = w_{(S)}^\top [H_{:t}; U_{:j}; H_{:t} \circ U_{:j}]

    where w(S)∈R6dw_{(S)} \in \mathbb{R}^{6d} is a trainable parameter vector, [;][;] denotes vector concatenation, and ∘\circ denotes elementwise multiplication.

    1. Context-to-Query (C2Q) Attention: Measures which query words are most relevant to each context token tt. The attention distribution over query words is:

    at=softmax(St:)∈RJa_t = \text{softmax}(S_{t:}) \in \mathbb{R}^J

    The attended query vector for context token tt is U~:t=∑j=1JatjU:j\tilde{U}_{:t} = \sum_{j=1}^J a_{tj} U_{:j}, yielding U~∈R2d×T\tilde{U} \in \mathbb{R}^{2d \times T}.

    1. Query-to-Context (Q2C) Attention: Identifies which context words are most similar to any query word. The attention weights over the context are:

    b=softmax(max⁡col(S))∈RTb = \text{softmax}\left(\max_{\text{col}}(S)\right) \in \mathbb{R}^T

    where max⁡col(S)\max_{\text{col}}(S) computes the maximum across query columns for each context row. The attended context representation is h~=∑t=1TbtH:t∈R2d\tilde{h} = \sum_{t=1}^T b_t H_{:t} \in \mathbb{R}^{2d}. This vector is replicated TT times across columns to produce H~∈R2d×T\tilde{H} \in \mathbb{R}^{2d \times T}.

    1. Query-Aware Representation Fusion: The contextual and attended representations are fused for each position tt via:

    G:t=β(H:t,U~:t,H~:t)=[H:t;U~:t;H:t∘U~:t;H:t∘H~:t]∈R8dG_{:t} = \beta(H_{:t}, \tilde{U}_{:t}, \tilde{H}_{:t}) = [H_{:t}; \tilde{U}_{:t}; H_{:t} \circ \tilde{U}_{:t}; H_{:t} \circ \tilde{H}_{:t}] \in \mathbb{R}^{8d}

    producing the combined query-aware context matrix G∈R8d×TG \in \mathbb{R}^{8d \times T}.

  3. Knowl 3 — Extractive Answer Span Prediction and Objective Function in BiDAF

    model/method

    In extractive question answering, the answer is a continuous span [k,l][k, l] within the context sequence of length TT. BiDAF models this prediction through a Modeling Layer and an Output Layer.

    The query-aware context matrix G∈R8d×TG \in \mathbb{R}^{8d \times T} is processed by a two-layer BiLSTM with output size dd per direction, producing M∈R2d×TM \in \mathbb{R}^{2d \times T}.

    The start index probability distribution p1∈RTp^1 \in \mathbb{R}^T over context positions is:

    p1=softmax(w(p1)⊤[G;M])p^1 = \text{softmax}(w_{(p1)}^\top [G; M])

    where w(p1)∈R10dw_{(p1)} \in \mathbb{R}^{10d} is a trainable weight vector.

    To predict the end index, MM is passed through an additional BiLSTM layer yielding M2∈R2d×TM^2 \in \mathbb{R}^{2d \times T}, and the end index probability distribution p2∈RTp^2 \in \mathbb{R}^T is:

    p2=softmax(w(p2)⊤[G;M2])p^2 = \text{softmax}(w_{(p2)}^\top [G; M^2])

    where w(p2)∈R10dw_{(p2)} \in \mathbb{R}^{10d} is a trainable weight vector.

    Training Objective: For a dataset of NN examples with ground truth start and end indices (yi1,yi2)(y_i^1, y_i^2), the network parameters θ\theta are trained by minimizing the cross-entropy loss:

    L(θ)=−1N∑i=1N(log⁡pyi11+log⁡pyi22)L(\theta) = -\frac{1}{N} \sum_{i=1}^N \left( \log p^1_{y_i^1} + \log p^2_{y_i^2} \right)

    Inference: At test time, the optimal answer span (k,l)(k, l) satisfies 1≤k≤l≤T1 \le k \le l \le T and maximizes the joint score pk1pl2p^1_k p^2_l, which is computed in O(T)O(T) time using dynamic programming.

  4. Knowl 4 — Character, Word, and Contextual Embedding Pipeline

    model/method

    BiDAF constructs word representations for context tokens {x1,…,xT}\{x_1, \dots, x_T\} and query tokens {q1,…,qJ}\{q_1, \dots, q_J\} across multiple granularities:

    1. Character Embeddings: Each character is mapped to a vector space. A 1D convolutional neural network (CNN) with 100 filters of width 5 operates across the character sequence of each word, followed by max-pooling across the entire word width to generate a fixed-size character representation per word.
    2. Word Embeddings: Pre-trained fixed 300-dimensional GloVe vectors are used for each token.
    3. Highway Network: For each token, the character-level CNN representation and the GloVe word embedding are concatenated and fed through a 2-layer Highway Network, outputting context matrix X∈Rd×TX \in \mathbb{R}^{d \times T} and query matrix Q∈Rd×JQ \in \mathbb{R}^{d \times J} (with d=100d = 100).
    4. Contextual BiLSTM: A bidirectional LSTM with output size dd per direction processes XX and QQ to capture temporal word dependencies, yielding context representations H∈R2d×TH \in \mathbb{R}^{2d \times T} and query representations U∈R2d×JU \in \mathbb{R}^{2d \times J} via concatenation of forward and backward hidden states.
  5. Knowl 5 — BiDAF Adaptation for Cloze-Style Comprehension

    model/method

    To apply BiDAF to cloze-style comprehension datasets (such as CNN and DailyMail) where the target answer is a single anonymized entity:

    1. Single-Index Prediction: Because answers are single words, the model only computes the start distribution p1p^1; the end distribution p2p^2 and its BiLSTM are omitted.
    2. Candidate Masking: All non-entity positions in the context are masked out prior to the final softmax layer.
    3. Multi-Occurrence Probability Pooling: If an entity candidate ee occurs multiple times in the context paragraph at index set Ie⊆{1,…,T}I_e \subseteq \{1, \dots, T\}, its aggregate probability is calculated by summing the predicted probabilities across all its occurrences:

    P(e)=∑t∈Iept1P(e) = \sum_{t \in I_e} p^1_t

    The loss minimized during training is the negative log probability of the ground-truth entity: −log⁡P(e∗)-\log P(e^*).

    1. Context Windowing: Each article is split into 19-word windows centered around each entity, and the RNNs do not propagate hidden states across sentence boundaries, allowing efficient parallelization.
  6. Knowl 6 — Performance on Stanford Question Answering Dataset (SQuAD)

    data/table

    BiDAF was evaluated on the SQuAD test set using Exact Match (EM) and F1 metrics against contemporary machine comprehension baselines.

    Single Model Ensemble
    Model EM F1 EM F1
    Logistic Regression Baseline 40.4 51.0 - -
    Dynamic Chunk Reader 62.5 71.0 - -
    Fine-Grained Gating 62.5 73.3 - -
    Match-LSTM 64.7 73.7 67.9 77.0
    Multi-Perspective Matching 65.5 75.1 68.2 77.2
    Dynamic Coattention Networks 66.2 75.9 71.6 80.4
    R-Net 68.4 77.5 72.1 79.7
    BiDAF (Ours) 68.0 77.3 73.3 81.1

    The 12-run BiDAF ensemble achieved an EM score of 73.3% and an F1 score of 81.1%, establishing state-of-the-art performance on the competitive SQuAD leaderboard.

  7. Knowl 7 — Ablation Study of BiDAF Components on SQuAD Development Set

    data/table

    An ablation study evaluated the impact of individual architectural components of BiDAF on the SQuAD development set.

    Model Variant EM F1
    No char embedding 65.0 75.4
    No word embedding 55.5 66.8
    No C2Q attention 57.2 67.7
    No Q2C attention 63.6 73.7
    Dynamic attention 63.5 73.6
    BiDAF (single) 67.7 77.3
    BiDAF (ensemble) 72.6 80.7

    Key observations:

    • Removing Context-to-Query (C2Q) attention caused the largest drop among attention ablations (>9.6>9.6 points in F1 and >10.5>10.5 points in EM).
    • Query-to-Context (Q2C) attention provided a substantial complementary gain (+3.6+3.6 F1).
    • Replacing the memory-less attention flow with dynamic attention (where attention is iteratively computed inside the modeling LSTM) reduced performance by 3.73.7 F1 points, confirming the advantage of pre-computed, memory-less attention flow.
    • Character embeddings provided crucial robustness for out-of-vocabulary and rare tokens (+1.9+1.9 F1).
  8. Knowl 8 — Performance on CNN and DailyMail Cloze-Style Benchmarks

    data/table

    BiDAF was tested on the CNN and DailyMail cloze datasets, where predictions are evaluated on accuracy across validation and test sets (* indicates ensemble models).

    CNN DailyMail
    Model val test val test
    Attentive Reader 61.6 63.0 70.5 69.0
    MemNN 63.4 66.8 - -
    AS Reader 68.6 69.5 75.0 73.9
    DER Network 71.3 72.9 - -
    Iterative Attention 72.6 73.3 - -
    EpiReader 73.4 74.0 - -
    Stanford AR 73.8 73.6 77.6 76.6
    GAReader 73.0 73.8 76.7 75.7
    AoA Reader 73.1 74.4 - -
    ReasoNet 72.9 74.7 77.6 76.6
    BiDAF (Ours) 76.3 76.9 80.3 79.6
    MemNN* 66.2 69.4 - -
    ASReader* 73.9 75.4 78.7 77.7
    Iterative Attention* 74.5 75.7 - -
    GA Reader* 76.4 77.4 79.1 78.1
    Stanford AR* 77.2 77.6 80.2 79.2

    A single-run BiDAF model achieved 76.9% on the CNN test set and 79.6% on the DailyMail test set, surpassing all prior single models on both datasets and exceeding prior ensemble models on the DailyMail test set.

  9. Knowl 9 — Ablation of Similarity and Fusion Function Formulations

    data/table

    The performance of BiDAF on the SQuAD development set was evaluated under alternative definitions of the similarity function α(h,u)\alpha(h, u) and the fusion function β(h,u~,h~)\beta(h, \tilde{u}, \tilde{h}):

    1. Dot product: α(h,u)=h⊤u\alpha(h, u) = h^\top u
    2. Linear: α(h,u)=wlin⊤[h;u]\alpha(h, u) = w_{\text{lin}}^\top [h; u]
    3. Bilinear: α(h,u)=h⊤Wbiu\alpha(h, u) = h^\top W_{\text{bi}} u
    4. Linear after MLP: α(h,u)=wlin⊤tanh⁡(Wmlp[h;u]+bmlp)\alpha(h, u) = w_{\text{lin}}^\top \tanh(W_{\text{mlp}}[h; u] + b_{\text{mlp}})
    5. MLP after concat fusion: β(h,u~,h~)=max⁡(0,Wmlp[h;u~;h∘u~;h∘h~]+bmlp)\beta(h, \tilde{u}, \tilde{h}) = \max(0, W_{\text{mlp}}[h; \tilde{u}; h \circ \tilde{u}; h \circ \tilde{h}] + b_{\text{mlp}})
    Function Variant EM F1
    Eqn. 1: dot product 65.5 75.5
    Eqn. 1: linear 59.5 69.7
    Eqn. 1: bilinear 61.6 71.8
    Eqn. 1: linear after MLP 66.2 76.4
    Eqn. 2: MLP after concat 67.1 77.0
    BiDAF (single, original) 68.0 77.3

    The formulation α(h,u)=w(S)⊤[h;u;h∘u]\alpha(h, u) = w_{(S)}^\top [h; u; h \circ u] without extra non-linearities in β\beta outperforms all other evaluated similarity and fusion variants.

  10. Knowl 10 — Error Analysis Breakdown for Extractive QA

    data/table

    A manual categorization of 50 randomly sampled Exact Match (EM) incorrect predictions on SQuAD revealed six primary failure modes:

    Error Type Ratio (%) Typical Cause
    Imprecise answer boundaries 50 Span partially matches but misses/includes extra boundary words
    Syntactic complications/ambiguities 28 Complex clausal structure or syntactic ambiguity in context/query
    Paraphrase problems 14 Lexical/semantic mismatches between query and text
    External knowledge 4 Question requires background facts not present in the text
    Multi-sentence reasoning 2 Synthesizing information across multiple context sentences
    Incorrect preprocessing 2 Tokenization errors in original data

    Half of all EM failures were due to slightly imprecise span boundary selection rather than a failure to locate the target answer region.

Coverage note — None was omitted; all key model architectures, equations, experimental benchmarks (SQuAD, CNN/DailyMail), ablations, and error analyses were extracted as self-contained knowls.

References

  1. 1.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In ICCV, 2015.
  2. 2.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. ICLR, 2015.
  3. 3.Danqi Chen, Jason Bolton, and Christopher D. Manning. A thorough examination of the cnn/daily mail reading comprehension task. In ACL, 2016.
  4. 4.Yiming Cui, Zhipeng Chen, Si Wei, Shijin Wang, Ting Liu, and Guoping Hu. Attention-over-attention neural networks for reading comprehension. arXiv preprint arXiv:1607.04423, 2016.
  5. 5.Bhuwan Dhingra, Hanxiao Liu, William W Cohen, and Ruslan Salakhutdinov. Gated-attention readers for text comprehension. arXiv preprint arXiv:1606.01549, 2016.
  6. 6.Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. Multimodal compact bilinear pooling for visual question answering and visual grounding. In EMNLP, 2016.
  7. 7.Karl Moritz Hermann, Tomas Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. Teaching machines to read and comprehend. In NIPS, 2015.
  8. 8.Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. The goldilocks principle: Reading children’s books with explicit memory representations. In ICLR, 2016.
  9. 9.Sepp Hochreiter and Jurgen Schmidhuber. Long short-term memory. Neural Computation, 1997.
  10. 10.Rudolf Kadlec, Martin Schmid, Ondrej Bajgar, and Jan Kleindienst. Text understanding with the attention sum reader network. In ACL, 2016.
  11. 11.Yoon Kim. Convolutional neural networks for sentence classification. In EMNLP, 2014.
  12. 12.Sosuke Kobayashi, Ran Tian, Naoaki Okazaki, and Kentaro Inui. Dynamic entity representation with max-pooling improves machine reading. In NAACL-HLT, 2016.
  13. 13.Kenton Lee, Tom Kwiatkowski, Ankur Parikh, and Dipanjan Das. Learning recurrent span representations for extractive question answering. arXiv preprint arXiv:1611.01436, 2016.
  14. 14.Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. Hierarchical question-image co-attention for visual question answering. In NIPS, 2016.
  15. 15.Mateusz Malinowski, Marcus Rohrbach, and Mario Fritz. Ask your neurons: A neural-based approach to answering questions about images. In ICCV, 2015.
  16. 16.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In EMNLP, 2014.
  17. 17.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100,000+ questions for machine comprehension of text. In EMNLP, 2016.
  18. 18.Matthew Richardson, Christopher JC Burges, and Erin Renshaw. Mctest: A challenge dataset for the open-domain machine comprehension of text. In EMNLP, 2013.
  19. 19.Yelong Shen, Po-Sen Huang, Jianfeng Gao, and Weizhu Chen. Reasonet: Learning to stop reading in machine comprehension. arXiv preprint arXiv:1609.05284, 2016.
  20. 20.Alessandro Sordoni, Phillip Bachman, and Yoshua Bengio. Iterative alternating neural attention for machine reading. arXiv preprint arXiv:1606.02245, 2016.
  21. 21.Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. JMLR, 2014.
  22. 22.Rupesh Kumar Srivastava, Klaus Greff, and Jurgen Schmidhuber. Highway networks. arXiv preprint arXiv:1505.00387, 2015.
  23. 23.Adam Trischler, Zheng Ye, Xingdi Yuan, and Kaheer Suleman. Natural language comprehension with the epireader. In EMNLP, 2016.
  24. 24.Shuohang Wang and Jing Jiang. Machine comprehension using match-lstm and answer pointer. arXiv preprint arXiv:1608.07905, 2016.
  25. 25.Jason Weston, Sumit Chopra, and Antoine Bordes. Memory networks. In ICLR, 2015.
  26. 26.Caiming Xiong, Stephen Merity, and Richard Socher. Dynamic memory networks for visual and textual question answering. In ICML, 2016a.
  27. 27.Caiming Xiong, Victor Zhong, and Richard Socher. Dynamic coattention networks for question answering. arXiv preprint arXiv:1611.01604, 2016b.
  28. 28.Huijuan Xu and Kate Saenko. Ask, attend and answer: Exploring question-guided spatial attention for visual question answering. In ECCV, 2016.
  29. 29.Zhilin Yang, Bhuwan Dhingra, Ye Yuan, Junjie Hu, William W Cohen, and Ruslan Salakhutdinov. Words or characters? fine-grained gating for reading comprehension. arXiv preprint arXiv:1611.01724, 2016.
  30. 30.Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. Stacked attention networks for image question answering. arXiv preprint arXiv:1511.02274, 2015.
  31. 31.Yang Yu, Wei Zhang, Kazi Hasan, Mo Yu, Bing Xiang, and Bowen Zhou. End-to-end reading comprehension with dynamic answer chunk ranking. arXiv preprint arXiv:1610.09996, 2016.
  32. 32.Matthew D Zeiler. Adadelta: an adaptive learning rate method. arXiv preprint arXiv:1212.5701, 2012.
  33. 33.Yuke Zhu, Oliver Groth, Michael S. Bernstein, and Li Fei-Fei. Visual7w: Grounded question answering in images. In CVPR, 2016.

Citation

MLA
Seo, M., et al. “Bidirectional Attention Flow for Machine Comprehension”. arXiv, 2016, http://arxiv.org/abs/1611.01603v6.
APA
Seo, M., Kembhavi, A., Farhadi, A., & Hajishirzi, H. (2016). Bidirectional Attention Flow for Machine Comprehension. arXiv. http://arxiv.org/abs/1611.01603v6
Chicago
Seo, M., A. Kembhavi, A. Farhadi, and H. Hajishirzi. 2016. “Bidirectional Attention Flow for Machine Comprehension”. arXiv. http://arxiv.org/abs/1611.01603v6.
Harvard
Seo, M. et al. (2016) “Bidirectional Attention Flow for Machine Comprehension”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1611.01603v6.
Vancouver
1. Seo M, Kembhavi A, Farhadi A, Hajishirzi H (2016) Bidirectional Attention Flow for Machine Comprehension. arXiv

BibTeX

@article{seo2016bidirectional,
  title = {Bidirectional Attention Flow for Machine Comprehension},
  author = {Seo, Minjoon and Kembhavi, Aniruddha and Farhadi, Ali and Hajishirzi, Hannaneh},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1611.01603v6},
  eprint = {1611.01603}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission