Grounded Multimodal Named Entity Recognition on Social Media

Jianfei YuZiyan LiJieming WangRui Xia

article2023ACL64 citations

Introduces the task of Grounded Multimodal Named Entity Recognition along with a benchmark Twitter dataset and a hierarchical index generation framework that jointly extracts text entities and locates their corresponding image regions to resolve visual ambiguity in social media posts.

Listen

Social media content is increasingly multimodal, combining text and imagery in ways that outpace human ability to manually review and categorize. While traditional Multimodal Named Entity Recognition identifies key entities in text using visual clues for context, it fails to link those entities to specific visual regions within the image. This limitation hinders the construction of complete multimodal knowledge graphs and makes entity disambiguation difficult when textual context alone is ambiguous.

To address this gap, the article introduces Grounded Multimodal Named Entity Recognition, a new task aimed at simultaneously identifying named entities in text, classifying their types, and locating their exact bounding-box groundings in the associated image. The article evaluates this task by creating a dedicated dataset and demonstrating an end-to-end model designed to extract multimodal entity triples in a single framework.

The authors constructed a benchmark dataset of 10,000 multimodal Twitter posts containing 16,778 entities, manually annotating visual bounding boxes for groundable mentions with high inter-annotator agreement. They extended several existing multimodal baseline systems using a two-stage pipeline that pairs traditional recognition with visual grounding. To overcome the pipeline's error propagation, they developed H-Index, a hierarchical index generation framework. H-Index uses a pre-trained sequence-to-sequence model to generate position indexes for entities, entity types, and groundability indicators, alongside an added visual output layer to predict exact image regions.

The experimental findings show that the proposed H-Index framework achieves an overall F1 score of 56.41%, outperforming the strongest baseline pipeline by 3.96 absolute percentage points. Across all tests, multimodal systems significantly outperformed text-only methods, confirming the value of visual grounding. In subtask evaluations, H-Index surpassed the best baseline on entity extraction and grounding by 5.44 percentage points, while performing competitively on the text-only recognition subtask. Ablation analyses confirmed that removing the hierarchical indicator prediction reduced overall F1 performance by 2.09 percentage points.

These results demonstrate that formulating grounded multimodal entity recognition as a unified, hierarchical index generation problem effectively avoids pipeline error propagation and enhances multimodal information extraction. For organizations relying on social media intelligence, knowledge graph construction, or content moderation, this unified approach offers improved accuracy and richer structural data extraction from complex multimedia posts.

Organizations developing multimodal artificial intelligence systems should consider adopting end-to-end generative index frameworks over traditional disjoint pipelines. Before full-scale deployment, practitioners should evaluate candidate region thresholds, as the study found grounding accuracy depends on selecting an appropriate number of proposed visual regions.

A key boundary condition of the study is that it focuses exclusively on visual grounding for entities explicitly mentioned in text, leaving unmentioned visual entities unannotated. While confidence in the reported improvements is high based on rigorous benchmark evaluations, users should note that the dataset reflects social media characteristics, meaning performance across other enterprise domains may require additional domain-specific data collection and tuning.

Yu et al (2023).pdf
Cover for Grounded Multimodal Named Entity Recognition on Social Media

Abstract

In recent years, Multimodal Named Entity Recognition (MNER) on social media has attracted considerable attention. However, existing MNER studies only extract entity-type pairs in text, which is useless for multimodal knowledge graph construction and insufficient for entity disambiguation. To solve these issues, in this work, we introduce a Grounded Multimodal Named Entity Recognition (GMNER) task. Given a text-image social post, GMNER aims to identify the named entities in text, their entity types, and their bounding box groundings in image (i.e., visual regions). To tackle the GMNER task, we construct a Twitter dataset based on two existing MNER datasets. Moreover, we extend four well-known MNER methods to establish a number of baseline systems and further propose a Hierarchical Index generation framework named H-Index, which generates the entity-type-region triples in a hierarchical manner with a sequence-to-sequence model. Experiment results on our annotated dataset demonstrate the superiority of our H-Index framework over baseline systems on the GMNER task. Our dataset annotation and source code are publicly released at https://github.com/NUSTM/GMNER.

Table of Contents

  • 1 Introduction
  • 2 Task Formulation
  • 3 Dataset
  • 4 Methodology
  • 4.1 Feature Extraction
  • 4.2 Design of Multimodal Indexes
  • 4.3 Index Generation Framework
  • 4.4 Entity Grounding
  • 4.5 Entity-Type-Region Triple Recovery
  • 5 Experiments
  • 5.1 Experimental Settings
  • 5.2 Baseline Systems
  • 5.3 Main Results
  • 5.4 In-Depth Analysis
  • 5.5 Case Study
  • 6 Related Work
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Grounded Multimodal Named Entity Recognition (GMNER) Task Formulation

    definition

    Given an input consisting of an nn-word text sequence s = (s_1, arscdots, s_n) and an accompanying image vv, the Grounded Multimodal Named Entity Recognition (GMNER) task requires extracting all multimodal entity triples:

    Y={(e1,t1,r1),…,(em,tm,rm)}Y = \{(e_1, t_1, r_1), \dots, (e_m, t_m, r_m)\}

    where for each triple (ei,ti,ri)(e_i, t_i, r_i):

    • eie_i is a named entity mention in the text ss.
    • tit_i is the semantic entity type belonging to a predefined set: {PER,LOC,ORG,MISC}\{\text{PER}, \text{LOC}, \text{ORG}, \text{MISC}\}.
    • rir_i is the visually grounded region in image vv. If entity eie_i has no grounded visual region in vv, ri=Noner_i = \text{None}; if entity eie_i is visually grounded, ri=(rix1,riy1,rix2,riy2)r_i = (r_i^{x_1}, r_i^{y_1}, r_i^{x_2}, r_i^{y_2}) is a 4-dimensional spatial coordinate vector specifying the top-left and bottom-right coordinates of the bounding box.
  2. Knowl 2 — Hierarchical Index (H-Index) Multimodal Input Representation and Index Serialization

    model/method

    The Hierarchical Index (H-Index) framework converts multimodal entity-type-region extraction into a unified position index generation problem using sequence-to-sequence modeling.

    Given text s=(s1,…,sn)s = (s_1, \dots, s_n) and image vv:

    1. Text Representation: The token sequence ss is embedded via the embedding matrix of a pretrained sequence-to-sequence model (BART) to yield text feature vectors T={e1,…,en}T = \{e_1, \dots, e_n\}, where each ei∈Rde_i \in \mathbb{R}^d.
    2. Visual Representation: Object detector VinVL extracts the top-KK candidate visual region features R={r1,…,rK}R = \{r_1, \dots, r_K\} with ri∈R2048r_i \in \mathbb{R}^{2048}, which are projected to text embedding dimension dd using a linear layer, producing visual region representations V={v1,…,vK}V = \{v_1, \dots, v_K\}, where vi∈Rdv_i \in \mathbb{R}^d.
    3. Unified Vocabulary of Multimodal Indexes:
      • Index 11: <groundable> indicator
      • Index 22: <ungroundable> indicator
      • Indexes 3,4,5,63, 4, 5, 6: <person>, <location>, <organization>, <miscellaneous> entity types
      • Indexes 77 to n+6n+6: word position indexes corresponding to words s1,…,sns_1, \dots, s_n.

    Each extracted entity is serialized into an index subsequence [p1,…,pl,u,t][p_1, \dots, p_l, u, t], where p1,…,plp_1, \dots, p_l are the text position indexes of the mention, u∈{1,2}u \in \{1, 2\} indicates groundability, and t∈{3,4,5,6}t \in \{3, 4, 5, 6\} indicates the entity type.

  3. Knowl 3 — H-Index Generative Model Architecture and Index Sequence Objective

    model/method

    The H-Index framework processes multimodal inputs using a BART encoder-decoder architecture.

    Encoder: The concatenation of text representations T∈Rn×dT \in \mathbb{R}^{n \times d} and visual representations V∈RK×dV \in \mathbb{R}^{K \times d} is encoded by the BART encoder:

    He=[HTe;HVe]=Encoder([T;V])H^e = [H_T^e; H_V^e] = \text{Encoder}([T; V])

    where HTe∈Rn×dH_T^e \in \mathbb{R}^{n \times d} and HVe∈RK×dH_V^e \in \mathbb{R}^{K \times d} are textual and visual hidden states, respectively.

    Decoder: At time step ii, the decoder hidden state hi∈Rdh_i \in \mathbb{R}^d is computed from encoder hidden representations HeH^e and previous target tokens y<iy_{<i}:

    hi=Decoder(He;y<i)h_i = \text{Decoder}(H^e; y_{<i})

    HˉTe=T+MLP(HTe)2\bar{H}_T^e = \frac{T + \text{MLP}(H_T^e)}{2}

    p(yi)=Softmax([C;HˉTe]⋅hi)p(y_i) = \text{Softmax}([C; \bar{H}_T^e] \cdot h_i)

    where C=TokenEmbed(c)C = \text{TokenEmbed}(c) denotes embeddings of the indicator indexes, entity type indexes, and special generation tokens (such as </s>).

    The index generation loss across NN samples of length MM is the cross-entropy objective:

    LT=−1NM∑j=1N∑i=1Mlog⁡p(yij)\mathcal{L}^T = -\frac{1}{NM} \sum_{j=1}^N \sum_{i=1}^M \log p(y_i^j)

  4. Knowl 4 — Visual Region Grounding Distribution and KL Divergence Objective

    equation

    For an entity predicted to be groundable at decoder step kk (where the generated token is the groundable indicator token with hidden state hkh_k), the distribution over the KK candidate visual regions from VinVL is computed as:

    HˉVe=V+MLP(HVe)2\bar{H}_V^e = \frac{V + \text{MLP}(H_V^e)}{2}

    p(zk)=Softmax(HˉVe⋅hk)p(z_k) = \text{Softmax}(\bar{H}_V^e \cdot h_k)

    Region Supervision: For each candidate visual region, the maximum Intersection over Union (IoU) with all ground-truth bounding boxes of that entity is calculated. Regions with IoU<0.5\text{IoU} < 0.5 have their IoU set to 00. Non-zero IoU values are normalized into a probability distribution g(zk)g(z_k).

    The visual grounding objective minimizes the Kullback-Leibler Divergence (KLD) loss:

    LV=1NE∑j=1N∑k=1Eg(zkj)log⁡g(zkj)p(zkj)\mathcal{L}^V = \frac{1}{NE} \sum_{j=1}^N \sum_{k=1}^E g(z_k^j) \log \frac{g(z_k^j)}{p(z_k^j)}

    where EE is the total number of groundable entities across NN training samples.

    The overall training loss of H-Index is:

    L=LT+LV\mathcal{L} = \mathcal{L}^T + \mathcal{L}^V

  5. Knowl 5 — Entity-Type-Region Triple Recovery Algorithm

    algorithm

    In the H-Index inference phase, greedy autoregressive decoding generates the predicted index sequence y^=[y^1,…,y^l]\hat{y} = [\hat{y}_1, \dots, \hat{y}_l], which is decoded into entity mention spans, groundability indicators, and entity types. For groundable entities, the predicted visual bounding box is chosen as the region maximizing p(z^k)p(\hat{z}_k).

    Input: Predicted index sequence y^=[y^1,…,y^l]\hat{y} = [\hat{y}_1, \dots, \hat{y}_l] where y^i∈[1,n+∣c∣]\hat{y}_i \in [1, n + |c|], and cc is the list containing indicator indexes, entity type indexes, and special tokens.
    Output: Set of entity-indicator-type triples EE
    E=∅E = \emptyset
    e=[]e = []
    i=1i = 1
    while i≤li \le l do
        yi=y^[i]y_i = \hat{y}[i]
        if yi≤∣c∣y_i \le |c| then
            if len(e)>0\text{len}(e) > 0 then
                if indexes in ee are in strictly ascending order then
                    if yi=1y_i = 1 or yi=2y_i = 2 then
                        E=E∪{(e,cyi,cyi+1)}E = E \cup \{(e, c_{y_i}, c_{y_{i+1}})\}
                    end if
                end if
            end if
            e=[]e = []
            i=i+2i = i + 2
        else
            e=e+[yi]e = e + [y_i]
            i=i+1i = i + 1
        end if
    end while
    return EE

    For each recovered tuple (ej,indicatorj,tj)(e_j, \text{indicator}_j, t_j) in EE:

    • If indicatorj=ungroundable\text{indicator}_j = \text{ungroundable} (index 2), the final predicted triple is (ej,tj,None)(e_j, t_j, \text{None}).
    • If indicatorj=groundable\text{indicator}_j = \text{groundable} (index 1), the visual region candidate rk=(rkx1,rky1,rkx2,rky2)r_k = (r_k^{x_1}, r_k^{y_1}, r_k^{x_2}, r_k^{y_2}) with the highest predicted probability arg⁡max⁡p(z^k)\arg\max p(\hat{z}_k) is assigned, giving (ek,tk,rk)(e_k, t_k, r_k).
  6. Knowl 6 — Evaluation Metrics for GMNER, MNER, and EEG

    definition

    Evaluation of GMNER requires evaluating exact string and type matching alongside spatial bounding box overlap.

    For a predicted triple (pe,pt,pr)(p_e, p_t, p_r) and gold triple (ge,gt,gr)(g_e, g_t, g_r):

    Ce={1,pe=ge0,otherwiseCt={1,pt=gt0,otherwiseC_e = \begin{cases} 1, & p_e = g_e \\ 0, & \text{otherwise} \end{cases} \qquad C_t = \begin{cases} 1, & p_t = g_t \\ 0, & \text{otherwise} \end{cases}

    Cr={1,pr=gr=None1,max⁡(IoU1,…,IoUj)>0.50,otherwiseC_r = \begin{cases} 1, & p_r = g_r = \text{None} \\ 1, & \max( \text{IoU}_1, \dots, \text{IoU}_j ) > 0.5 \\ 0, & \text{otherwise} \end{cases}

    where IoUj\text{IoU}_j is the Intersection over Union between the predicted bounding box prp_r and the jj-th ground-truth bounding box gr,jg_{r,j} associated with the gold entity.

    A predicted triple is counted as correct if and only if:

    correct={1,Ce=1 and Ct=1 and Cr=10,otherwise\text{correct} = \begin{cases} 1, & C_e = 1 \text{ and } C_t = 1 \text{ and } C_r = 1 \\ 0, & \text{otherwise} \end{cases}

    Precision (Pre.), Recall (Rec.), and F1 score are computed as:

    Pre.=#correct#predict,Rec.=#correct#gold,F1=2×Pre.×Rec.Pre.+Rec.\text{Pre.} = \frac{\text{\#correct}}{\text{\#predict}}, \quad \text{Rec.} = \frac{\text{\#correct}}{\text{\#gold}}, \quad \text{F1} = \frac{2 \times \text{Pre.} \times \text{Rec.}}{\text{Pre.} + \text{Rec.}}

    Two subtasks are evaluated with the same metrics:

    • Multimodal NER (MNER): Evaluates entity-type pairs (ei,ti)(e_i, t_i) via Ce∧CtC_e \wedge C_t.
    • Entity Extraction & Grounding (EEG): Evaluates entity-region pairs (ei,ri)(e_i, r_i) via Ce∧CrC_e \wedge C_r.
  7. Knowl 7 — Twitter-GMNER Dataset Statistics and Visual Grounding Distributions

    data/table

    The Twitter-GMNER dataset contains 10,000 multimodal tweet samples annotated with named entities, four entity types (PER, LOC, ORG, MISC), and visual bounding boxes annotated by three independent annotators with an inter-annotator agreement Fleiss' κ=0.84\kappa = 0.84.

    Split #Tweet #Entity #Groundable Entity #Box
    Train 7,000 11,782 4,694 5,680
    Dev 1,500 2,453 986 1,166
    Test 1,500 2,543 1,036 1,244
    Total 10,000 16,778 6,716 8,090

    Key distribution properties:

    • Approximately 60% of all entities (10,062 of 16,778) have no grounded bounding box in the accompanying image (r=Noner = \text{None}).
    • Image-level box counts: 44.2% of tweets contain 0 bounding boxes (unrelated image), 41.8% contain 1 box, 9.7% contain 2 boxes, 2.0% contain 3 boxes, 1.0% contain 4 boxes, and 1.2% contain 5 or more boxes.
    • Most PER entities are groundable in images, whereas LOC, ORG, and MISC entities (especially LOC) are predominantly ungroundable.
  8. Knowl 8 — Entity-Aware Visual Grounding (EVG) Pipeline Baseline Architecture

    model/method

    The Entity-aware Visual Grounding (EVG) pipeline baseline decouples GMNER into two sequential stages:

    1. Entity-Type Extraction: An existing MNER model extracts entity-type pairs (ei,ti)(e_i, t_i) from text ss.
    2. Visual Grounding with EVG:
      • Textual input formatting: [[CLS],s,[SEP],ei,[SEP],ti,[SEP]][[\text{CLS}], s, [\text{SEP}], e_i, [\text{SEP}], t_i, [\text{SEP}]] is processed by a pretrained BERTbase\text{BERT}_{\text{base}} model to yield text representations TT.
      • Candidate visual regions V={v1,…,vK}V = \{v_1, \dots, v_K\} are extracted via Faster R-CNN or VinVL.
      • Cross-Modal Transformer (CMT) models cross-modal interactions: H=CMT(V,T,T)={h1,…,hK}H = \text{CMT}(V, T, T) = \{h_1, \dots, h_K\}, where visual regions VV serve as queries and text TT serves as keys/values.
      • Region classification: For each visual region representation hj∈Rdh_j \in \mathbb{R}^d, a sigmoid layer outputs probability p(yj)=sigmoid(w⊤hj)p(y_j) = \text{sigmoid}(w^\top h_j), where w∈Rdw \in \mathbb{R}^d.
      • Inference decision: The region with highest probability is predicted as the visual grounding if its probability exceeds a tuned threshold; otherwise, the grounded region is predicted as None\text{None}.
  9. Knowl 9 — Empirical Comparison on GMNER, MNER, and EEG Benchmarks

    data/table

    Evaluation of text-only baselines (predicting None\text{None} for region grounding), pipeline multimodal systems, and the end-to-end H-Index model on the Twitter-GMNER test set.

    Methods GMNER MNER EEG
    Pre. Rec. F1 Pre. Rec. F1 Pre. Rec. F1
    Text
    HBiLSTM-CRF-None 43.56 40.69 42.07 78.80 72.61 75.58 49.17 45.92 47.49
    BERT-None 42.18 43.76 42.96 77.26 77.41 77.30 46.76 48.52 47.63
    BERT-CRF-None 42.73 44.88 43.78 77.23 78.64 77.93 46.92 49.28 48.07
    BARTNER-None 44.61 45.04 44.82 79.67 79.98 79.83 48.77 49.23 48.99
    Text+Image
    GVATT-RCNN-EVG 49.36 47.80 48.57 78.21 74.39 76.26 54.19 52.48 53.32
    UMT-RCNN-EVG 49.16 51.48 50.29 77.89 79.28 78.58 53.55 56.08 54.78
    UMT-VinVL-EVG 50.15 52.52 51.31 77.89 79.28 78.58 54.35 56.91 55.60
    UMGF-VinVL-EVG 51.62 51.72 51.67 79.02 78.64 78.83 55.68 55.80 55.74
    ITA-VinVL-EVG 52.37 50.77 51.56 80.40 78.37 79.37 56.57 54.84 55.69
    BARTMNER-VinVL-EVG 52.47 52.43 52.45 80.65 80.14 80.39 55.68 55.63 55.66
    H-Index (Ours) 56.16 56.67 56.41 79.37 80.10 79.73 60.90 61.46 61.18

    H-Index achieves 56.41% F1 on GMNER, outperforming the best pipeline baseline (BARTMNER-VinVL-EVG at 52.45%) by 3.96 percentage points, and outperforms all baselines on Entity Extraction & Grounding (EEG) by 5.44 percentage points in F1 (61.18% vs 55.74%).

  10. Knowl 10 — Ablation and Candidate Region Sensitivity Analysis of H-Index

    data/table

    Ablation experiments on the Twitter-GMNER test set demonstrate the contributions of the KLD grounding loss and the hierarchical prediction mechanism.

    Methods Pre. Rec. F1
    H-Index 56.16 56.67 56.41
    - replace KLD Loss with CE Loss 55.88 53.72 54.78
    - w/o Hierarchical Prediction 55.83 52.89 54.32

    Key observations:

    • Replacing Kullback-Leibler Divergence (KLD) loss with Cross-Entropy (CE) loss reduces F1 from 56.41% to 54.78% (-1.63 points), indicating that soft IoU supervision across visual regions captures inter-region relations better than hard binary labels.
    • Removing the hierarchical structure (predicting groundability indicators prior to visual region grounding and using a flat binary classification token instead) causes a 2.09 percentage point F1 drop to 54.32%.
    • Sensitivity analysis of top-KK VinVL proposals across K∈[2,20]K \in [2, 20] shows that performance on GMNER and EEG steadily improves as KK increases up to K=18K=18, where both reach optimal F1, because small KK values fail to cover ground-truth visual regions.
  11. Knowl 11 — Limitations of the GMNER Formulation and Approaches

    limitation

    The GMNER task formulation and current modeling approaches have two main limitations:

    1. Text-Conditioned Grounding Scope: The task only identifies visual regions corresponding to entities explicitly mentioned in the textual post. Real-world visual entities present in the image but unmentioned in text are not recognized or grounded.
    2. Dependence on Off-the-shelf Base Components: The baseline and H-Index architectures rely on general off-the-shelf components (BART, VinVL) without cross-modal pretraining tailored specifically for joint entity-type-region generation.

Coverage note — No substantial contributed material was omitted. Qualitative case study examples were excluded in favor of comprehensive quantitative ablation, sensitivity, and baseline comparison tables.

References

  1. 1.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of CVPR, pages 6077–6086.
  2. 2.Timothy Baldwin, Marie-Catherine de Marneffe, Bo Han, Young-Bum Kim, Alan Ritter, and Wei Xu. 2015. Shared tasks of the 2015 workshop on noisy user-generated text: Twitter lexical normalization and named entity recognition. In Proceedings of the Workshop on Noisy User-generated Text, pages 126–135.
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer.
  4. 4.Long Chen, Wenbo Ma, Jun Xiao, Hanwang Zhang, and Shih-Fu Chang. 2021a. Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 1036–1044.
  5. 5.Shuguang Chen, Gustavo Aguilar, Leonardo Neves, and Thamar Solorio. 2021b. Can images help recognize entities? a study of the role of images for multimodal ner. In Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021), pages 87–96.
  6. 6.Xiang Chen, Ningyu Zhang, Lei Li, Shumin Deng, Chuanqi Tan, Changliang Xu, Fei Huang, Luo Si, and Huajun Chen. 2022a. Hybrid transformer with multi-level fusion for multimodal knowledge graph completion. In Proceedings of SIGIR '22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 904–915.
  7. 7.Xiang Chen, Ningyu Zhang, Lei Li, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. 2022b. Good visual guidance make a better extractor: Hierarchical visual prefix for multimodal entity and relation extraction. In Findings of the Association for Computational Linguistics: NAACL 2022.
  8. 8.Jason PC Chiu and Eric Nichols. 2016. Named entity recognition with bidirectional lstm-cnns. Transactions of the Association for Computational Linguistics, 4:357–370.
  9. 9.Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li. 2021. Transvg: End-to-end visual grounding with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1769–1779.
  10. 10.Leon Derczynski, Eric Nichols, Marieke van Erp, and Nut Limsopatham. 2017. Results of the wnut2017 shared task on novel and emerging entity recognition. In Proceedings of the 3rd Workshop on Noisy User-generated Text, pages 140–147.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL, pages 4171–4186.
  12. 12.Jenny Rose Finkel, Trond Grenager, and Christopher Manning. 2005. Incorporating non-local information into information extraction systems by gibbs sampling. In Proceedings of ACL.
  13. 13.Joseph L Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychological bulletin, 76(5):378.
  14. 14.Michel Naim Gerguis, Cherif Salama, and M Watheq El-Kharashi. 2016. Asu: An experimental study on applying deep learning in twitter named entity recognition. In Proceedings of the 2nd Workshop on Noisy User-generated Text (WNUT), pages 188–196.
  15. 15.Kevin Gimpel, Nathan Schneider, Brendan O'Connor, Dipanjan Das, Daniel Mills, Jacob Eisenstein, Michael Heilman, Dani Yogatama, Jeffrey Flanigan, and Noah A Smith. 2010. Part-of-speech tagging for twitter: Annotation, features, and experiments. Technical report, Carnegie-Mellon Univ Pittsburgh Pa School of Computer Science.
  16. 16.Meihuizi Jia, Lei Shen, Xin Shen, Lejian Liao, Meng Chen, Xiaodong He, Zhendong Chen, and Jiaqi Li. 2023. Mner-qg: An end-to-end mrc framework for multimodal named entity recognition with query grounding. In Proceedings of AAAI.
  17. 17.Meihuizi Jia, Xin Shen, Lei Shen, Jinhui Pang, Lejian Liao, Yang Song, Meng Chen, and Xiaodong He. 2022. Query prior matters: A mrc framework for multimodal named entity recognition. In Proceedings of ACM MM, pages 3549–3558.
  18. 18.Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016. Neural architectures for named entity recognition. In Proceedings of NAACL-HLT.
  19. 19.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of ACL, pages 7871–7880.
  20. 20.Jing Li, Aixin Sun, Jianglei Han, and Chenliang Li. 2020a. A survey on deep learning for named entity recognition. IEEE Transactions on Knowledge and Data Engineering, 34(1):50–70.
  21. 21.Xiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han, Fei Wu, and Jiwei Li. 2020b. A unified mrc framework for named entity recognition. In Proceedings of ACL, pages 5849–5859.
  22. 22.Nut Limsopatham and Nigel Collier. 2016. Bidirectional LSTM for named entity recognition in twitter messages. In Proceedings of the 2nd Workshop on Noisy User-generated Text (WNUT).
  23. 23.Bill Yuchen Lin, Frank F Xu, Zhiyi Luo, and Kenny Zhu. 2017. Multi-channel bilstm-crf model for emerging named entity recognition in social media. In Proceedings of the 3rd Workshop on Noisy User-generated Text.
  24. 24.Ye Liu, Hui Li, Alberto Garcia-Duran, Mathias Niepert, Daniel Onoro-Rubio, and David S Rosenblum. 2019. Mmkg: multi-modal knowledge graphs. In The Semantic Web: 16th International Conference, ESWC 2019, pages 459–474.
  25. 25.Di Lu, Leonardo Neves, Vitor Carvalho, Ning Zhang, and Heng Ji. 2018. Visual attention model for name tagging in multimodal social media. In Proceedings of ACL, pages 1990–1999.
  26. 26.Xuezhe Ma and Eduard Hovy. 2016. End-to-end sequence labeling via bi-directional lstm-cnns-crf. In Proceedings of ACL.
  27. 27.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. 2016. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 11–20.
  28. 28.Seungwhan Moon, Leonardo Neves, and Vitor Carvalho. 2018. Multimodal named entity recognition for short social media posts. In Proceedings of NAACL.
  29. 29.Lev Ratinov and Dan Roth. 2009. Design challenges and misconceptions in named entity recognition. In Proceedings of CoNLL.
  30. 30.Joseph Redmon and Ali Farhadi. 2018. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767.
  31. 31.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28.
  32. 32.Alan Ritter, Sam Clark, Mausam, and Oren Etzioni. 2011. Named entity recognition in tweets: An experimental study. In Proceedings of EMNLP, pages 1524–1534.
  33. 33.Benjamin Strauss, Bethany Toma, Alan Ritter, Marie-Catherine De Marneffe, and Wei Xu. 2016. Results of the wnut16 named entity recognition shared task. In Proceedings of the 2nd Workshop on Noisy User-generated Text (WNUT), pages 138–144.
  34. 34.Chanchal Suman, Saichethan Miriyala Reddy, Sriparna Saha, and Pushpak Bhattacharyya. 2021. Why pay more? a simple and efficient named entity recognition system for tweets. Expert Systems with Applications, 167:114101.
  35. 35.Lin Sun, Jiquan Wang, Kai Zhang, Yindu Su, and Fangsheng Weng. 2021. Rpbert: a text-image relation propagation-based bert model for multimodal ner. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 13860–13868.
  36. 36.Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of ACL, pages 6558—-6569.
  37. 37.Xinyu Wang, Min Gui, Yong Jiang, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, and Kewei Tu. 2022a. ITA: Image-text alignments for multi-modal named entity recognition. In Proceedings of NAACL, pages 3176–3189.
  38. 38.Xuwu Wang, Junfeng Tian, Min Gui, Zhixu Li, Rui Wang, Ming Yan, Lihan Chen, and Yanghua Xiao. 2022b. Wikidiverse: A multimodal entity linking dataset with diversified contextual topics and entity types. In Proceedings of ACL, pages 4785–4797.
  39. 39.Xuwu Wang, Junfeng Tian, Min Gui, Zhixu Li, Jiabo Ye, Ming Yan, and Yanghua Xiao. 2022c. : Prompt-based entity-related visual clue extraction and integration for multimodal named entity recognition. In International Conference on Database Systems for Advanced Applications, pages 297–305.
  40. 40.Bo Xu, Shizhou Huang, Chaofeng Sha, and Hongya Wang. 2022. Maf: A general matching and alignment framework for multimodal named entity recognition. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining, pages 1215–1223.
  41. 41.Hang Yan, Tao Gui, Junqi Dai, Qipeng Guo, Zheng Zhang, and Xipeng Qiu. 2021. A unified generative framework for various ner subtasks. In Proceedings of ACL-IJCNLP, pages 5808–5822.
  42. 42.Sibei Yang, Guanbin Li, and Yizhou Yu. 2020. Graph-structured referring expression reasoning in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9952–9961.
  43. 43.Zhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang, Dong Yu, and Jiebo Luo. 2019. A fast and accurate one-stage approach to visual grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4683–4693.
  44. 44.Jiabo Ye, Junfeng Tian, Ming Yan, Xiaoshan Yang, Xuwu Wang, Ji Zhang, Liang He, and Xin Lin. 2022. Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15502–15512.
  45. 45.Jianfei Yu, Jing Jiang, Li Yang, and Rui Xia. 2020. Improving multimodal named entity recognition via entity span detection with unified multimodal transformer. In Proceedings of ACL, pages 3342–3352.
  46. 46.Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. 2018a. Mattnet: Modular attention network for referring expression comprehension. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1307–1315.
  47. 47.Zhou Yu, Jun Yu, Chenchao Xiang, Zhou Zhao, Qi Tian, and Dacheng Tao. 2018b. Rethinking diversified and discriminative proposal generation for visual grounding. In Proceedings of IJCAI, pages 1114–1120.
  48. 48.Dong Zhang, Suzhong Wei, Shoushan Li, Hanqian Wu, Qiaoming Zhu, and Guodong Zhou. 2021a. Multi-modal graph fusion for named entity recognition with targeted visual guidance. In Proceedings of AAAI, pages 14347–14355.
  49. 49.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. 2021b. Vinvl: Revisiting visual representations in vision-language models. In Proceedings of CVPR, pages 5579–5588.
  50. 50.Qi Zhang, Jinlan Fu, Xiaoyu Liu, and Xuanjing Huang. 2018. Adaptive co-attention network for named entity recognition in tweets. In Proceedings of AAAI, pages 5674–5681.
  51. 51.Fei Zhao, Chunhui Li, Zhen Wu, Shangyu Xing, and Xinyu Dai. 2022. Learning from different text-image pairs: A relation-enhanced graph convolutional network for multimodal ner. In Proceedings of the 30th ACM International Conference on Multimedia, pages 3983–3992.
  52. 52.Changmeng Zheng, Zhiwei Wu, Tao Wang, Yi Cai, and Qing Li. 2020. Object-aware multimodal named entity recognition in social media posts with adversarial learning. IEEE Transactions on Multimedia, 23:2520–2532.

Citation

MLA
Yu, J., et al. “Grounded Multimodal Named Entity Recognition on Social Media”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 9141–54, https://doi.org/10.18653/v1/2023.acl-long.508.
APA
Yu, J., Li, Z., Wang, J., & Xia, R. (2023). Grounded Multimodal Named Entity Recognition on Social Media. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9141–9154. https://doi.org/10.18653/v1/2023.acl-long.508
Chicago
Yu, J., Z. Li, J. Wang, and R. Xia. 2023. “Grounded Multimodal Named Entity Recognition on Social Media”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9141–54. https://doi.org/10.18653/v1/2023.acl-long.508.
Harvard
Yu, J. et al. (2023) “Grounded Multimodal Named Entity Recognition on Social Media”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 9141–9154. Available at: https://doi.org/10.18653/v1/2023.acl-long.508.
Vancouver
1. Yu J, Li Z, Wang J, Xia R (2023) Grounded Multimodal Named Entity Recognition on Social Media. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 9141–9154

BibTeX

@inproceedings{yu-etal-2023-grounded,
    title = "Grounded Multimodal Named Entity Recognition on Social Media",
    author = "Yu, Jianfei  and
      Li, Ziyan  and
      Wang, Jieming  and
      Xia, Rui",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.508/",
    doi = "10.18653/v1/2023.acl-long.508",
    pages = "9141--9154"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/