Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task

Stan Weixian LeiDifei GaoJay Zhangjie WuYuxuan WangWei LiuMengmi ZhangMike Zheng Shou

article2023AAAI60 citations

Proposes a real-data-free continual learning framework that replays scene graphs instead of images to prevent catastrophic forgetting in visual question answering models across evolving visual environments and question types.

Listen

Visual question answering systems aim to answer open-ended questions about images, but real-world deployments require these models to continuously expand their capabilities and adapt to new environments. Existing systems are typically trained once on static datasets, causing them to rapidly forget previous knowledge when updated on new tasks—a challenge known as catastrophic forgetting. Addressing this issue through continual learning is critical for scalable artificial intelligence deployments, yet previous continual learning benchmarks largely overlooked the multi-modal reasoning and dynamic task expansions required for visual question answering.

The article establishes a benchmark to evaluate continual learning in visual question answering across expanding visual environments and functional skills. In addition, it develops and evaluates a data-free replay framework that allows models to retain historical knowledge without storing past user images or data.

To evaluate these capabilities, the researchers constructed a new benchmark consisting of two tracks: a scene-incremental track spanning six visual domains (such as workplaces and sports) and a function-incremental track spanning six distinct operational capabilities (such as object recognition and scene text reading). The authors then evaluated a proposed method, Scene Graph as Prompt for symbolic replay, which uses structured, text-like scene graphs as symbolic representations of past images. A language-based symbolic replay model generates synthetic scene graphs paired with questions and answers from minimal prompts, and a unified multimodal model trains on both new inputs and these replayed synthetic triplets.

The experimental findings show that the proposed symbolic replay method outperforms existing data-free continual learning approaches by a wide margin. In the functional expansion setting, the proposed method achieved average accuracies between 38.65% and 45.97% across various task sequences, exceeding existing real-data-free baselines by up to 13 percentage points. Furthermore, it performed competitively with methods that save actual historical images, matching the performance of real-data replay systems while requiring up to 100 times less storage (saving 612 kilobytes of prompts compared to 60 megabytes of stored images). Ablation studies showed that the method remains highly efficient, with a model trained on only 1% of scene graph annotations still surpassing previous pseudo-replay techniques by nearly 5 percentage points.

These results demonstrate that structured symbolic representations can effectively bridge language and visual reasoning without retaining sensitive user images. This allows organizations to continuously upgrade deployed artificial intelligence models while complying with strict data privacy regulations, minimizing storage infrastructure, and avoiding costly full-model retraining.

Decision-makers implementing continuous visual intelligence systems should consider symbolic replay over standard parameter-regularization methods, which struggle to decouple multimodal representations. For optimal performance, teams should calibrate replay sample volumes to balance knowledge retention and current task accuracy without oversaturating training distributions.

The article notes that in domain-shifting visual environments (the scene-incremental track), symbolic replay exhibits a slight performance gap compared to storing raw images, as text-based graphs do not capture every fine-grained visual detail. While confidence in the symbolic replay mechanism is high for functional reasoning tasks, organizations deploying in purely vision-heavy domain adaptation settings should anticipate moderate trade-offs between absolute visual accuracy and strict data privacy compliance.

Cover for Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task

Table of Contents

  • Introduction
  • Related Work
  • CLOVE Benchmark
  • Task Formulation
  • CLOVE-scene Setting
  • CLOVE-function Setting
  • Evaluation Metric
  • Method
  • Overview of Continual Learning Pipeline
  • Unified VQA Transformer
  • Baselines
  • Results and Analyses
  • Experiments
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — CLOVE benchmark for continual visual question answering

    definition

    CLOVE (Continual Learning On Visual quEstion answering) defines continual learning for VQA as sequentially learning a series of multimodal tasks while retaining performance on earlier image-question-answer distributions. A task sequence is T=(T_1,dots,T_N), where task TiT_i is trained on a dataset DiD_i of triplets (v,q,a)(v,q,a) consisting of an image vv, question qq, and answer aa.

    CLOVE contains two settings:

    • CLOVE-scene models deployment in new visual environments and contains six scene tasks: ShopAndDinning, Workplace, HomeOrHotel, Transportation, SportAndLeisure, and Outdoors.
    • CLOVE-function models the acquisition of new VQA abilities and contains six function tasks: object recognition, attribute recognition, relation reasoning, logic reasoning, knowledge reasoning, and scene text recognition.

    For an answer prediction, CLOVE follows VQA v2 soft scoring. If nan_a of ten human annotations agree with the predicted answer, the per-question accuracy is Acc⁡(a)=min⁡(na/3,1)\operatorname{Acc}(a)=\min(n_a/3,1). If ak,ja_{k,j} is the accuracy on task TjT_j after sequential training through task TkT_k, the average accuracy after task TkT_k is Ak=1k∑j=1kak,jA_k=\frac{1}{k}\sum_{j=1}^{k}a_{k,j}.

  2. Knowl 2 — Construction of the CLOVE scene and function task sequences

    experimental setup

    The six CLOVE-scene datasets are created from GQA images using the second-level scene taxonomy of SUN. An off-the-shelf scene classifier first partitions images; low-confidence images and images containing objects unusually frequent for the assigned scene are removed. A random sample of 100 images per scene was evaluated by three human workers, yielding a mean scene-assignment accuracy of 91.0%. Each scene task contains both scene-specific questions, whose answers are unique to that scene, and questions about concepts shared across scenes, such as color or material. Common and unique answers are sampled in similar proportions, task sizes are balanced, and answer distributions are smoothed following GQA.

    CLOVE-function combines samples from GQA, CRIC, and TextVQA. Object recognition, attribute recognition, relation reasoning, and logic reasoning are sourced from GQA; knowledge reasoning comes from CRIC; and scene text recognition comes from TextVQA. Questions from datasets with functional programs are assigned to a function task according to the operations in their reasoning program: object recognition uses operations such as Select, Query, and Choose name; attribute recognition uses Query, Verify, Choose, and Filter; relation reasoning uses Relate, Verify, and Choose relation; logic reasoning uses Different, Same, Common, and Choose; knowledge reasoning uses knowledge-graph lookup; and scene text recognition uses OCR-oriented recognition. The operation sets are not mutually exclusive because different functions can share basic VQA operations. The task distributions are balanced to reduce effects from unequal dataset sizes.

  3. Knowl 3 — Scene Graph as Prompt for symbolic replay

    model/method

    Scene Graph as Prompt (SGP) is a real-data-free replay framework for continual VQA. It contains a Symbolic Replay Model (SRM), denoted by SS, and a Unified VQA model (UniVQA), denoted by UU.

    Before learning a new task, the SRM uses a short scene-graph relationship as a prompt and autoregressively generates a completed pseudo scene graph together with a question-answer pair. The generated triplet replaces the unavailable past image-question-answer example during replay. The current task's real samples and the generated samples are then used to update both the SRM and UniVQA. The SRM preserves symbolic visual relations and their associated question-answering patterns, while UniVQA learns to answer questions from images, scene graphs, OCR tokens, knowledge representations, and replayed language-only inputs. Because only symbolic prompts rather than past images are retained, SGP is designed for settings in which storing real data is disallowed for privacy reasons.

  4. Knowl 4 — Joint symbolic replay model for scene graphs and question-answer pairs

    model/method

    The SRM is a DistilGPT2-based autoregressive language model trained to represent image scene graphs and their question-answering patterns. A scene graph is serialized as a token sequence G=(g1,…,gM)G=(g_1,\ldots,g_M), with a separation token between relationships. For a training collection DSGD_{\mathrm{SG}} of scene-graph sequences and SRM parameters θ\theta, scene-graph completion uses next-token loss

    LSG(θ)=−∑G∈DSG∑m=1Mlog⁡Pθ(gm∣g1,…,gm−1).\mathcal{L}_{\mathrm{SG}}(\theta)=-\sum_{G\in D_{\mathrm{SG}}}\sum_{m=1}^{M}\log P_{\theta}(g_m\mid g_1,\ldots,g_{m-1}).

    For each annotated question-answer example, let GqaG_{qa} be the scene-graph relationships relevant to its question and answer, and let DqaD_{qa} contain tuples (Gqa,q,a)(G_{qa},q,a). The supervised question-answer loss is

    LQA(θ)=−∑(Gqa,q,a)∈Dqalog⁡Pθ(q,a∣Gqa).\mathcal{L}_{\mathrm{QA}}(\theta)=-\sum_{(G_{qa},q,a)\in D_{qa}}\log P_{\theta}(q,a\mid G_{qa}).

    The joint SRM objective is LSRM=LQA+λLSG\mathcal{L}_{\mathrm{SRM}}=\mathcal{L}_{\mathrm{QA}}+\lambda\mathcal{L}_{\mathrm{SG}}, where λ\lambda weights scene-graph completion. During training, the serialized input contains a generation token, scene-graph relationships, a question token, the question, an answer token, and the answer. During replay, the SRM receives a generation token, a sampled scene-graph prompt, and a separator; it completes the scene graph and generates a question and answer conditioned on the same prompt.

  5. Knowl 5 — Privacy-preserving scene-graph prompt memory

    algorithm

    SGP does not retain complete past scene graphs or real images as an external episodic memory. For each task, it counts the frequencies of objects, attributes, and relations appearing in the task's training data. It then samples one to three scene-graph items using these frequencies and stores only the sampled items as prompts for future replay.

    At a later task, a prompt is serialized after the generation token and separator token. The SRM autoregressively completes the scene graph, then generates a related question and answer. Prompts from each previous task are used to produce pseudo-replay samples, so the method retains a compact symbolic trace of past visual contexts without retaining the original image-question-answer triplets.

  6. Knowl 6 — Unified multimodal VQA Transformer

    model/method

    UniVQA is a multimodal Transformer that accepts heterogeneous inputs encountered across CLOVE tasks. Question words, scene-graph text, and knowledge representations are embedded with language features from a pretrained BERT model. Visual object features are extracted with Faster R-CNN and combine appearance and location information. OCR token features are extracted when scene text recognition is required.

    All features are projected into a common representation space and processed by stacked multimodal Transformer layers. An autoregressive decoder predicts each answer word either from the normal answer vocabulary or by copying an OCR token through a dynamic pointer network. For current-task examples, UniVQA uses object features, question features, and task-specific features such as OCR or knowledge inputs. It also receives a language representation of an offline-generated scene graph. For replayed examples, it receives the language representation of the SRM-generated scene graph instead of image features. During training, current-task object features are randomly masked with probability 0.150.15, encouraging the model to use scene-graph information.

  7. Knowl 7 — Sequential training with balanced symbolic replay

    model/method

    When training task TiT_i for i>1i>1, UniVQA is optimized on the union of current-task data and SRM-generated samples from all earlier tasks. If DiD_i is the current dataset, Si−1S_{i-1} denotes the distribution of replayed triplets generated from tasks 11 through i−1i-1, ϕ\phi denotes UniVQA parameters, and L\mathcal{L} is the answer-prediction loss, the objective is

    LVQA(ϕ)=E(v,q,a)∼Di[L(U(v,q;ϕ),a)]+E(v′,q′,a′)∼Si−1[L(U(v′,q′;ϕ),a′)].\mathcal{L}_{\mathrm{VQA}}(\phi)=\mathbb{E}_{(v,q,a)\sim D_i}\left[\mathcal{L}(U(v,q;\phi),a)\right]+\mathbb{E}_{(v',q',a')\sim S_{i-1}}\left[\mathcal{L}(U(v',q';\phi),a')\right].

    The replay ratio γ\gamma is the number of generated samples relative to the number of current-task samples. The experiments generally use γ=1.5\gamma=1.5; when training TiT_i, γ∣Di∣\gamma|D_i| replay samples are generated in total and distributed equally among the i−1i-1 preceding tasks. The SRM is likewise trained sequentially using current ground-truth scene graphs together with symbolic-replayed samples.

  8. Knowl 8 — Continual VQA evaluation protocol and baselines

    experimental setup

    The experiments evaluate six randomly selected task orders for both CLOVE settings. The scene-order notation abcdefabcdef represents ShopAndDinning →\rightarrow Workplace →\rightarrow HomeOrHotel →\rightarrow Transportation →\rightarrow SportAndLeisure →\rightarrow Outdoors. The function-order notation oarlksoarlks represents object recognition →\rightarrow attribute recognition →\rightarrow relation reasoning →\rightarrow logic reasoning →\rightarrow knowledge reasoning →\rightarrow scene text recognition. Other reported permutations are bdfcaebdfcae, beacfdbeacfd, beadcfbeadcf, bedfcabedfca, ecdfabecdfab for scenes and roslakroslak, rklsaorklsao, rsolakrsolak, lkosralkosra, kaorlskaorls for functions.

    SGP is compared with sequential fine-tuning, online EWC, MAS, LAMOL-m adapted to multimodal VQA, VQG replay using saved images and answers, and two real-data replay methods. Real-rnd randomly stores samples, whereas Real-kmeans selects samples closest to cluster centroids in UniVQA feature space. Offline training on all tasks is reported as an upper bound. Unless otherwise stated, results use the final model after the last task and average accuracy as the metric.

  9. Knowl 9 — SGP substantially improves average accuracy without storing real data

    data/table

    The following values are average accuracy in percent after the final task for six scene and six function task orders. The comparison tests whether symbolic scene-graph replay preserves prior VQA abilities better than fine-tuning, regularization, alternative pseudo-replay, and real-data replay. SGP is the strongest real-data-free method in every listed order and is close to real-data replay on CLOVE-function, while real-data replay remains stronger on CLOVE-scene.

    Could not parse LaTeX table

    The results also show that EWC and MAS can perform no better than fine-tuning on some orders, indicating that parameter-importance regularization does not reliably separate multimodal representation and reasoning knowledge. Pseudo-replay methods improve over fine-tuning but remain substantially below SGP.

  10. Knowl 10 — Scene-graph replay is the main source of SGP's continual-learning benefit

    data/table

    An ablation under replay ratio γ=0.9\gamma=0.9 compares random prompts, ground-truth prompts, and which elements are replayed. The results show that replaying only generated questions and answers is weaker than replaying a generated scene graph together with the question and answer. Replacing random prompts with ground-truth scene-graph prompts further improves performance.

    Could not parse LaTeX table

    Using ground-truth prompts instead of randomly sampled prompts raises average accuracy by 3.01 percentage points on CLOVE-scene and 2.8 percentage points on CLOVE-function. Experiments that vary the replay ratio show low performance when too few samples are generated, with performance becoming comparatively stable once γ\gamma exceeds approximately 0.70.7. Oracle-style replay using predicted scene graphs paired with ground-truth questions and answers performs better than ordinary SGP, indicating that the quality of generated scene-graph-question-answer triplets remains an important bottleneck.

  11. Knowl 11 — Memory efficiency and limitations of symbolic replay

    limitation

    SGP requires relatively few scene-graph annotations to train its replay model. In CLOVE-function order oarlksoarlks, using 1%, 5%, 50%, and 75% of the available scene-graph annotations produced average accuracies of 37.68%, 41.94%, 43.27%, and 44.15%, respectively. The 1% setting corresponds to roughly 200 VQA samples and still exceeds VQG by 4.9 percentage points; using 50% remains close to Real-rnd.

    The external memory is also compact: storing 612 KB of scene-graph prompts gives performance comparable to Real-rnd storing 24 MB on CLOVE-scene and 60 MB on CLOVE-function, corresponding to approximately 40.2-fold and 100.4-fold less storage, respectively.

    The method has a setting-dependent weakness. On CLOVE-scene, SGP is below real-data replay because disjoint visual domains make it difficult for generated scene graphs to preserve all visual knowledge carried by images. The paper identifies three unresolved challenges for continual VQA: separating multimodal knowledge with regularization, finding discriminative representations for selecting real replay samples, and generating sufficiently faithful alternatives to complex images.

Coverage note — Implementation details of individual baseline methods and qualitative examples from the paper's illustrations were omitted because they do not add independent load-bearing contributions beyond the benchmark, SGP method, ablations, and reported results.

References

  1. 1.Aljundi, R.; Babiloni, F.; Elhoseiny, M.; Rohrbach, M.; and Tuytelaars, T. 2018. Memory aware synapses: Learning what (not) to forget. In Proceedings of the European Conference on Computer Vision (ECCV), 139–154.
  2. 2.Aljundi, R.; Chakravarty, P.; and Tuytelaars, T. 2017. Expert gate: Lifelong learning with a network of experts. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3366–3375.
  3. 3.Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, 6077–6086.
  4. 4.Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015. VQA: Visual Question Answering. In Proceedings of the IEEE International Conference on Computer Vision (ICCV).
  5. 5.Biesialska, M.; Biesialska, K.; and Costa-jussa, M. R. `2020. Continual Lifelong Learning in Natural Language Processing: A Survey. In Proceedings of the 28th International Conference on Computational Linguistics, 6523–6541. Barcelona, Spain (Online): International Committee on Computational Linguistics.
  6. 6.Chaudhry, A.; Dokania, P. K.; Ajanthan, T.; and Torr, P. H. 2018. Riemannian walk for incremental learning: Understanding forgetting and intransigence. In Proceedings of the European Conference on Computer Vision (ECCV), 532–547.
  7. 7.Chen, Y.-C.; Li, L.; Yu, L.; El Kholy, A.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020. Uniter: Universal image-text representation learning. In European conference on computer vision, 104–120. Springer.
  8. 8.Chen, Z.; Ma, N.; and Liu, B. 2015. Lifelong Learning for Sentiment Classification. international joint conference on natural language processing.
  9. 9.Gao, D.; Li, K.; Wang, R.; Shan, S.; and Chen, X. 2020. Multi-modal graph neural network for joint reasoning on vision and scene text. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 12746–12756.
  10. 10.Gao, D.; Wang, R.; Shan, S.; and Chen, X. 2019. CRIC: A vqa dataset for compositional reasoning on vision and commonsense. arXiv preprint arXiv:1908.02962.
  11. 11.Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017. Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. In Conference on Computer Vision and Pattern Recognition (CVPR).
  12. 12.Greco, C.; Plank, B.; Fernandez, R.; and Bernardi, R. ´ 2019. Psycholinguistics Meets Continual Learning: Measuring Catastrophic Forgetting in Visual Question Answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3601–3605. Florence, Italy: Association for Computational Linguistics.
  13. 13.Hu, J.; Ruder, S.; Siddhant, A.; Neubig, G.; Firat, O.; and Johnson, M. 2020a. Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation. In International Conference on Machine Learning, 4411–4421. PMLR.
  14. 14.Hu, R.; Singh, A.; Darrell, T.; and Rohrbach, M. 2020b. Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.
  15. 15.Huang, Y.; Zhang, Y.; Chen, J.; Wang, X.; and Yang, D. 2021. Continual learning for text classification with information disentanglement based regularization. arXiv preprint arXiv:2104.05489.
  16. 16.Hudson, D. A.; and Manning, C. D. 2019. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 6700–6709.
  17. 17.Jiang, H.; Misra, I.; Rohrbach, M.; Learned-Miller, E.; and Chen, X. 2020. In defense of grid features for visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10267–10276.
  18. 18.Johnson, J.; Hariharan, B.; van der Maaten, L.; Fei-Fei, L.; Lawrence Zitnick, C.; and Girshick, R. 2017. CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  19. 19.Kirkpatrick, J.; Pascanu, R.; Rabinowitz, N.; Veness, J.; Desjardins, G.; Rusu, A. A.; Milan, K.; Quan, J.; Ramalho, T.; Grabska-Barwinska, A.; et al. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13): 3521–3526.
  20. 20.Krishna, R.; Bernstein, M.; and Fei-Fei, L. 2019. Information Maximizing Visual Question Generation. In IEEE Conference on Computer Vision and Pattern Recognition.
  21. 21.Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123(1): 32–73.
  22. 22.Lee, S. 2017. Toward Continual Learning for Conversational Agents. arXiv: Computation and Language.
  23. 23.Li, Z.; and Hoiem, D. 2017. Learning without forgetting. IEEE transactions on pattern analysis and machine intelligence, 40(12): 2935–2947.
  24. 24.Liu, Y.; Schiele, B.; and Sun, Q. 2021. Adaptive aggregation networks for class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2544–2553.
  25. 25.Lopez-Paz, D.; and Ranzato, M. 2017. Gradient episodic memory for continual learning. Advances in neural information processing systems, 30.
  26. 26.Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32.
  27. 27.Mallya, A.; and Lazebnik, S. 2018. Packnet: Adding multiple tasks to a single network by iterative pruning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 7765–7773.
  28. 28.Marino, K.; Rastegari, M.; Farhadi, A.; and Mottaghi, R. 2019. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, 3195–3204.
  29. 29.Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8): 9.
  30. 30.Rebuffi, S.-A.; Bilen, H.; and Vedaldi, A. 2017. Learning multiple visual domains with residual adapters. Advances in neural information processing systems, 30.
  31. 31.Rebuffi, S.-A.; Kolesnikov, A.; Sperl, G.; and Lampert, C. H. 2017. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2001–2010.
  32. 32.Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28.
  33. 33.Sauer, A.; Schwarz, K.; and Geiger, A. 2022. Stylegan-xl: Scaling stylegan to large diverse datasets. In Special Interest Group on Computer Graphics and Interactive Techniques Conference Proceedings, 1–10.
  34. 34.Schwarz, J.; Czarnecki, W.; Luketina, J.; Grabska-Barwinska, A.; Teh, Y. W.; Pascanu, R.; and Hadsell, R. 2018. Progress & compress: A scalable framework for continual learning. In International Conference on Machine Learning, 4528–4537. PMLR.
  35. 35.Serra, J.; Suris, D.; Miron, M.; and Karatzoglou, A. 2018. Overcoming catastrophic forgetting with hard attention to the task. In International Conference on Machine Learning, 4548–4557. PMLR.
  36. 36.Shin, H.; Lee, J. K.; Kim, J.; and Kim, J. 2017. Continual learning with deep generative replay. Advances in neural information processing systems, 30.
  37. 37.Singh, A.; Natarjan, V.; Shah, M.; Jiang, Y.; Chen, X.; Parikh, D.; and Rohrbach, M. 2019. Towards VQA Models That Can Read. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 8317–8326.
  38. 38.Su, W.; Zhu, X.; Cao, Y.; Li, B.; Lu, L.; Wei, F.; and Dai, J. 2019. Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530.
  39. 39.Sun, F.-K.; Ho, C.-H.; and Lee, H.-Y. 2020. LAMOL: Language Modeling Is All You Need for Lifelong Language Learning. In International Conference on Learning Representations.
  40. 40.Tang, K.; Niu, Y.; Huang, J.; Shi, J.; and Zhang, H. 2020. Unbiased Scene Graph Generation from Biased Training. In Conference on Computer Vision and Pattern Recognition.
  41. 41.Van de Ven, G. M.; and Tolias, A. S. 2018. Generative replay with feedback connections as a general strategy for continual learning. arXiv preprint arXiv:1809.10635.
  42. 42.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  43. 43.Xiao, J.; Hays, J.; Ehinger, K. A.; Oliva, A.; and Torralba, A. 2010. Sun database: Large-scale scene recognition from abbey to zoo. In 2010 IEEE computer society conference on computer vision and pattern recognition, 3485–3492. IEEE.

Citation

MLA
Lei, S. W., et al. “Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task”. arXiv, 2022, http://arxiv.org/abs/2208.12037v2.
APA
Lei, S. W., Gao, D., Wu, J. Z., Wang, Y., Liu, W., Zhang, M., & Shou, M. Z. (2022). Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task. arXiv. http://arxiv.org/abs/2208.12037v2
Chicago
Lei, S. W., D. Gao, J. Z. Wu, et al. 2022. “Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task”. arXiv. http://arxiv.org/abs/2208.12037v2.
Harvard
Lei, S.W. et al. (2022) “Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2208.12037v2.
Vancouver
1. Lei SW, Gao D, Wu JZ, Wang Y, Liu W, Zhang M, Shou MZ (2022) Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task. arXiv

BibTeX

@article{lei2022symbolic,
  title = {Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task},
  author = {Lei, Stan Weixian and Gao, Difei and Wu, Jay Zhangjie and Wang, Yuxuan and Liu, Wei and Zhang, Mengmi and Shou, Mike Zheng},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2208.12037v2},
  eprint = {2208.12037}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF