Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA Task
Stan Weixian LeiDifei GaoJay Zhangjie WuYuxuan WangWei LiuMengmi ZhangMike Zheng Shou
Proposes a real-data-free continual learning framework that replays scene graphs instead of images to prevent catastrophic forgetting in visual question answering models across evolving visual environments and question types.
Visual question answering systems aim to answer open-ended questions about images, but real-world deployments require these models to continuously expand their capabilities and adapt to new environments. Existing systems are typically trained once on static datasets, causing them to rapidly forget previous knowledge when updated on new tasks—a challenge known as catastrophic forgetting. Addressing this issue through continual learning is critical for scalable artificial intelligence deployments, yet previous continual learning benchmarks largely overlooked the multi-modal reasoning and dynamic task expansions required for visual question answering.
The article establishes a benchmark to evaluate continual learning in visual question answering across expanding visual environments and functional skills. In addition, it develops and evaluates a data-free replay framework that allows models to retain historical knowledge without storing past user images or data.
To evaluate these capabilities, the researchers constructed a new benchmark consisting of two tracks: a scene-incremental track spanning six visual domains (such as workplaces and sports) and a function-incremental track spanning six distinct operational capabilities (such as object recognition and scene text reading). The authors then evaluated a proposed method, Scene Graph as Prompt for symbolic replay, which uses structured, text-like scene graphs as symbolic representations of past images. A language-based symbolic replay model generates synthetic scene graphs paired with questions and answers from minimal prompts, and a unified multimodal model trains on both new inputs and these replayed synthetic triplets.
The experimental findings show that the proposed symbolic replay method outperforms existing data-free continual learning approaches by a wide margin. In the functional expansion setting, the proposed method achieved average accuracies between 38.65% and 45.97% across various task sequences, exceeding existing real-data-free baselines by up to 13 percentage points. Furthermore, it performed competitively with methods that save actual historical images, matching the performance of real-data replay systems while requiring up to 100 times less storage (saving 612 kilobytes of prompts compared to 60 megabytes of stored images). Ablation studies showed that the method remains highly efficient, with a model trained on only 1% of scene graph annotations still surpassing previous pseudo-replay techniques by nearly 5 percentage points.
These results demonstrate that structured symbolic representations can effectively bridge language and visual reasoning without retaining sensitive user images. This allows organizations to continuously upgrade deployed artificial intelligence models while complying with strict data privacy regulations, minimizing storage infrastructure, and avoiding costly full-model retraining.
Decision-makers implementing continuous visual intelligence systems should consider symbolic replay over standard parameter-regularization methods, which struggle to decouple multimodal representations. For optimal performance, teams should calibrate replay sample volumes to balance knowledge retention and current task accuracy without oversaturating training distributions.
The article notes that in domain-shifting visual environments (the scene-incremental track), symbolic replay exhibits a slight performance gap compared to storing raw images, as text-based graphs do not capture every fine-grained visual detail. While confidence in the symbolic replay mechanism is high for functional reasoning tasks, organizations deploying in purely vision-heavy domain adaptation settings should anticipate moderate trade-offs between absolute visual accuracy and strict data privacy compliance.
- Paper: Learning to Prompt for Continual Learning, Zifeng Wang et al. (2021). Introduces prompt-based continual learning without storing past data or updating core weights, establishing key prompt mechanics leveraged in data-free continual multimodal setups.
- Paper: Continual Learning with Deep Generative Replay, Hanul Shin et al. (2017). Establishes the foundational concept of deep generative replay for continual learning to overcome catastrophic forgetting without storing real historical images.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Provides the foundational distillation methodology for adapting neural networks to sequential tasks without access to legacy training data.
- Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). Defines key replay baseline dynamics and logits-matching strategies for general continual learning scenarios across task-incremental and domain-incremental boundaries.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). Presents a unified transformer baseline for vision-and-language tasks like VQA, illustrating how cross-modal visual and textual features are jointly aligned.
- Paper: Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering, Yash Goyal et al. (2016). Establishes standard benchmarks and diagnostic foundations for understanding multimodal visual reasoning and language dependencies in VQA.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). Provides a comprehensive, unified theoretical survey categorizing replay, prompt, and representation-based continual learning methods evaluated in specialized benchmark studies.
- Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). Extends continual learning evaluation to massive long-horizon memorization horizons by investigating how multiple retention mechanisms compose beyond single-method setups.
- Paper: Class-Incremental Exemplar Compression for Class-Incremental Learning, Zilin Luo et al. (2023). Develops complementary visual exemplar compression techniques to optimize historical storage trade-offs in incremental learning.
