VizCopilot: Fostering Appropriate Reliance on Enterprise Chatbots with Context Visualization
Sam Yu-Te LeeJingya ChenAlbert CalzarettoRichard LeeSamir PassiAlice FerngMihaela Vorvoreanu
Demonstrates that integrating document visualization with topic modeling allows knowledge workers to inspect and adjust retrieved context in enterprise chatbots, improving answer relevance and user prompting strategies.
Enterprise artificial intelligence chatbots promise to streamline knowledge work by synthesizing vast corporate data archives, yet they frequently generate answers that are plausible but misaligned with actual user intent. This creates a critical operational risk known as overreliance, where users uncritically accept flawed recommendations that can silently degrade decision-making, increase compliance risk, and harm performance. Conventional conversational interfaces leave underlying retrieval processes obscured, forcing users to rely on trial-and-error prompting without providing transparent mechanisms to inspect or steer the contextual data feeding the model.
The article evaluates whether visual context engineering—giving knowledge workers direct visual oversight and control over retrieved data—fosters appropriate reliance on enterprise chatbots. It demonstrates this approach through a research prototype, VizCopilot, which couples a conversational assistant with an interactive treemap visualization.
To assess this paradigm, the authors conducted a qualitative Research-through-Design study comparing VizCopilot against a baseline pure-text chat interface. Fourteen participants with prior enterprise chatbot experience completed realistic information synthesis tasks. The testing leveraged a synthetic corporate corpus representing approximately 1,000 employees and 10,000 messy enterprise records, complete with realistic metadata conflicts, duplicate items, and unstructured content.
The investigation produced four primary findings. First, interactive visual scaffolding enabled users to rapidly detect context misalignment, such as missing topics or extraneous data, without manually reviewing every file. Second, direct manipulation allowed users to successfully correct errors: for instance, 10 out of 14 participants detected and resolved a subtle error where the chatbot conflated two distinct employees sharing the same name by inspecting the visual file view. Third, visual context substantially improved user agency and efficiency; participants required fewer follow-up prompts to achieve target results and naturally adapted their queries using highlighted keywords. Finally, participants showed persistent skepticism toward automated subtopic text summaries, consistently preferring direct access to raw evidence for critical verification.
These findings indicate that integrating interactive visual structures can meaningfully augment human oversight and reduce overreliance without creating excessive cognitive strain. By exposing intermediate retrieval steps, the interface counters anthropomorphic assumptions about artificial intelligence, encouraging users to treat the tool as an inspectable search-and-synthesis system. This balance between automation and direct manipulation offers a viable pathway toward meeting emerging regulatory requirements for human oversight in high-risk automated systems.
Organizations developing or deploying enterprise chatbots should consider incorporating group-level visual context controls alongside conversational prompts to improve retrieval alignment. System designers must also provide explicit uncertainty indicators to trigger manual verification on deceptively simple queries and improve tools for rapid raw-document inspection, such as keyword highlighting. Moving forward, teams should pilot these interfaces across larger real-world data environments to evaluate technical scalability and refine onboarding workflows that ease initial interface complexity.
The conclusions should be interpreted within the boundaries of an exploratory qualitative study using a synthetic dataset and brief participant sessions. While long-term longitudinal studies are needed before scaling across production enterprise suites, the article provides moderate to high confidence that visual context engineering is a superior design direction for maintaining human agency and reliable system oversight.
- Paper: From Local to Global: A Graph RAG Approach to Query-Focused Summarization, Darren Edge et al. (2024). Introduces GraphRAG for hierarchically summarizing and structuring document collections, providing essential context for the information synthesis and topic-level aggregation challenges VizCopilot addresses.
- Paper: Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models, Wenhao Yu et al. (2024). Examines how retrieval-augmented language models handle irrelevant or noisy context, motivating VizCopilot's focus on context misalignment and verification.
- Paper: Knowledge-Grounded Dialogue Generation with a Unified Knowledge Representation, Yu Li et al. (2022). Establishes techniques for grounding conversational dialogue systems in heterogeneous external knowledge bases before answer generation.
- Paper: LaMDA: Language Models for Dialog Applications, Romal Thoppilan et al. (2022). Provides foundational principles for integrating external retrieval tools into conversational dialogue models to improve factual grounding.
- Paper: LM Agents for Coordinating Multi-User Information Gathering, Harsh Jhamtani et al. (2025). Extends single-user enterprise context gathering and synthesis to collaborative, multi-user information retrieval environments.
- Paper: YES AND: A Generative AI Multi-Agent Framework for Enhancing Diversity of Thought in Individual Ideation for Problem-Solving Through Confidence-Based Agent Turn-Taking, Pratik Ghosh et al. (2025). Explores multi-agent collaborative ideation frameworks in enterprise tasks, continuing the design paradigm of structured human-AI collaboration.
- Paper: Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context, Keivan Alizadeh et al. (2026). Applies self-reflective uncertainty estimation and program search to mitigate context degradation over long, heterogeneous enterprise documents.
