Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts

Yunfan ZhouXiwen CaiQiming ShiYanwei HuangHaotian LiHuamin QuDi WengYingcai Wu

article2025International Conference on Human Factors in Computing Systems6 citationsBest Paper Award, CHI 2025

Introduces Xavier, a computational notebook assistant that integrates tabular data schemas and values into code completion to generate context-aware suggestions and provide real-time transformation previews.

Listen

Data wrangling—cleaning, transforming, and integrating raw datasets—is a critical bottleneck in data science workflows. While data analysts increasingly rely on artificial intelligence coding assistants such as GitHub Copilot, current tools primarily focus on language grammar and code semantics. They frequently ignore dataset metadata such as table schemas, column names, and unique data values. This lack of data awareness leads to hallucinated column names, syntactical mistakes, and constant workflow interruptions as analysts must repeatedly write exploratory code or open raw files to inspect their data.

The article demonstrates and evaluates Xavier, a computational notebook extension designed to improve the authoring of tabular data wrangling scripts in Python Pandas. The objective of the research is to maintain continuous data context awareness by dynamically linking active dataset properties with code suggestions and live visual feedback.

The authors designed Xavier using a modular architecture consisting of a code context manager, a data context manager, and a completion generator powered by the open-source Llama3-70B model. Alongside an initial preliminary observational study of nine data professionals, the researchers evaluated the system through a counterbalanced, mixed-design user study with 16 data analysts. Participants completed standardized data wrangling tasks using both Xavier and a baseline tool that lacked integrated data context and dynamic visual feedback, measuring completion time, error rates, context switches, and perceived mental workload.

The evaluation revealed three primary outcomes. First, analysts using Xavier experienced significantly fewer context switches and made fewer coding and data errors (both p < 0.001) compared to the baseline tool. Second, subjective workload assessments showed lower mental demand, effort, and frustration when using Xavier. Third, while accuracy and workflow continuity improved markedly, the difference in total task completion time was not statistically significant, likely due to the concise nature of the experimental scripting tasks. Participants strongly favored shorter, highly accurate code completions (such as column names and parameters) over multi-line statement generations, and they noted that dynamic schema highlighting and instant data transformation previews greatly enhanced their confidence.

These findings indicate that integrating real-time dataset metadata directly into code generation prompts and user interfaces resolves core usability challenges in data analysis. Providing immediate visual previews substantially lowers the cognitive burden of verifying machine-generated code and eliminates disruptive manual data inspection. For organizations employing data teams, adopting data-aware programming assistants can reduce script errors, improve code quality, and lessen developer fatigue without requiring major workflow overhauls.

Organizations developing or deploying developer tools should prioritize tight integration between active runtime data and code assistance models. Tool designers should emphasize shorter, high-precision recommendations, control generation length, and implement persistent side-panel previews. Further technical work is recommended to optimize model latency, determine optimal data sampling thresholds for large enterprise databases, and expand parsing support beyond Python Pandas to other data analysis libraries.

The findings are bounded by the study's laboratory setting, which utilized standardized datasets and concise scripts of approximately 10 lines. Further research through long-term field deployments and eye-tracking studies is needed to definitively confirm productivity gains and validate attention dynamics on complex, large-scale enterprise workflows.

Cover for Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts

Abstract

Data analysts frequently employ code completion tools in writing custom scripts to tackle complex tabular data wrangling tasks. However, existing tools do not sufficiently link the data contexts such as schemas and values with the code being edited. This not only leads to poor code suggestions, but also frequent interruptions in coding processes as users need additional code to locate and understand relevant data. We introduce Xavier, a tool designed to enhance data wrangling script authoring in computational notebooks. Xavier maintains users' awareness of data contexts while providing data-aware code suggestions. It automatically highlights the most relevant data based on the user's code, integrates both code and data contexts for more accurate suggestions, and instantly previews data transformation results for easy verification. To evaluate the effectiveness and usability of Xavier, we conducted a user study with 16 data analysts, showing its potential to streamline data wrangling scripts authoring.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Data Wrangling Tools
  • 2.2 Code Assistants in Computational Notebooks
  • 2.3 Code Completion Tools
  • 2.4 Live Programming Tools
  • 3 Preliminary Study
  • 3.1 Participants
  • 3.2 Apparatus and Materials
  • 3.3 Procedure
  • 3.4 Findings
  • 3.4.1 Definition of Two Kinds of Activities
  • 3.4.2 User Behavior and Challenges
  • 3.5 User Requirements
  • 4 Xavier
  • 4.1 Overview
  • 4.2 Data Context-Aware Code Completion
  • 4.2.1 Code context detection
  • 4.2.2 Data context organization
  • 4.2.3 Code completion generation
  • 4.3 Automatic Data Context Highlighting
  • 4.4 Real-time Transformation Preview
  • 5 User Study
  • 5.1 Participants
  • 5.2 Apparatus and Materials
  • 5.3 Procedure
  • 5.4 Results
  • 5.4.1 Quantitative Results
  • 5.4.2 Qualitative Results
  • 6 Discussion
  • 6.1 Lessons Learned
  • 6.2 Limitations and Future Work
  • 6.2.1 Functionalities
  • 6.2.2 Evaluations
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Taxonomy of Data Contexts for Tabular Data Wrangling

    definition

    In the context of AI-assisted tabular data wrangling, data contexts comprise runtime tabular metadata and values structured across three granularities of tabular data objects:

    • Table-level data contexts: Basic structural metadata representing a table as a whole, including the table or variable identifier, column schemas and names, and matrix dimensions (shape).
    • Column-level data contexts: Detailed statistical and semantic metadata describing individual columns. Common attributes across all columns include data type, null value count, and sortedness. For categorical columns, this includes unique values, value frequencies, cardinality, and string format patterns. For numerical columns, this includes numerical range (minimum and maximum bounds), representative sample data points, and numerical formatting.
    • Row-level data contexts: Ordered subsets of sample rows that preserve the multidimensional and spatial relationships among adjacent table cells across columns.

    These context levels are combined according to operator requirements; for example, relational joins combine table-level column names with row-level sample records across multiple tables to infer key columns, whereas unary string operations rely on column-level format representations.

  2. Knowl 2 — Data Context-Aware Code Completion Pipeline in Xavier

    model/method

    The Xavier code completion pipeline generates data wrangling script recommendations by conditioning a large language model (LLM) or a prefix-matching engine on runtime data metadata combined with editor code context. The workflow consists of three stages:

    1. Code Context Detection: The system identifies the cursor position relative to data wrangling library function signatures (such as Python Pandas). When the cursor is inside a function signature, active positional and keyword arguments are parsed to determine unfilled parameters. When outside a signature, child-to-parent Abstract Syntax Tree (AST) pattern matching analyzes incomplete expressions (such as filtering brackets or chained function calls) to identify target transformation operators and missing AST nodes.
    2. Data Context Organization: Data contexts pre-calculated from active DataFrames in the notebook kernel are filtered in two successive steps:
      • Type Matching: Operators are classified into types (such as DataFrame-level, Series-level, or general) to filter out irrelevant context granularities (e.g., restricting Series string operations strictly to column-level contexts).
      • Context Selection by Operands: Specific table and column metadata are prioritized based on the missing AST nodes (e.g., dynamically prioritizing column unique values over full table samples once a column identifier is typed). To adhere to LLM context limits, high-cardinality categorical columns are sampled (e.g., up to 50 unique values) with support for dynamic prefix filtering.
    3. Code Completion Generation: A structured prompt containing four components is submitted to an LLM (such as Llama3-70B):
      • Code Context: All wrangling code preceding the cursor.
      • Data Context: Textual serialization of the filtered table-, column-, and row-level metadata.
      • Task Instruction: Step-by-step reasoning instructions to infer transformation arguments.
      • Format Control: Syntax constraints enforcing syntactically valid completions.

    For token-level completion, column and table identifiers are suggested directly via prefix matching against active table schemas.

  3. Knowl 3 — Non-Execution Real-Time Transformation Previews for Tabular Operations

    model/method

    Xavier provides non-execution visual previews of prospective data transformations within a split-window notebook side panel as the user types or navigates code completion suggestions. When all required arguments for an operation are present in a candidate completion, the system computes and renders the result using one of three preview modalities:

    1. Column Format Transformation Preview: For unary column transformations (such as string replacements, formatting, or null value imputation via .fillna()), a new preview column highlighted in yellow is rendered immediately to the right of the original column, with transformed table cells displayed in bold typeface.
    2. Table Filtering Transformation Preview: For boolean indexing and filtering operations, table rows evaluated as false (to be dropped) are marked in red, while the filter values matched in the user's code are highlighted in bold within remaining cells.
    3. Structural and Table-Level Transformation Preview: For global transformations that alter table geometry or instantiate new tables (such as sorting, aggregation, or multi-table joins), the interface simultaneously displays both the original table state and the newly computed table state side by side.
  4. Knowl 4 — Context-Adaptive Data Highlighting and Column Anchoring

    model/method

    In Xavier, the notebook data view dynamically alters its visual state to match the user's partial code and focused completion item:

    • Adaptive Schema and Row Display: When partial code or active completion suggestions contain only table identifiers (e.g., pd.concat([df1, ), the data view unfolds the schemas of the referenced DataFrames while collapsing unrelated tables and concealing sample rows. When column identifiers are present, the panel expands the sample row view (showing the top 15 rows) and applies distinct background highlighting to the referenced and candidate column headers and cells.
    • Floating Column Anchoring: When a DataFrame has more columns than can fit within the horizontal width of the side panel, highlighted columns currently referenced in the editor or completion popup are dynamically anchored (floated) to the right edge of the viewport. This keeps the relevant columns visible without requiring horizontal scrolling across the table.
  5. Knowl 5 — Activity Taxonomy of Tabular Data Wrangling Scripting

    definition

    Authoring tabular data wrangling scripts consists of interleaved activities across two primary categories:

    • Data Inspection (DI): Examining tabular data states, divided into:
      • Profiling (DI1DI_1): Examining raw datasets, table schemas, unique values, or printing intermediate variables (e.g., via df.head()) to understand the data and decide subsequent wrangling operations.
      • Verifying (DI2DI_2): Checking the correctness and output structure of a transformation after running code.
    • Code Authoring (CA): Writing programming instructions, divided into:
      • Creating (CA1CA_1): Writing new transformation logic or temporary profiling code.
      • Modifying (CA2CA_2): Editing existing lines, tuning parameters, or debugging errors.

    Without data-aware assistance, users frequently depart from the baseline iterative workflow (DI1→CA1→DI2DI_1 \to CA_1 \to DI_2) due to naming errors and hallucinations, creating dense cycles of debugging (CA2CA_2) and re-verification (DI2DI_2) accompanied by frequent context switches between code cells and external data views.

  6. Knowl 6 — Quantitative Evaluation of Xavier on Context Switches, Errors, and Workload

    empirical result

    A counterbalanced mixed-design user study (N=16N = 16 data analysts) evaluated Xavier against a baseline tool sharing the same underlying language model (Llama3-70B) without data context and equipped with an AutoProfiler continuous profiling panel. Statistical analysis using the Scheirer-Ray-Hare nonparametric test demonstrated:

    • Context Switches: Users experienced significantly fewer context switches (pauses in coding to inspect the side panel, write manual profiling code, or view raw files) when using Xavier compared to the baseline (p<0.001p < 0.001).
    • Errors: Users encountered significantly fewer coding and data errors (such as referencing non-existent columns or generating malformed outputs) when using Xavier (p<0.001p < 0.001).
    • Perceived Workload: Across all six NASA-TLX dimensions (mental demand, physical demand, temporal demand, performance, effort, and frustration), participants reported lower average workload scores for Xavier on a 7-point Likert scale.
    • Task Completion Time: No statistically significant main effect of the tool was observed on task completion time, which was attributed to the relatively short length of the benchmark tasks (approximately 10 lines of code per task).
  7. Knowl 7 — Granularity Preferences and Trust Dynamics in AI Code Completion

    empirical result

    User study interviews (N=16N = 16) revealed distinct preferences and trust behaviors when interacting with data-aware code completion tools:

    • Preference for Short-Granularity Completions: Participants consistently favored short completions (such as column names, individual data values, and function arguments) over full-line or multi-line transformation blocks. Shorter suggestions exhibited lower generation latency and reduced the cognitive effort needed to verify correctness.
    • Trust Fragility from Early Hallucinations: When the completion engine failed to generate an accurate multi-token statement on initial attempts, users quickly lost confidence in long completions and thereafter skipped long suggestions in favor of typing manual operators and accepting only single tokens.
    • Verification Utility of Dynamic Views: 14 out of 16 participants reported that automatic transformation previews helped prevent errors, while 13 out of 16 stated that automatic column highlighting provided intuitive reference without causing visual distraction during typing.
  8. Knowl 8 — Data Context Sampling and Latency Trade-Offs in LLM Prompts

    limitation

    Serializing runtime tabular contexts into LLM prompt templates is constrained by LLM context windows and inference latency. When columns have high cardinality, Xavier limits categorical metadata to a randomly sampled subset (at most 50 unique values per column). While client-side prefix matching allows users to target values beyond the sample, optimal sampling budgets for tabular prompts remain unbenchmarked. Furthermore, combining code contexts with serialized table data increases token lengths and network latency, causing generation delays that lead users to skip multi-token suggestions.

  9. Knowl 9 — Library-Specific AST Dependency in Wrangling Code Assistance

    limitation

    The code context manager and completion triggering mechanisms in Xavier rely on manually constructed syntax rules and Abstract Syntax Tree (AST) pattern matching tailored specifically to the Python Pandas API. Extending this data-aware completion, highlighting, and preview architecture to other programming languages (such as R) or alternative tabular libraries (such as dplyr, Polars, or SQL) requires re-engineering syntax matching rules or developing a unified intermediate domain-specific language (DSL).

Coverage note — None was omitted; all key contributions, preliminary study findings, system methods, user study results, interaction dynamics, and limitations were extracted.

References

  1. 1.Alfred V. Aho, Monica S. Lam, Ravi Sethi, and Jeffrey D. Ullman. 2006. Compilers: Principles, Techniques, and Tools (2nd Edition). Addison-Wesley Longman Publishing Co., Inc., USA.
  2. 2.Shraddha Barke, Michael B James, and Nadia Polikarpova. 2023. Grounded Copilot: How Programmers Interact with Code-Generating Models. Proceedings of the ACM on Programming Languages 7, OOPSLA1 (2023), 85–111.
  3. 3.Rohan Bavishi, Caroline Lemieux, Roy Fox, Koushik Sen, and Ion Stoica. 2019. AutoPandas: Neural-Backed Generators for Program Synthesis. Proceedings of the ACM on Programming Languages 3, OOPSLA (2019), 1–27.
  4. 4.Marcel Bruch, Martin Monperrus, and Mira Mezini. 2009. Learning from Examples to Improve Code Completion Systems. In Proceedings of the 7th Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on The Foundations of Software Engineering. ACM, New York, NY, USA, 213–222. https://doi.org/10.1145/1595696.1595728
  5. 5.Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. 2021. Extracting Training Data from Large Language Models. In Proceedings of the 30th USENIX Security Symposium. USENIX Association, Virtual Event, USA, 2633–2650. https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
  6. 6.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating Large Language Models Trained on Code. arXiv preprint arXiv:2107.03374 abs/2107.03374 (2021).
  7. 7.Ran Chen, Di Weng, Yanwei Huang, Xinhuan Shu, Jiayi Zhou, Guodao Sun, and Yingcai Wu. 2023. Rigel: Transforming Tabular Data by Declarative Mapping. IEEE Transactions on Visualization and Computer Graphics 29, 1 (2023), 128–138. https://doi.org/10.1109/TVCG.2022.3209385
  8. 8.Ruijia Cheng, Titus Barik, Alan Leung, Fred Hohman, and Jeffrey Nichols. 2024. BISCUIT: Scaffolding LLM-Generated Code with Ephemeral UIs in Computational Notebooks. In Proceedings of IEEE Symposium on Visual Languages and Human-Centric Computing. IEEE Computer Society, Los Alamitos, CA, USA, 13–23. https://doi.org/10.1109/VL/HCC60511.2024.00012
  9. 9.Bhavya Chopra, Anna Fariha, Sumit Gulwani, Austin Z. Henley, Daniel Perelman, Mohammad Raza, Sherry Shi, Danny Simmons, and Ashish Tiwari. 2023. CoWrangler: Recommender System for Data-Wrangling Scripts. In Proceedings of Companion of the International Conference on Management of Data. ACM, New York, NY, USA, 147–150. https://doi.org/10.1145/3555041.3589722
  10. 10.Robert A DeLine. 2021. Glinda: Supporting Data Science with Live Programming, GUIs and a Domain-specific Language. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 11 pages. https://doi.org/10.1145/3411764.3445267
  11. 11.Ian Drosos, Titus Barik, Philip J. Guo, Robert DeLine, and Sumit Gulwani. 2020. Wrex: A Unified Programming-by-Example Interaction for Synthesizing Readable Code for Data Scientists. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–12.
  12. 12.Will Epperson, Vaishnavi Gorantla, Dominik Moritz, and Adam Perer. 2023. Dead or Alive: Continuous Data Profiling for Interactive Data Science. IEEE Transactions on Visualization and Computer Graphics 30, 1 (2023), 197–207. https://doi.org/10.1109/TVCG.2023.3327367
  13. 13.Kasra Ferdowsi, Ruanqianqian Huang, Michael B James, Nadia Polikarpova, and Sorin Lerner. 2024. Validating AI-Generated Code with Live Programming. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–8.
  14. 14.Kasra Ferdowsi, Jack Williams, Ian Drosos, Andrew D. Gordon, Carina Negreanu, Nadia Polikarpova, Advait Sarkar, and Benjamin Zorn. 2023. COLDECO: An End User Spreadsheet Inspection Tool for AI-Generated Code. In Proceedings of IEEE Symposium on Visual Languages and Human-Centric Computing. IEEE Computer Society, Los Alamitos, CA, USA, 82–91. https://doi.org/10.1109/VL-HCC57772.2023.00017
  15. 15.Kasra Ferdowsifard, Shraddha Barke, Hila Peleg, Sorin Lerner, and Nadia Polikarpova. 2021. LooPy: Interactive Program Synthesis with Control Structures. Proceedings of the ACM on Programming Languages 5, OOPSLA, Article 153 (2021), 29 pages. https://doi.org/10.1145/3485530
  16. 16.Kasra Ferdowsifard, Allen Ordookhanians, Hila Peleg, Sorin Lerner, and Nadia Polikarpova. 2020. Small-Step Live Programming by Example. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology. ACM, New York, NY, USA, 614–626. https://doi.org/10.1145/3379337.3415869
  17. 17.Python Software Foundation. 2024. Full Grammar Specification of Python. https://docs.python.org/3/reference/grammar.html. Last accessed on 2024-11-09.
  18. 18.GitHub, Inc. 2024. GitHub Copilot. https://github.com/features/copilot. Last accessed on 2024-11-09.
  19. 19.Sumit Gulwani. 2011. Automating String Processing in Spreadsheets Using Input-Output Examples. In Proceedings of the ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages. ACM, New York, NY, USA, 317–330.
  20. 20.Philip J. Guo, Sean Kandel, Joseph M. Hellerstein, and Jeffrey Heer. 2011. Proactive Wrangling: Mixed-Initiative End-User Programming of Data Transformation Scripts. In Proceedings of the ACM Symposium on User Interface Software and Technology. ACM, New York, NY, USA, 65–74.
  21. 21.Tihomir Gvero, Viktor Kuncak, Ivan Kuraj, and Ruzica Piskac. 2013. Complete Completion Using Types and Weights. In Proceedings of the 34th ACM SIGPLAN Conference on Programming Language Design and Implementation. ACM, New York, NY, USA, 27–38. https://doi.org/10.1145/2491956.2462192
  22. 22.Sandra G. Hart and Lowell E. Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research. In Human Mental Workload. Advances in Psychology, Vol. 52. North-Holland, 139–183. https://doi.org/10.1016/S0166-4115(08)62386-9
  23. 23.Vincent J. Hellendoorn and Premkumar Devanbu. 2017. Are Deep Neural Networks the Best Choice for Modeling Source Code?. In Proceedings of the 11th Joint Meeting on Foundations of Software Engineering. ACM, New York, NY, USA, 763–773. https://doi.org/10.1145/3106237.3106290
  24. 24.R. Hill and J. Rideout. 2004. Automatic Method Completion. In Proceedings of the 19th International Conference on Automated Software Engineering. IEEE Computer Society, USA, 228–235.
  25. 25.Abram Hindle, Earl T. Barr, Zhendong Su, Mark Gabel, and Premkumar Devanbu. 2012. On the Naturalness of Software. In Proceedings of the 34th International Conference on Software Engineering. IEEE Press, Zurich, Switzerland, 837–847.
  26. 26.Joshua Horowitz and Jeffrey Heer. 2023. Engraft: An API for Live, Rich, and Composable Programming. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. ACM, New York, NY, USA, Article 72, 18 pages. https://doi.org/10.1145/3586183.3606733
  27. 27.Daqing Hou and David M. Pletcher. 2011. An Evaluation of the Strategies of Sorting, Filtering, and Grouping API Methods for Code Completion. In Proceedings of the 27th IEEE International Conference on Software Maintenance. IEEE Computer Society, USA, 233–242.
  28. 28.Yanwei Huang, Yunfan Zhou, Ran Chen, Changhao Pan, Xinhuan Shu, Di Weng, and Yingcai Wu. 2024. Interactive Table Synthesis with Natural Language. IEEE Transactions on Visualization and Computer Graphics 30, 9 (2024), 6130–6145. https://doi.org/10.1109/TVCG.2023.3329120
  29. 29.Naman Jain, Skanda Vaidyanath, Arun Iyer, Nagarajan Natarajan, Suresh Parthasarathy, Sriram Rajamani, and Rahul Sharma. 2022. Jigsaw: Large Language Models Meet Program Synthesis. In Proceedings of the 44th International Conference on Software Engineering. ACM, New York, NY, USA, 1219–1231. https://doi.org/10.1145/3510003.3510203
  30. 30.Zhongjun Jin, Michael R. Anderson, Michael J. Cafarella, and H. V. Jagadish. 2017. Foofah: Transforming Data By Example. In Proceedings of the ACM International Conference on Management of Data. ACM, New York, NY, USA, 683–698.
  31. 31.Sean Kandel, Andreas Paepcke, Joseph M. Hellerstein, and Jeffrey Heer. 2011. Wrangler: Interactive Visual Specification of Data Transformation Scripts. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 3363–3372.
  32. 32.Hyeonsu Kang and Philip J. Guo. 2017. Omnicode: A Novice-Oriented Live Programming Environment with Always-On Run-Time Value Visualizations. In Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology. ACM, New York, NY, USA, 737–745. https://doi.org/10.1145/3126594.3126632
  33. 33.Rafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton, and Andrea Janes. 2020. Big Code != Big Vocabulary: Open-Vocabulary Models for Source Code. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering. ACM, New York, NY, USA, 1073–1085. https://doi.org/10.1145/3377811.3380342
  34. 34.Stephen Kasica, Charles Berret, and Tamara Munzner. 2021. Table Scraps: An Actionable Framework for Multi-Table Data Wrangling From An Artifact Study of Computational Journalism. IEEE Transactions on Visualization and Computer Graphics 27, 2 (2021), 957–966. https://doi.org/10.1109/TVCG.2020.3030462
  35. 35.Majeed Kazemitabaar, Jack Williams, Ian Drosos, Tovi Grossman, Austin Zachary Henley, Carina Negreanu, and Advait Sarkar. 2024. Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. ACM, New York, NY, USA, Article 92, 19 pages. https://doi.org/10.1145/3654777.3676345
  36. 36.Mary Beth Kery, Donghao Ren, Fred Hohman, Dominik Moritz, Kanit Wongsuphasawat, and Kayur Patel. 2020. mage: Fluid Moves Between Code and Graphical Work in Computational Notebooks. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology. ACM, New York, NY, USA, 140–151.
  37. 37.Nodira Khoussainova, YongChul Kwon, Magdalena Balazinska, and Dan Suciu. 2010. SnipSuggest: Context-Aware Autocompletion for SQL. Proceedings of the VLDB Endowment 4, 1 (2010), 22–33. https://doi.org/10.14778/1880172.1880175
  38. 38.Klaus Krippendorff. 2018. Content Analysis: An Introduction to Its Methodology. SAGE Publications, Thousand Oaks, CA.
  39. 39.Sam Lau, Sruti Srinivasa Srinivasa Ragavan, Ken Milne, Titus Barik, and Advait Sarkar. 2021. TweakIt: Supporting End-User Programmers Who Transmogrify Code. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, Article 311, 12 pages. https://doi.org/10.1145/3411764.3445265
  40. 40.Yun Young Lee, Sam Harwell, Sarfraz Khurshid, and Darko Marinov. 2013. Temporal Code Completion and Navigation. In Proceedings of the 35th International Conference on Software Engineering. IEEE Press, San Francisco, CA, USA, 1181–1184.
  41. 41.Sorin Lerner. 2020. Projection Boxes: On-the-fly Reconfigurable Visualization for Live Programming. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–7. https://doi.org/10.1145/3313831.3376494
  42. 42.Haotian Li, Lu Ying, Haidong Zhang, Yingcai Wu, Huamin Qu, and Yun Wang. 2023. Notable: On-the-fly Assistant for Data Storytelling in Computational Notebooks. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, Article 173, 16 pages. https://doi.org/10.1145/3544548.3580965
  43. 43.Jian Li, Yue Wang, Michael R. Lyu, and Irwin King. 2018. Code Completion with Neural Attention and Pointer Networks. In Proceedings of the 27th International Joint Conference on Artificial Intelligence. AAAI Press, Stockholm, Sweden, 4159–4165.
  44. 44.Xingjun Li, Yizhi Zhang, Justin Leung, Chengnian Sun, and Jian Zhao. 2023. EDAssistant: Supporting Exploratory Data Analysis in Computational Notebooks with In Situ Code Search and Recommendation. ACM Transactions on Interactive Intelligent Systems 13, 1, Article 1 (2023), 27 pages. https://doi.org/10.1145/3545995
  45. 45.Jenny T. Liang, Chenyang Yang, and Brad A. Myers. 2024. A Large-Scale Survey on the Usability of AI Programming Assistants: Successes and Challenges. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering. ACM, New York, NY, USA, Article 52, 13 pages. https://doi.org/10.1145/3597503.3608128
  46. 46.Fang Liu, Ge Li, Bolin Wei, Xin Xia, Zhiyi Fu, and Zhi Jin. 2022. A Unified Multi-Task Learning Model for AST-Level and Token-Level Code Completion. Empirical Softw. Engg. 27, 4 (2022), 91. https://doi.org/10.1007/s10664-022-10140-7
  47. 47.Fang Liu, Ge Li, Yunfei Zhao, and Zhi Jin. 2021. Multi-Task Learning Based Pre-trained Language Model for Code Completion. In Proceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering. ACM, New York, NY, USA, 473–485. https://doi.org/10.1145/3324884.3416591
  48. 48.Zhongsu Luo, Kai Xiong, Jiajun Zhu, Ran Chen, Xinhuan Shu, Di Weng, and Yingcai Wu. 2025. Ferry: Toward Better Understanding of Input/Output Space for Data Wrangling Scripts. IEEE Transactions on Visualization and Computer Graphics 31, 1 (2025), 1202–1212. https://doi.org/10.1109/TVCG.2024.3456328
  49. 49.Irv Lustig and Princeton Consultants. 2014. Pandas Cheat Sheet. https://pandas.pydata.org/Pandas_Cheat_Sheet.pdf. Last accessed on 2024-11-09.
  50. 50.Andrew M Mcnutt, Chenglong Wang, Robert A Deline, and Steven M. Drucker. 2023. On the Design of AI-powered Code Assistants for Notebooks. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–16. https://doi.org/10.1145/3544548.3580940
  51. 51.Meta. 2024. Llama 3. https://llama.meta.com/docs/model-cards-and-prompt-formats/meta-llama-3/. Last accessed on 2024-11-09.
  52. 52.Microsoft. 2024. Data Wrangler. https://marketplace.visualstudio.com/items?itemName=ms-toolsai.datawrangler. Last accessed on 2024-11-09.
  53. 53.Microsoft. 2024. IntelliSense in Visual Studio Code. https://code.visualstudio.com/docs/editor/intellisense. Last accessed on 2024-11-09.
  54. 54.Hussein Mozannar, Gagan Bansal, Adam Fourney, and Eric Horvitz. 2024. Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, Article 142, 16 pages. https://doi.org/10.1145/3613904.3641936
  55. 55.Michael Muller, Ingrid Lange, Dakuo Wang, David Piorkowski, Jason Tsay, Q. Vera Liao, Casey Dugan, and Thomas Erickson. 2019. How Data Science Workers Work with Data: Discovery, Capture, Curation, Design, Creation. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–15.
  56. 56.Kirill Müller, Romain François, and Hadley Wickham. 2018. dplyr. https://dplyr.tidyverse.org/. Last accessed on 2024-11-09.
  57. 57.Anh Tuan Nguyen, Tung Thanh Nguyen, Hoan Anh Nguyen, Ahmed Tamrawi, Hung Viet Nguyen, Jafar Al-Kofahi, and Tien N. Nguyen. 2012. Graph-Based Pattern-Oriented, Context-Sensitive Source Code Completion. In Proceedings of the 34th International Conference on Software Engineering. IEEE Press, Zurich, Switzerland, 69–79.
  58. 58.Takayuki Omori, Hiroaki Kuwabara, and Katsuhisa Maruyama. 2012. A Study on Repetitiveness of Code Completion Operations. In Proceedings of the 28th IEEE International Conference on Software Maintenance. IEEE Computer Society, USA, 584–587.
  59. 59.The pandas development team. 2024. pandas-dev/pandas: Pandas 2.2.3. https://doi.org/10.5281/zenodo.13819579. Last accessed on 2024-11-09.
  60. 60.Daniel Perelman, Sumit Gulwani, Thomas Ball, and Dan Grossman. 2012. Type-Directed Completion of Partial Expressions. In Proceedings of the 33rd ACM SIGPLAN Conference on Programming Language Design and Implementation. ACM, New York, NY, USA, 275–286. https://doi.org/10.1145/2254064.2254098
  61. 61.Deepthi Raghunandan, Zhe Cui, Kartik Krishnan, Segen Tirfe, Shenzhi Shi, Tejaswi Darshan Shrestha, Leilani Battle, and Niklas Elmqvist. 2024. Lodestar: Supporting Rapid Prototyping of Data Science Workflows through Data-Driven Analysis Recommendations. Information Visualization 23, 1 (2024), 21–39. https://doi.org/10.1177/14738716231190429
  62. 62.Vijayshankar Raman and Joseph M Hellerstein. 2001. Potter’s Wheel: An Interactive Data Cleaning System. In Proceedings of the VLDB Endowment. Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 381–390.
  63. 63.Veselin Raychev, Pavol Bielik, and Martin Vechev. 2016. Probabilistic Model for Code with Decision Trees. In Proceedings of the ACM SIGPLAN International Conference on Object-Oriented Programming, Systems, Languages, and Applications. ACM, New York, NY, USA, 731–747. https://doi.org/10.1145/2983990.2984041
  64. 64.Veselin Raychev, Martin Vechev, and Eran Yahav. 2014. Code Completion with Statistical Language Models. In Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation. ACM, New York, NY, USA, 419–428. https://doi.org/10.1145/2594291.2594321
  65. 65.Romain Robbes and Michele Lanza. 2008. How Program History Can Improve Code Completion. In Proceedings of the 23rd IEEE/ACM International Conference on Automated Software Engineering. IEEE Computer Society, USA, 317–326.
  66. 66.Johan Rosenkilde. 2024. How GitHub Copilot is Getting Better at Understanding Your Code. https://github.blog/ai-and-ml/github-copilot/how-github-copilot-is-getting-better-at-understanding-your-code/. Last accessed on 2024-11-09.
  67. 67.Adam Rule, Aurélien Tabard, and James D. Hollan. 2018. Exploration and Explanation in Computational Notebooks. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–12. https://doi.org/10.1145/3173574.3173606
  68. 68.Bahador Saket, Alex Endert, and Çağatay Demiralp. 2019. Task-Based Effectiveness of Basic Visualizations. IEEE Transactions on Visualization and Computer Graphics 25, 7 (2019), 2505–2512. https://doi.org/10.1109/TVCG.2018.2829750
  69. 69.Advait Sarkar, Andrew D. Gordon, Carina Negreanu, Christian Poelitz, Sruti Srinivasa Ragavan, and Ben Zorn. 2022. What Is It Like to Program with Artificial Intelligence?. In Proceedings of the 33rd Annual Conference of the Psychology of Programming Interest Group. 127–153.
  70. 70.Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020. IntelliCode Compose: Code Generation Using Transformer. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. ACM, New York, NY, USA, 1433–1443. https://doi.org/10.1145/3368089.3417058
  71. 71.Alexey Svyatkovskiy, Ying Zhao, Shengyu Fu, and Neel Sundaresan. 2019. Pythia: AI-assisted Code Completion System. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. ACM, New York, NY, USA, 2727–2735. https://doi.org/10.1145/3292500.3330699
  72. 72.The pandas development team. 2024. Pandas API Reference. https://pandas.pydata.org/docs/reference/index.html. Last accessed on 2024-11-09.
  73. 73.Priyan Vaithilingam, Tianyi Zhang, and Elena L. Glassman. 2022. Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models. In Proceedings of Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–7. https://doi.org/10.1145/3491101.3519665
  74. 74.April Yi Wang, Will Epperson, Robert A DeLine, and Steven M Drucker. 2022. Diff in the Loop: Supporting Data Comparison in Exploratory Data Analysis. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1–10.
  75. 75.Chenglong Wang, Yu Feng, Rastislav Bodik, Isil Dillig, Alvin Cheung, and Amy J. Ko. 2021. Falx: Synthesis-Powered Visualization Authoring. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 106:1–106:15.
  76. 76.Chaozheng Wang, Junhao Hu, Cuiyun Gao, Yu Jin, Tao Xie, Hailiang Huang, Zhenyu Lei, and Yuetang Deng. 2023. How Practitioners Expect Code Completion?. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. ACM, New York, NY, USA, 1294–1306. https://doi.org/10.1145/3611643.3616280
  77. 77.Zijie J. Wang, David Munechika, Seongmin Lee, and Duen Horng Chau. 2024. SuperNOVA: Design Strategies and Opportunities for Interactive Visualization in Computational Notebooks. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, Article 304, 17 pages. https://doi.org/10.1145/3613905.3650848
  78. 78.Thomas Weber and Sven Mayer. 2024. From Computational to Conversational Notebooks. arXiv preprint (2024). https://doi.org/10.48550/ARXIV.2406.10636
  79. 79.Hadley Wickham, Davis Vaughan, Maximilian Girlich, and Posit Software. 2024. tidyr. https://tidyr.tidyverse.org/. Last accessed on 2024-11-09.
  80. 80.Liwenhan Xie, Chengbo Zheng, Haijun Xia, Huamin Qu, and Chen Zhu-Tian. 2024. WaitGPT: Monitoring and Steering Conversational LLM Agent in Data Analysis with On-the-Fly Code Visualization. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. ACM, New York, NY, USA, Article 119, 14 pages. https://doi.org/10.1145/3654777.3676374
  81. 81.Cong Yan and Yeye He. 2020. Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks. In Proceedings of the ACM SIGMOD International Conference on Management of Data. ACM, New York, NY, USA, 1539–1554. https://doi.org/10.1145/3318464.3389738
  82. 82.Litao Yan, Alyssa Hwang, Zhiyuan Wu, and Andrew Head. 2024. Ivie: Lightweight Anchored Explanations of Just-Generated Code. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, Article 140, 15 pages. https://doi.org/10.1145/3613904.3642239
  83. 83.Amy X. Zhang, Michael Muller, and Dakuo Wang. 2020. How do Data Science Workers Collaborate? Roles, Workflows, and Tools. In Proceedings of the ACM on Human-Computer Interaction. ACM, New York, NY, USA, 23 pages.
  84. 84.Chunqi Zhao, I-Chao Shen, Tsukasa Fukusato, Jun Kato, and Takeo Igarashi. 2022. ODEN: Live Programming for Neural Network Architecture Editing. In Proceedings of the 27th International Conference on Intelligent User Interfaces. ACM, New York, NY, USA, 392–404. https://doi.org/10.1145/3490099.3511120

Citation

MLA
Zhou, Y., et al. “Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts”. arXiv, 2025, http://arxiv.org/abs/2503.02639v1.
APA
Zhou, Y., Cai, X., Shi, Q., Huang, Y., Li, H., Qu, H., Weng, D., & Wu, Y. (2025). Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts. arXiv. http://arxiv.org/abs/2503.02639v1
Chicago
Zhou, Y., X. Cai, Q. Shi, et al. 2025. “Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts”. arXiv. http://arxiv.org/abs/2503.02639v1.
Harvard
Zhou, Y. et al. (2025) “Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2503.02639v1.
Vancouver
1. Zhou Y, Cai X, Shi Q, Huang Y, Li H, Qu H, Weng D, Wu Y (2025) Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts. arXiv

BibTeX

@article{zhou2025xavier,
  title = {Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts},
  author = {Zhou, Yunfan and Cai, Xiwen and Shi, Qiming and Huang, Yanwei and Li, Haotian and Qu, Huamin and Weng, Di and Wu, Yingcai},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2503.02639v1},
  eprint = {2503.02639}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF