DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation

Nitay CalderonEyal Ben-DavidAmir FederRoi Reichart

article2022ACL56 citations

Proposes DoCoGen, an unsupervised controllable generation framework that transforms multi-sentence texts across domains while preserving their task labels, enabling effective data augmentation for low-resource domain adaptation without requiring parallel text or target task annotations.

Listen

Modern natural language processing models frequently struggle when applied to new environments or domains where labeled training data is scarce or nonexistent. In real-world applications, collecting and annotating large volumes of text for every potential operational setting is costly, labor-intensive, and often impractical. When trained on limited data from a single domain, models tend to rely on domain-specific shortcuts and spurious correlations rather than true underlying patterns, leading to severe performance drops when deployed across different subject areas or unseen domains.

The article introduces and evaluates DoCoGen, a framework designed to generate domain-counterfactual text examples to improve domain adaptation in low-resource environments. The primary objective is to demonstrate that an unsupervised controllable text generation method can automatically transform existing labeled source-domain examples into coherent, label-preserving text in other domains without requiring parallel training data or task labels.

The researchers designed a two-step framework consisting of domain corruption and domain-oriented reconstruction. In the corruption stage, the system identifies and masks domain-specific terms based on domain affinity statistics. In the reconstruction stage, a generative language model uses learned domain orientation vectors to fill the masked gaps with terminology appropriate for the target domain while preserving the original task label. To evaluate this approach, the authors tested the generated data on binary sentiment classification across six review domains and multi-label intent prediction across 14 dialogue domains under two distinct settings: unsupervised adaptation, where unlabeled data from the target domain is available, and any-domain adaptation, where the target domain is completely unseen during training. Evaluations were conducted across dozens of transfer configurations using 100 labeled source examples.

The evaluation yielded several key findings. First, human intrinsic evaluation confirmed that DoCoGen generates coherent multi-sentence text, achieving a 93% domain relevance score and an 80% label preservation rate, substantially outperforming alternative generative approaches. Second, augmenting small labeled datasets with DoCoGen-generated text consistently outperformed standard transfer baselines and traditional text augmentation techniques, achieving average accuracy gains of 1.3% to 1.9% in sentiment tasks and 1.5% to 1.6% in intent classification. Third, combining DoCoGen with an existing state-of-the-art transfer model improved performance and reduced model variance by 42%, indicating significantly greater output stability. Finally, experiments revealed that the performance benefit of domain-counterfactual augmentation is most pronounced in low-resource regimes, tapering off as base models reach roughly 85% accuracy on larger labeled datasets.

These findings indicate that targeted counterfactual generation provides an effective, data-centric remedy for domain shift and overfitting in scarce-data scenarios. By replacing domain-specific keywords while keeping sentiment or intent intact, the approach forces models to rely on domain-general features rather than local artifacts. For organizations deploying language models, this strategy can lower annotation costs, shorten deployment timelines for new domains, and mitigate operational risks associated with brittle out-of-domain performance.

Decision-makers and engineering teams working in data-constrained settings should consider incorporating domain-counterfactual augmentation pipelines prior to fine-tuning downstream classifiers. For target-domain filtering, teams should evaluate domain-specific characteristics: using a domain-classifier filter proved beneficial for domain-distinct tasks like sentiment classification but caused minor degradation in conversational intent prediction where utterances naturally contain fewer domain-specific terms. Before wide operational rollout, practitioners should conduct pilot evaluations to assess whether the target task relies heavily on domain-specific syntax.

The study's primary limitations include input truncation to short text spans (96 tokens) due to computational constraints and diminishing returns once labeled data exceeds roughly 250 to 500 examples. Confidence in the reported results is high within the tested benchmark datasets and low-resource boundaries, though caution is warranted when applying the approach to long-form documents or tasks where domain vocabulary and task labels are deeply entangled.

arXiv: 2202.12350hitaytech/DoCoGen
Cover for DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation

Abstract

Natural language processing (NLP) algorithms have become very successful, but they still struggle when applied to out-of-distribution examples. In this paper we propose a controllable generation approach in order to deal with this domain adaptation (DA) challenge. Given an input text example, our DoCoGen algorithm generates a domain-counterfactual textual example (D-CON) – that is similar to the original in all aspects, including the task label, but its domain is changed to a desired one. Importantly, DoCoGen is trained using only unlabeled examples from multiple domains – no NLP task labels or parallel pairs of textual examples and their domain-counterfactuals are required. We show that DoCoGen can generate coherent counterfactuals consisting of multiple sentences. We use the D-CONs generated by DoCoGen to augment a sentiment classifier and a multi-label intent classifier in 20 and 78 DA setups, respectively, where source-domain labeled data is scarce. Our model outperforms strong baselines and improves the accuracy of a state-of-the-art unsupervised DA algorithm.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Domain-Counterfactual Examples
  • 4 DoCoGen: Domain Counterfactual Generation
  • 4.1 Domain Corruption
  • 4.2 Domain-Oriented Reconstruction
  • 4.3 Filtering Mechanism
  • 5 Intrinsic Evaluation
  • 6 Experimental Setup
  • 6.1 Tasks and Domains
  • 6.2 Models and Baselines
  • 7 Results
  • 8 Conclusions
  • Acknowledgements
  • References
  • A Additional Generated Examples
  • B Implementation Details
  • B.1 URLs of Code and Data
  • B.2 Hyperparameters and Setups
  • B.3 Masking
  • C Ablation Results
  • C.1 Standard Deviations

Knowls

  1. Knowl 1 — Domain-counterfactual text preserves the task label while changing domain

    definition

    A domain-counterfactual example (D-CON) for a text xx is a coherent, human-like text x′x' that changes the domain of xx to a specified destination domain D′D' while aiming to preserve its other relevant properties, especially its task label. If (x,y)(x,y) is a labeled example from a source domain and YY denotes the task label, the intended target is a sample x′∼PD′(X∣Y=y)x' \sim P_{D'}(X\mid Y=y): text from domain D′D' conditional on the original label yy. Thus, for example, a negative review should remain negative when transformed from a kitchen-products review to an electronics review. DoCoGen produces such examples automatically from unlabeled multi-domain text, without task-label supervision or parallel original–counterfactual pairs.

  2. Knowl 2 — Domain affinity determines which n-grams are masked

    equation

    For an n-gram ww and domain DD, let cw,Dc_{w,D} be the number of unlabeled examples in DD containing ww, let nDn_D be the number of unlabeled examples in DD, and let NN be the number of domains. With equal domain priors, DoCoGen estimates the domain posterior by normalizing the smoothed within-domain frequencies: qD(w)=(cw,D+αn)/nD∑j=1N(cw,j+αn)/njq_D(w)=\frac{(c_{w,D}+\alpha_n)/n_D}{\sum_{j=1}^{N}(c_{w,j}+\alpha_n)/n_j}, where αn\alpha_n is an order-specific smoothing constant. Define H(D∣w)=−∑j=1Nqj(w)log⁡qj(w)H(D\mid w)=-\sum_{j=1}^{N}q_j(w)\log q_j(w) and the affinity ρ(w,D)=qD(w)(1−H(D∣w)log⁡N)\rho(w,D)=q_D(w)\left(1-\frac{H(D\mid w)}{\log N}\right). The masking score for moving from source domain DD to destination domain D′D' is m(w,D,D′)=ρ(w,D)−ρ(w,D′)m(w,D,D')=\rho(w,D)-\rho(w,D'). DoCoGen masks an n-gram when m(w,D,D′)>τm(w,D,D')>\tau: source-associated n-grams are favored for masking, while terms already associated with the destination are protected. In the reported experiments, n-grams were stemmed, only those found in at least 10 unlabeled examples were scored, τ=0.08\tau=0.08, and αn\alpha_n was 1, 5, and 7 for unigrams, bigrams, and trigrams, respectively.

  3. Knowl 3 — DoCoGen generates text by domain corruption and reconstruction

    model/method

    DoCoGen takes an input text xx, its source domain DD, and a destination domain D′D'. It first masks domain-specific spans: unigrams are considered before bigrams and then trigrams, and a longer span is considered only if it does not contain a previously masked shorter span; the masking decision uses the domain-difference score and threshold defined above. The resulting masked text is encoded by a T5 encoder–decoder together with a domain-orientation vector for D′D', concatenated with the token embeddings. At inference, the model generates a reconstruction intended to retain the input's task-relevant content while expressing it in the destination domain. Decoding uses beam search with beam size 4 and restricts output tokens to tokens from the original input or tokens that meet the model's destination-domain relevance criterion. For example, the airline text “The entertainment system failed twice but the crew reactivated it quick” can be transformed into kitchen-related text such as “The heating system failed twice but the thermostat reactivated it quick”; a term such as “system” can remain when it is also appropriate to the destination. The method is designed for coherent multi-sentence texts, although preservation of domain and label is an empirical goal rather than a guarantee.

  4. Knowl 4 — Learned orientation vectors provide unsupervised domain control

    model/method

    DoCoGen augments the pretrained T5 embedding matrix with KK learnable orientation vectors for each domain; the experiments use K=4K=4. The vectors are initialized from T5 embeddings of the domain name and the domain's top K−1K-1 representative words, ranked using log⁡(cw,D+1)ρ(w,D)\log(c_{w,D}+1)\rho(w,D). During unsupervised training, each unlabeled text xx is corrupted using a randomly selected destination different from its actual domain DD, while the reconstruction target remains the original text xx. The conditioning vector is selected for DD, randomly choosing the domain-name vector or a representative-word vector whose word occurs in xx. This reconstruction objective trains the vectors to condition generation on domain semantics without labeled task examples or manually paired D-CONs; selecting different vectors for one destination also permits varied outputs. The reported model uses T5-base, trains for 20 epochs with AdamW at learning rate 5×10−55\times10^{-5} and weight decay 10−510^{-5}, and selects the checkpoint with the highest destination-domain accuracy on generated held-out examples.

  5. Knowl 5 — F-DoCoGen filters generated examples by domain and basic quality

    model/method

    F-DoCoGen is DoCoGen followed by a quality filter. A domain classifier is trained on human-written unlabeled examples with their domain identities. A generated D-CON is retained only if this classifier predicts the requested destination domain, the text contains at least four words, and its word overlap with the original text is at least 25%. The filter targets examples that fail to change domain or have too little lexical continuity with their source; it can also remove valid examples when the text is not strongly domain-specific.

  6. Knowl 6 — Evaluation covers low-resource UDA and unseen-domain adaptation

    experimental setup

    The paper evaluates binary sentiment classification and multi-label intent prediction with 100 labeled examples in each source-domain training set and averages results over 25 random seeds and training-set samples. Sentiment uses reviews from six domains: Airline (A), DVD (D), Electronics (E), Kitchen (K), Books (B), and Restaurant (R). The four domains A, D, E, and K provide unlabeled data for adaptation; the study has 12 directed UDA source–target pairs among these domains and 8 ADA setups in which the target is B or R. For sentiment augmentation, each labeled example receives four generated examples per each of the four unlabeled domains, or 16 D-CONs. Intent prediction uses MANtIS conversation utterances with the five most common of its eight intent labels. Six domains—Apple (AP), DBA (DB), Electronics (EL), Physics (PH), Statistics (ST), and askubuntu (UB)—are used as adaptation domains, yielding 30 UDA pairs with distinct source and target domains; the other eight MANtIS domains are ADA targets, yielding 48 setups. Sentiment is scored by accuracy and intent prediction by macro-F1. Task classifiers use a T5 encoder and a linear output layer, except PERL, which uses BERT.

  7. Knowl 7 — Human ratings find DoCoGen outputs more fluent and label-preserving than a VAE baseline

    empirical result

    In an intrinsic evaluation, five nearly native English speakers rated 60 DoCoGen D-CONs generated from 20 original reviews across four domains, alongside VAE-generated reviews and original reviews. Domain relevance (D.REL) and label preservation (L.PRES) are percentages; linguistic acceptability (ACCPT) is on a 1–5 scale; WER is the reported minimum number of word substitutions, deletions, and insertions needed to make an example logical and grammatical. DoCoGen scored 93.0 D.REL, 80.0 L.PRES, 4.01 ACCPT, and 0.17 WER; the VAE scored 90.0, 46.0, 2.11, and 0.54, respectively. Original reviews scored 99.0, 88.0, 4.73, and 0.10. These ratings indicate that DoCoGen usually shifts text to the intended domain and preserves its label more often than the VAE baseline, while its fluency ratings approach those of human-written reviews.

  8. Knowl 8 — D-CON augmentation improves sentiment accuracy, and complements PERL in UDA

    empirical result

    Across 12 sentiment UDA setups, F-DoCoGen averages 81.8% accuracy, compared with 79.9% for the strongest non-augmentation baseline, DANN; it beats the baselines in 10 of 12 setups. DoCoGen without filtering averages 81.5%. Across 8 sentiment ADA setups, F-DoCoGen averages 82.8%, compared with 81.5% for the strongest baseline, RM-RR, and beats all baselines in every setup; unfiltered DoCoGen averages 82.1%. F-DoCoGen also outperforms the strongest orientation-vector ablation, RM-OV, by reported average error reductions of 11.2% in UDA and 5.0% in ADA. In UDA, combining D-CON augmentation with PERL gives 82.6% average accuracy versus 81.9% for PERL alone, outperforming PERL in 8 of 12 setups; their reported average standard deviations are 2.1 and 3.6, respectively. For F-DoCoGen, the reported average standard deviation improves by 22.0% in UDA and 27.5% in ADA over the best-performing baseline, indicating greater stability under these experimental conditions.

  9. Knowl 9 — DoCoGen yields the best average intent-prediction F1 among non-oracle methods

    empirical result

    On MANtIS intent prediction, unfiltered DoCoGen is the strongest non-oracle model across all evaluated source–target setups. Its mean macro-F1 is 75.4 in UDA and 74.6 in ADA, compared with 73.8 and 73.1 for the strongest baseline, DANN—gains of 1.6 and 1.5 points, respectively. F-DoCoGen averages 75.1 in UDA and 74.4 in ADA, still above every baseline but slightly below unfiltered DoCoGen. The paper attributes this small filtering penalty to intent examples that are not domain-specific and are therefore liable to be rejected by the domain filter. Relative to the stronger generation ablation, RM-OV, DoCoGen achieves reported error reductions of 8.0% in UDA and 7.6% in ADA.

  10. Knowl 10 — Augmentation gains diminish once the unaugmented classifier reaches a performance plateau

    empirical result

    A training-size analysis in sentiment classification varied the labeled source training set from 25 to 1,000 examples for Electronics and Kitchen, evaluating both UDA and ADA. The accuracy benefit from F-DoCoGen augmentation diminished when the unaugmented classifier exceeded approximately 85% accuracy and approached a plateau. This pattern is consistent with the paper's hypothesis that domain-counterfactual augmentation is most useful in low-resource conditions, where classifiers may rely more heavily on domain-specific correlations.

Coverage note — The full source–target score breakdowns, complete stability tables, and additional generated examples are omitted; aggregate comparisons and the principal stability results capture the main contribution without reproducing extensive appendix data.

References

  1. 1.Eyal Ben-David, Nadav Oved, and Roi Reichart. 2021. PADA: A prompt-based autoregressive approach for adaptation to unseen domains. CoRR, abs/2102.12206.
  2. 2.Eyal Ben-David, Carmel Rabinovitz, and Roi Reichart. 2020. PERL: pivot-based domain adaptation for pre-trained deep contextualized embedding models. Trans. Assoc. Comput. Linguistics, 8:504–521.
  3. 3.John Blitzer, Mark Dredze, and Fernando Pereira. 2007. Biographies, bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification. In ACL 2007, Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics, June 23-30, 2007, Prague, Czech Republic. The Association for Computational Linguistics.
  4. 4.John Blitzer, Ryan T. McDonald, and Fernando Pereira. 2006. Domain adaptation with structural correspondence learning. In EMNLP 2006, Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing, 22-23 July 2006, Sydney, Australia, pages 120–128. ACL.
  5. 5.Samuel R. Bowman, Luke Vilnis, Oriol Vinyals, Andrew M. Dai, Rafal Józefowicz, and Samy Bengio. 2016. Generating sentences from a continuous space. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, CoNLL 2016, Berlin, Germany, August 11-12, 2016, pages 10–21. ACL.
  6. 6.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  7. 7.Jiaao Chen, Derek Tam, Colin Raffel, Mohit Bansal, and Diyi Yang. 2021. An empirical survey of data augmentation for limited data learning in NLP. CoRR, abs/2106.07499.
  8. 8.Minmin Chen, Zhixiang Eddie Xu, Kilian Q. Weinberger, and Fei Sha. 2012. Marginalized denoising autoencoders for domain adaptation. In Proceedings of the 29th International Conference on Machine Learning, ICML 2012, Edinburgh, Scotland, UK, June 26 - July 1, 2012. icml.cc / Omnipress.
  9. 9.Hal Daumé III and Daniel Marcu. 2006. Domain adaptation for statistical classifiers. J. Artif. Intell. Res., 26:101–126.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics.
  11. 11.Sergey Edunov, Michael Ott, Michael Auli, and David Grangier. 2018. Understanding back-translation at scale. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pages 489–500. Association for Computational Linguistics.
  12. 12.Amir Feder, Katherine A. Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E. Roberts, Brandon M. Stewart, Victor Veitch, and Diyi Yang. 2021a. Causal inference in natural language processing: Estimation, prediction, interpretation and beyond. CoRR, abs/2109.00725.
  13. 13.Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart. 2021b. Causalm: Causal model explanation through counterfactual language models. Computational Linguistics, 47(2):333–386.
  14. 14.Steven Y. Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard H. Hovy. 2021. A survey of data augmentation approaches for NLP. In Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, volume ACL/IJCNLP 2021 of Findings of ACL, pages 968–988. Association for Computational Linguistics.
  15. 15.Steven Y. Feng, Aaron W. Li, and Jesse Hoey. 2019. Keep calm and switch on! preserving sentiment and fluency in semantic text exchange. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 2701–2711. Association for Computational Linguistics.
  16. 16.Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor S. Lempitsky. 2016. Domain-adversarial training of neural networks. The journal of machine learning research, 17:59:1–59:35.
  17. 17.Matt Gardner, William Merrill, Jesse Dodge, Matthew E. Peters, Alexis Ross, Sameer Singh, and Noah A. Smith. 2021. Competency problems: On finding and removing artifacts in language data. CoRR, abs/2104.08646.
  18. 18.Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard S. Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. 2020. Shortcut learning in deep neural networks. CoRR, abs/2004.07780.
  19. 19.Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011. Domain adaptation for large-scale sentiment classification: A deep learning approach. In Proceedings of the 28th International Conference on Machine Learning, ICML 2011, Bellevue, Washington, USA, June 28 - July 2, 2011, pages 513–520. Omnipress.
  20. 20.Xiaochuang Han and Jacob Eisenstein. 2019. Unsupervised domain adaptation of contextualized embeddings for sequence labeling. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 4237–4247. Association for Computational Linguistics.
  21. 21.Jiayuan Huang, Alexander J. Smola, Arthur Gretton, Karsten M. Borgwardt, and Bernhard Schölkopf. 2006. Correcting sample selection bias by unlabeled data. In Advances in Neural Information Processing Systems 19, Proceedings of the Twentieth Annual Conference on Neural Information Processing Systems, Vancouver, British Columbia, Canada, December 4-7, 2006, pages 601–608. MIT Press.
  22. 22.Touseef Iqbal and Shaima Qureshi. 2020. The survey: Text generation models in deep learning. Journal of King Saud University-Computer and Information Sciences.
  23. 23.Nitish Joshi and He He. 2021. An investigation of the (in)effectiveness of counterfactually augmented data. CoRR, abs/2107.00753.
  24. 24.Divyansh Kaushik, Eduard H. Hovy, and Zachary Chase Lipton. 2020. Learning the difference that makes A difference with counterfactually-augmented data. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  25. 25.Divyansh Kaushik, Amrith Setlur, Eduard H. Hovy, and Zachary Chase Lipton. 2021. Explaining the efficacy of counterfactually augmented data. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  26. 26.Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019. CTRL: A conditional transformer language model for controllable generation. CoRR, abs/1909.05858.
  27. 27.Daniel Khashabi, Tushar Khot, and Ashish Sabharwal. 2020. More bang for your buck: Natural perturbation for robust question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 163–170. Association for Computational Linguistics.
  28. 28.Sosuke Kobayashi. 2018. Contextual augmentation: Data augmentation by words with paradigmatic relations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 2 (Short Papers), pages 452–457. Association for Computational Linguistics.
  29. 29.Ashutosh Kumar, Satwik Bhattamishra, Manik Bhandari, and Partha P. Talukdar. 2019. Submodular optimization-based diverse paraphrasing and its effectiveness in data augmentation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 3609–3619. Association for Computational Linguistics.
  30. 30.Entony Lekhtman, Yftah Ziser, and Roi Reichart. 2021. DILBERT: customized pre-training for domain adaptation withcategory shift, with an application to aspect extraction. CoRR, abs/2109.00571.
  31. 31.Jiwei Li, Michel Galley, Chris Brockett, Georgios P. Spithourakis, Jianfeng Gao, and William B. Dolan. 2016. A persona-based neural conversation model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Volume 1: Long Papers. The Association for Computer Linguistics.
  32. 32.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  33. 33.Clara Meister, Ryan Cotterell, and Tim Vieira. 2020. If beam search is the answer, what was the question? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 2173–2185. Association for Computational Linguistics.
  34. 34.Nathan Ng, Kyunghyun Cho, and Marzyeh Ghassemi. 2020. SSMBA: self-supervised manifold based data augmentation for improving out-of-domain robustness. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 1268–1283. Association for Computational Linguistics.
  35. 35.Quang Nguyen. 2015. The airline review dataset.
  36. 36.Sinno Jialin Pan, Xiaochuan Ni, Jian-Tao Sun, Qiang Yang, and Zheng Chen. 2010. Cross-domain sentiment classification via spectral feature alignment. In Proceedings of the 19th International Conference on World Wide Web, WWW 2010, Raleigh, North Carolina, USA, April 26-30, 2010, pages 751–760. ACM.
  37. 37.Gustavo Penha, Alexandru Balan, and Claudia Hauff. 2019. Introducing mantis: a novel multi-domain information seeking dialogues dataset. CoRR, abs/1912.04639.
  38. 38.Maria Pontiki, Dimitris Galanis, Haris Papageorgiou, Ion Androutsopoulos, Suresh Manandhar, Mohammad Al-Smadi, Mahmoud Al-Ayyoub, Yanyan Zhao, Bing Qin, Orphée De Clercq, Véronique Hoste, Marianna Apidianaki, Xavier Tannier, Natalia V. Loukachevitch, Evgeniy V. Kotelnikov, Núria Bel, Salud María Jiménez Zafra, and Gülsen Eryigit. 2016. Semeval-2016 task 5: Aspect based sentiment analysis. In Proceedings of the 10th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2016, San Diego, CA, USA, June 16-17, 2016, pages 19–30. The Association for Computer Linguistics.
  39. 39.Shrimai Prabhumoye, Alan W. Black, and Ruslan Salakhutdinov. 2020. Exploring controllable text generation techniques. In Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, pages 1–14. International Committee on Computational Linguistics.
  40. 40.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  41. 41.Roi Reichart and Ari Rappoport. 2007. Self-training for enhancement and domain adaptation of statistical parsers trained on small datasets. In ACL 2007, Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics, June 23-30, 2007, Prague, Czech Republic. The Association for Computational Linguistics.
  42. 42.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 3980–3990. Association for Computational Linguistics.
  43. 43.Brian Roark and Michiel Bacchiani. 2003. Supervised and unsupervised PCFG adaptation to novel domains. In Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics, HLT-NAACL 2003, Edmonton, Canada, May 27 - June 1, 2003. The Association for Computational Linguistics.
  44. 44.Daniel Rosenberg, Itai Gat, Amir Feder, and Roi Reichart. 2021. Are VQA systems rad? measuring robustness to augmented data with focused interventions. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 2: Short Papers), Virtual Event, August 1-6, 2021, pages 61–70. Association for Computational Linguistics.
  45. 45.Guy Rotman and Roi Reichart. 2019. Deep contextualized self-training for low resource dependency parsing. Trans. Assoc. Comput. Linguistics, 7:695–713.
  46. 46.Alexander M Rush, Roi Reichart, Michael Collins, and Amir Globerson. 2012. Improved parsing and pos tagging using inter-sentence consistency constraints. In Proceedings of the 2012 joint conference on empirical methods in natural language processing and computational natural language learning, pages 1434–1444.
  47. 47.Giuseppe Russo, Nora Hollenstein, Claudiu Cristian Musat, and Ce Zhang. 2020. Control, generate, augment: A scalable framework for multi-attribute text generation. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, volume EMNLP 2020 of Findings of ACL, pages 351–366. Association for Computational Linguistics.
  48. 48.Indira Sen, Mattia Samory, Fabian Flöck, Claudia Wagner, and Isabelle Augenstein. 2021. How does counterfactually augmented data impact models for social computing constructs? CoRR, abs/2109.07022.
  49. 49.Dinghan Shen, Asli Celikyilmaz, Yizhe Zhang, Liqun Chen, Xin Wang, Jianfeng Gao, and Lawrence Carin. 2019. Towards generating long and coherent text with multi-level latent variable models. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 2079–2089. Association for Computational Linguistics.
  50. 50.Victor Veitch, Alexander D’Amour, Steve Yadlowsky, and Jacob Eisenstein. 2021. Counterfactual invariance to spurious correlations: Why and how to pass stress tests. CoRR, abs/2106.00545.
  51. 51.Tianlu Wang, Xuezhi Wang, Yao Qin, Ben Packer, Kang Li, Jilin Chen, Alex Beutel, and Ed Chi. 2020. Cat-gen: Improving robustness in NLP models via controlled adversarial text generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 5141–5146. Association for Computational Linguistics.
  52. 52.Zhao Wang and Aron Culotta. 2020. Identifying spurious correlations for robust text classification. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, volume EMNLP 2020 of Findings of ACL, pages 3431–3440. Association for Computational Linguistics.
  53. 53.Zirui Wang, Adams Wei Yu, Orhan Firat, and Yuan Cao. 2021. Towards zero-label language learning. CoRR, abs/2109.09193.
  54. 54.Jason W. Wei and Kai Zou. 2019. EDA: easy data augmentation techniques for boosting performance on text classification tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 6381–6387. Association for Computational Linguistics.
  55. 55.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, EMNLP 2020 - Demos, Online, November 16-20, 2020, pages 38–45. Association for Computational Linguistics.
  56. 56.Dustin Wright and Isabelle Augenstein. 2020. Transformer based multi-source domain adaptation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 7963–7974. Association for Computational Linguistics.
  57. 57.Tongshuang Wu, Marco Túlio Ribeiro, Jeffrey Heer, and Daniel S. Weld. 2021. Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 6707–6723. Association for Computational Linguistics.
  58. 58.Qizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong, and Quoc Le. 2020. Unsupervised data augmentation for consistency training. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  59. 59.Yi Yang and Jacob Eisenstein. 2014. Fast easy unsupervised domain adaptation with marginalized structured dropout. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics, ACL 2014, June 22-27, 2014, Baltimore, MD, USA, Volume 2: Short Papers, pages 538–544. The Association for Computer Linguistics.
  60. 60.Jianfei Yu, Chenggong Gong, and Rui Xia. 2021. Cross-domain review generation for aspect-based sentiment analysis. In Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, volume ACL/IJCNLP 2021 of Findings of ACL, pages 4767–4777. Association for Computational Linguistics.
  61. 61.Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, pages 649–657.
  62. 62.Yftah Ziser and Roi Reichart. 2017. Neural structural correspondence learning for domain adaptation. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), Vancouver, Canada, August 3-4, 2017, pages 400–410. Association for Computational Linguistics.
  63. 63.Yftah Ziser and Roi Reichart. 2018a. Deep pivot-based modeling for cross-language cross-domain transfer with minimal guidance. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pages 238–249. Association for Computational Linguistics.
  64. 64.Yftah Ziser and Roi Reichart. 2018b. Pivot based language modeling for improved neural domain adaptation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 1241–1251. Association for Computational Linguistics.
  65. 65.Yftah Ziser and Roi Reichart. 2019. Task refinement learning for improved accuracy and stability of unsupervised domain adaptation. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 5895–5906. Association for Computational Linguistics.

Citation

MLA
Calderon, N., et al. “DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 7727–46, https://doi.org/10.18653/v1/2022.acl-long.533.
APA
Calderon, N., Ben-David, E., Feder, A., & Reichart, R. (2022). DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7727–7746. https://doi.org/10.18653/v1/2022.acl-long.533
Chicago
Calderon, N., E. Ben-David, A. Feder, and R. Reichart. 2022. “DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7727–46. https://doi.org/10.18653/v1/2022.acl-long.533.
Harvard
Calderon, N. et al. (2022) “DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7727–7746. Available at: https://doi.org/10.18653/v1/2022.acl-long.533.
Vancouver
1. Calderon N, Ben-David E, Feder A, Reichart R (2022) DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7727–7746

BibTeX

@inproceedings{calderon-etal-2022-docogen,
    title = "{D}o{C}o{G}en: {D}omain Counterfactual Generation for Low Resource Domain Adaptation",
    author = "Calderon, Nitay  and
      Ben-David, Eyal  and
      Feder, Amir  and
      Reichart, Roi",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.533/",
    doi = "10.18653/v1/2022.acl-long.533",
    pages = "7727--7746"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/