PAED: Zero-Shot Persona Attribute Extraction in Dialogues

Luyao ZhuWei LiRui MaoVlad PandeleaErik Cambria

article2023ACL48 citations

Introduces a refined persona attribute triplet dataset alongside a generation-based framework that uses a variational autoencoder hard negative sampling strategy to extract structured user personas in zero-shot dialogue settings.

Listen

Personalized dialogue systems rely on extracting user profiles and preferences directly from conversational interactions. Existing benchmarks for structured persona extraction often suffer from ambiguous relation categories, missing annotations, and noisy automated pairing. Furthermore, in real-world deployments, models routinely encounter persona traits and relations that were never seen during training, creating a generalized zero-shot extraction challenge where subtle linguistic shifts can completely reverse meaning.

The article establishes a high-quality persona attribute extraction dataset and introduces a generalized zero-shot extraction framework designed to accurately recognize unseen persona relations from dialogue utterances.

The researchers developed the PersonaExt dataset by refining previous multi-turn conversation corpora. They manually corrected 1,896 triplet labels, refined 6,357 utterance-triplet pairs across 105 distinct relation categories, and applied a strict consensus filtering rule requiring dual classifier agreement. To extract unseen relations, the authors implemented a generation-based framework utilizing a prompt-tuned generator to synthesize training instances for novel relations, alongside an extractor model. The extractor is enhanced with a Meta-Variational Autoencoder sampler and a contrastive structured constraint loss to identify and separate semantically similar, easily confused dialogue samples.

The evaluation demonstrates that the proposed framework consistently outperforms existing baselines across zero-shot benchmark settings. On the PersonaExt dataset, it achieved an average accuracy gain of 1.06 percentage points over the strongest baseline. In multi-triplet extraction on the FewRel benchmark, it outperformed baseline approaches by 3.18 percentage points in F1 score and achieved a 3.22 percentage point higher precision, substantially reducing false positive errors. Ablation experiments further revealed that the Meta-Variational Autoencoder hard negative sampling strategy exceeded alternative samplers by an average of 2.66 percentage points in extraction accuracy.

These results show that handling hard negative samples and enforcing rigorous structured contrastive learning allows models to generalize reliably to novel attributes without extensive retraining. In production conversational agents, higher precision in persona extraction significantly reduces the risk of generating inaccurate, hallucinated, or contradictory user profiles, thereby improving user experience and conversational consistency.

Organizations developing personalized conversational artificial intelligence should adopt conservative label-matching protocols and contrastive extraction frameworks when handling user attributes. When deploying zero-shot extractors, teams should utilize greedy or structured search decoding rather than open-ended sampling strategies to preserve extraction accuracy.

Confidence in these findings is high for single-turn explicit persona extractions across predefined benchmarks. However, key operational limitations remain: the framework is evaluated primarily on explicit statements and does not currently extract multi-sentence contextual attributes or resolve speaker identities when complex pronoun references span multiple conversational turns.

Cover for PAED: Zero-Shot Persona Attribute Extraction in Dialogues

Abstract

Persona attribute extraction is critical for personalized human-computer interaction. Dialogue is an important medium that communicates and delivers persona information. Although there is a public dataset for triplet-based persona attribute extraction from conversations, its automatically generated labels present many issues, including unspecific relations and inconsistent annotations. We fix such issues by leveraging more reliable text-label matching criteria to generate high-quality data for persona attribute extraction. We also propose a contrastive learning- and generation-based model with a novel hard negative sampling strategy for generalized zero-shot persona attribute extraction. We benchmark our model with state-of-the-art baselines on our dataset and a public dataset, showing outstanding accuracy gains. Our sampling strategy also exceeds others by a large margin in persona attribute extraction.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Persona Extraction
  • 2.2 Relation Triplet Extraction
  • 2.3 Hard Negative Sampling
  • 3 PersonaExt Construction
  • 3.1 Automatic Intersection Assignment
  • 3.2 Attribute Triplet Label Correction
  • 4 Generalized Zero-Shot PAED
  • 4.1 Task Definition
  • 4.2 Persona Attribute Generator
  • 4.3 Persona Attribute Extractor
  • 4.4 Meta-VAE Sampler
  • 4.4.1 Meta-VAE
  • 4.4.2 Sampling Criteria
  • 4.5 Contrastive Structured Constraint
  • 5 Experiments
  • 5.1 Datasets
  • 5.2 Baselines
  • 5.3 Setups
  • 5.4 Experimental Results
  • 5.5 Ablation Study
  • 5.6 Revisiting Meta-VAE Sampler with CSC
  • 5.7 Case Study
  • 5.8 Exploration of Experimental Settings
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • Ethics Statement
  • References
  • A Performance of finetuning with Meta-VAE sampler and CSC
  • B Implementation details
  • C Meta-VAE Sampling
  • D Dataset Annotation
  • D.1 Statistics of Annotated Sentences
  • D.2 Relation Types
  • D.3 Annotation Rules for Selected Relation Types
  • E Discussion of data and code

Knowls

  1. Knowl 1 — PersonaExt construction and annotation correction

    model/method

    PersonaExt is a dialogue persona-attribute extraction dataset built from PersonaChat dialogues and the triplet annotations in Dialogue NLI. A triplet is (s,r,o)(s,r,o), where ss is the subject, rr the persona relation, and oo the object. To reduce false labels, a persona triplet is assigned to a dialogue utterance only when both a fine-tuned BERT classifier and a TF-IDF classifier predict entailment between the utterance and the persona sentence (or source utterance). This intersection rule is deliberately conservative: it can omit valid labels, but avoids assigning a triplet based on only one classifier.

    The creators manually corrected 1,896 persona sentences, focusing on negations and underspecified relation types such as other, have, and like. The corrected labels distinguish, for example, never_do from have_no, and retain quantities in objects. Corrected persona triplets were propagated to utterances using the two-classifier rule; Snowball stemming was then used to remove over-annotations and align triplet entities with the utterance. In total, 6,357 utterance–triplet pairs were processed.

  2. Knowl 2 — Human evaluation of PersonaExt annotation quality

    empirical result

    PersonaExt's utterance–triplet annotations were evaluated against those of Wu et al. (2020) on 150 randomly selected utterances. Two English-speaking graduate students scored relation and object annotations separately for consistency and specificity on a binary scale. PersonaExt scored higher on every dimension:

    Dataset Object consistency Object specificity Relation consistency Relation specificity
    Wu et al. (2020) 0.70 0.68 0.68 0.54
    PersonaExt 0.97 0.95 0.89 0.83

    For label correction, one experienced native-English annotator produced annotations and two side annotators reviewed them. Side-annotator Cohen's κ\kappa was 0.72; both side annotators supported 82.8% of the main annotator's annotations, and at least one supported 89.4% of newly generated annotations. The final dataset used a newly generated label when at least one side annotator supported it, and otherwise retained the Dialogue NLI label.

  3. Knowl 3 — Generalized zero-shot PAED framework

    model/method

    The paper formulates persona attribute extraction in dialogues (PAED) as generalized zero-shot learning. Let D=(U,Y)D=(U,Y) be a dataset of dialogue utterances UU and persona triplets YY, with each triplet y=(s,r,o)y=(s,r,o). Training relations RsR_s and test relations RtR_t are disjoint, but the test relation set is available during training. At inference, the extractor searches over both seen and unseen relations, Rs∪RtR_s\cup R_t, so a test utterance may be assigned either a training relation or an unseen test relation.

    The framework has a persona attribute generator (PAG) and a persona attribute extractor (PAE). PAG is first trained on seen data and then prompt-tuned with unseen relation labels to generate structured synthetic examples from a RELATION: r prompt, producing a context, subject, and object. PAE is first trained on seen data and then trained on PAG's synthetic examples; given a CONTEXT: u prompt, it generates subject, object, and relation. PAG uses causal next-token language modeling, while PAE uses sequence-to-sequence training.

  4. Knowl 4 — Relation-conditioned Meta-VAE

    model/method

    Meta-VAE is a shared variational autoencoder used to represent the utterance distribution associated with each persona relation without training a separate VAE for every relation. It maps a relation label to an embedding with a learned fully connected layer, conditions the utterance encoder on that relation embedding, and uses GRUs as encoder and decoder. The encoder produces a Gaussian approximate posterior over a continuous latent variable; the prior is a standard normal. Training combines a KL-divergence term between posterior and prior with an utterance reconstruction, or next-token generation, term.

    The model has 2.4 million parameters. The reported configuration uses a hidden-state dimension of 100, latent dimension of 128, two encoder layers, two decoder layers, and bidirectional encoding.

  5. Knowl 5 — Meta-VAE hard-negative sampling

    algorithm

    The Meta-VAE sampler identifies relation classes whose utterance distributions are close and uses utterances from those classes as hard negatives. For a relation class, the VAE represents an utterance with a diagonal-Gaussian latent distribution Pi=N(μi,Σi)P_i=\mathcal{N}(\mu_i,\Sigma_i); the sampler measures the distance from relation ii to relation jj with DKL(Pi∥Pj)D_{\mathrm{KL}}(P_i\|P_j). Here μi\mu_i is the latent mean vector and Σi\Sigma_i is the diagonal latent covariance matrix for relation ii. For each relation, it selects the kk closest other relations, then randomly draws one utterance from each selected class for every training utterance of the original class. The paper uses k=3k=3.

    Input: Training utterances D with relation labels; trained Meta-VAE distance function d; number of relation classes n; number of negatives per relation k = 3
    Output: For each training utterance, an index list containing that utterance and k hard-negative utterances
    Build rel2utt, mapping each relation to its utterance indices, and utt2rel, mapping each utterance index to its relation.
    Initialize an n by n distance matrix dist.
    For each relation i from 1 to n:
        For each relation j from 1 to n:
            Randomly draw an utterance u_i from relation i and u_j from relation j.
            Set dist[i, j] = d(u_i, u_j), using KL divergence between their latent Gaussians.
    Set dist[i, i] to infinity.
    For each relation i, select the k other relations with smallest distance from i.
    For each utterance index idx in D:
        Start a list with idx.
        Find relation i = utt2rel[idx].
        For each of the k closest relations to i:
            Randomly draw one utterance index from that relation and append it to the list.
        Add the list to the output iterator.
    Return the iterator.
  6. Knowl 6 — Contrastive Structured Constraint for separating hard samples

    model/method

    The Contrastive Structured Constraint (CSC) trains the extractor to distinguish an utterance from hard-negative utterances while keeping the candidate triplet fixed. For a positive example with context utu_t and triplet (s+,r+,o+)(s^+,r^+,o^+), the structured input contains that context, subject, object, and relation. A negative example replaces only the context with a Meta-VAE-selected utterance ut,j−u^-_{t,j}, while retaining (s+,r+,o+)(s^+,r^+,o^+). Thus, the model must judge whether the same triplet fits the positive context rather than a semantically related context.

    PAE represents each structured input by the hidden state of its final input token and passes that representation through a fully connected layer. CSC applies a symmetric KL-divergence-based loss to the positive and negative output vectors, encouraging them to move apart rather than toward fixed positive and negative target polarities. The extractor is trained with this constraint alongside its sequence-to-sequence extraction loss.

  7. Knowl 7 — Evaluation design and implementation settings

    experimental setup

    Experiments evaluate multiple-triplet extraction on FewRel and single-triplet extraction on FewRel and PersonaExt. FewRel has 56,000 samples, 72,964 entities, 80 relations, and mean length 24.95; PersonaExt has 35,078 samples, 3,295 entities, 105 relations, and mean length 13.44. For each dataset, the number of unseen relations is n∈{5,10,15}n\in\{5,10,15\}. Five random relation-set folds are created for each setting, keeping training, validation, and test sentences in disjoint relation sets; each fold is run three times. Multi-triplet extraction is scored by micro-precision, micro-recall, and micro-F1; single-triplet extraction is scored by accuracy.

    PAG is GPT-2 with 124 million parameters, PAE is BART with 140 million parameters, and Meta-VAE has 2.4 million parameters. The models are initially fine-tuned for five epochs with AdamW, selecting parameters by validation loss. Batch sizes are 128 for PAG and 32 for PAE; learning rates are 3×10−53\times10^{-5} for PAG, 6×10−56\times10^{-5} for PAE, and 0.005 for Meta-VAE; the warmup ratio is 0.2. PAG generates 250 examples per relation using validation and test relation labels as prompts, after which PAE is fine-tuned on the synthetic examples. Greedy decoding is the default for single-triplet extraction, and triplet-search decoding is used for multiple-triplet extraction.

  8. Knowl 8 — Zero-shot extraction results

    empirical result

    The proposed framework was compared with TableSequence (TS) and RelationPrompt (RP). The following values are the reported results; FewRel multiple-triplet columns are precision, recall, and F1, and the single-triplet columns and PersonaExt column are accuracy. Results are averages over the five relation-set folds and three runs per fold.

    Unseen relations Model FewRel multi P. FewRel multi R. FewRel multi F1. FewRel single Acc. PersonaExt Acc.
    5 TS 15.23 1.91 3.40 11.82 –
    5 RP 20.80 24.32 22.34 22.27 38.95
    5 Ours 25.79 34.54 29.47 24.46 40.01
    10 TS 28.93 3.60 6.37 12.54 –
    10 RP 21.59 28.68 24.61 23.18 26.29
    10 Ours 23.31 27.42 25.15 22.89 28.09
    15 TS 19.03 1.99 3.48 11.65 –
    15 RP 17.73 23.20 20.08 18.97 27.25
    15 Ours 20.68 23.39 21.95 19.47 27.57

    On PersonaExt, the proposed model exceeds RP in all three settings, with a mean accuracy gain of 1.06 percentage points. On FewRel multiple-triplet extraction, it has higher F1 than RP in all three settings and the paper reports an average F1 gain of 3.18 points; its precision is also higher in every setting. For FewRel single-triplet extraction, it outperforms RP at five and fifteen unseen relations, but not at ten.

  9. Knowl 9 — Hard-negative ablation and representation evidence

    empirical result

    On PersonaExt, the authors compared Meta-VAE hard-negative sampling with alternative samplers while keeping CSC and the random seed the same. Each reported accuracy is averaged over three runs. “w/o HNS” removes both Meta-VAE sampling and CSC; “Rand” randomly selects negative utterances from different relations. The table reports accuracy for each number of unseen relations, with the decrease from the full model shown where reported.

    Sampler n=5 n=10 n=15
    Ours 39.91 32.47 23.10
    w/o HNS 38.21 ( 1.70) 31.86 ( 0.61) 22.09 ( 1.01)
    SpERT 24.38 ( 15.53) 31.79 ( 0.68) 22.47 ( 0.63)
    RSAN 37.41 ( 2.50) 30.65 ( 1.82) 21.67 ( 1.43)
    GenTaxo 38.66 ( 1.25) 30.55 ( 1.92) 22.15 ( 0.95)
    Rand 37.59 ( 2.32) 30.69 ( 1.78) 22.01 ( 1.09)

    Meta-VAE exceeds the other samplers by 2.66 accuracy points on average and exceeds the strongest alternative, GenTaxo, by 1.37 points on average. The authors also visualized PAE representations before and after CSC fine-tuning using PCA: before fine-tuning, each positive example was close to its three Meta-VAE-selected negatives; after fine-tuning, positive and negative representations were dispersed. This provides empirical evidence that the sampler selects semantically close negatives and CSC separates their representations.

  10. Knowl 10 — Stated limitations of the dataset and extractor

    limitation

    PersonaExt's annotation scheme is not designed for implicit persona attributes or for extracting multiple persona triplets from one utterance. The authors give the example of inferring a liking-for-animals triplet from a statement about walking one's own dogs and neighbors' dogs; their current scheme does not formalize such implicit inference. The extraction framework also does not use complementary information from surrounding dialogue turns. Consequently, when input contains multiple utterances and speakers, pronouns and speaker changes can prevent the model from associating an extracted triplet with the correct speaker.

Coverage note — The smaller decoding-strategy and synthetic-data-size sweeps, the individual case-study examples, and the full inventory of 105 relation labels are omitted because they are secondary analyses or dataset detail rather than central method and benchmark findings.

References

  1. 1.Anton Alekseev and Sergey I Nikolenko. 2016. Predicting the age of social network users from user-generated texts with word embeddings. In 2016 IEEE Artificial Intelligence and Natural Language Conference (AINL), pages 1–11. IEEE.
  2. 2.Dzmitry Bahdanau, Kyung Hyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations, ICLR 2015.
  3. 3.Yoshua Bengio, Réjean Ducharme, and Pascal Vincent. 2000. A neural probabilistic language model. Advances in Neural Information Processing Systems, 13.
  4. 4.Christopher M Bishop. 1998. Latent variable models. In Proceedings of the NATO Advanced Study Institute on Learning in Graphical Models, pages 371–403. Springer.
  5. 5.Erik Cambria, Qian Liu, Sergio Decherchi, Frank Xing, and Kenneth Kwok. 2022. SenticNet 7: A commonsense-based neurosymbolic AI framework for explainable sentiment analysis. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 3829–3839.
  6. 6.Yu Cao, Wei Bi, Meng Fang, Shuming Shi, and Dacheng Tao. 2022. A model-agnostic data manipulation method for persona-based dialogue generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7984–8002.
  7. 7.Oscar Chang, Lampros Flokas, and Hod Lipson. 2019. Principled weight initialization for hypernetworks. In International Conference on Learning Representations.
  8. 8.Chih-Yao Chen and Cheng-Te Li. 2021. ZS-BERT: Towards zero-shot relation extraction with attribute representation learning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3470–3479. Association for Computational Linguistics.
  9. 9.Xiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng, Yunzhi Yao, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. 2022. KnowPrompt: Knowledge-aware prompt-tuning with synergistic optimization for relation extraction. In Proceedings of the ACM Web Conference 2022, pages 2778–2788.
  10. 10.Yew Ken Chia, Lidong Bing, Soujanya Poria, and Luo Si. 2022. RelationPrompt: Leveraging prompts to generate synthetic data for zero-shot relation triplet extraction. In Findings of the Association for Computational Linguistics: ACL 2022, pages 45–57.
  11. 11.Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. In NIPS 2014 Workshop on Deep Learning, December 2014.
  12. 12.Morgane Ciot, Morgan Sonderegger, and Derek Ruths. 2013. Gender inference of Twitter users in non-English contexts. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1136–1145.
  13. 13.Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1):37–46.
  14. 14.Raminta Daniulaityte, Lu Chen, Francois R Lamy, Robert G Carlson, Krishnaprasad Thirunarayan, Amit Sheth, et al. 2016. “when ‘bad’ is ‘good”’: identifying personal communication and sentiment in drug-related tweets. JMIR Public Health and Surveillance, 2(2):e6327.
  15. 15.Bi’an Du, Xiang Gao, Wei Hu, and Xin Li. 2021. Self-contrastive learning with hard negative sampling for self-supervised point cloud learning. In Proceedings of the 29th ACM International Conference on Multimedia, pages 3133–3142.
  16. 16.Markus Eberts and Adrian Ulges. 2020. Span-based joint entity and relation extraction with transformer pre-training. In ECAI 2020, pages 2006–2013. IOS Press.
  17. 17.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898.
  18. 18.Zhiqiang Geng, Yanhui Zhang, and Yongming Han. 2021. Joint entity and relation extraction model based on rich semantics. Neurocomputing, 429:132–140.
  19. 19.Jia-Chen Gu, Zhenhua Ling, Yu Wu, Quan Liu, Zhigang Chen, and Xiaodan Zhu. 2021. Detecting speaker personas from conversational texts. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1126–1136.
  20. 20.Qipeng Guo, Yuqing Yang, Hang Yan, Xipeng Qiu, and Zheng Zhang. 2022. DORE: Document ordered relation extraction based on generative framework. arXiv e-prints, pages arXiv–2210.
  21. 21.Pankaj Gupta, Hinrich Schütze, and Bernt Andrassy. 2016. Table filling multi-task recurrent neural network for joint entity and relation extraction. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 2537–2547.
  22. 22.David Ha, Andrew M. Dai, and Quoc V. Le. 2017. Hypernetworks. In International Conference on Learning Representations.
  23. 23.Xu Han, Hao Zhu, Pengfei Yu, Ziyun Wang, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2018. FewRel: A large-scale supervised few-shot relation classification dataset with state-of-the-art evaluation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4803–4809.
  24. 24.Kai He, Yucheng Huang, Rui Mao, Tieliang Gong, Chen Li, and Erik Cambria. 2023. Virtual prompt pre-training for prototype-based few-shot relation extraction. Expert Systems with Applications, 213:118927.
  25. 25.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the knowledge in a neural network. stat, 1050:9.
  26. 26.Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, pages 4171–4186.
  27. 27.Diederik P Kingma and Max Welling. 2014. Auto-encoding variational Bayes. In International Conference on Learning Representations.
  28. 28.Solomon Kullback and Richard A Leibler. 1951. On information and sufficiency. The Annals of Mathematical Statistics, 22(1):79–86.
  29. 29.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059.
  30. 30.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880.
  31. 31.Qi Li and Heng Ji. 2014. Incremental joint extraction of entity mentions and relations. In 52nd Annual Meeting of the Association for Computational Linguistics, ACL 2014, pages 402–412. Association for Computational Linguistics (ACL).
  32. 32.Wei Li, Luyao Zhu, Rui Mao, and Erik Cambria. 2023. SKIER: A symbolic knowledge integrated model for conversational emotion recognition. Proceedings of the AAAI Conference on Artificial Intelligence.
  33. 33.Shengcai Liao and Ling Shao. 2022. Graph sampling based deep metric learning for generalizable person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7359–7368.
  34. 34.Ilya Loshchilov and Frank Hutter. 2018. Decoupled weight decay regularization. In International Conference on Learning Representations.
  35. 35.Ximing Lu, Sean Welleck, Peter West, Liwei Jiang, Jungo Kasai, Daniel Khashabi, Ronan Le Bras, Lianhui Qin, Youngjae Yu, Rowan Zellers, et al. 2022. NEUROLOGIC A*esque decoding: Constrained text generation with lookahead heuristics. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 780–799.
  36. 36.Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015. Effective approaches to attention-based neural machine translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 1412–1421.
  37. 37.Rui Mao, Qian Liu, Kai He, Wei Li, and Erik Cambria. 2023. The biases of pre-trained language models: An empirical study on prompt-based sentiment analysis and emotion detection. IEEE Transactions on Affective Computing.
  38. 38.Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Jie Ma, Alessandro Achille, Rishita Anubhai, Cicero Nogueira dos Santos, Bing Xiang, and Stefano Soatto. 2020. Structured prediction as translation between augmented natural languages. In International Conference on Learning Representations.
  39. 39.Daniel Preo¸tiuc-Pietro, Vasileios Lampos, and Nikolaos Aletras. 2015. An analysis of the user occupational class through Twitter content. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1754–1764.
  40. 40.Pengda Qin, Weiran Xu, and William Yang Wang. 2018. DSGAN: Generative adversarial training for distant supervision relation extraction. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 496–505.
  41. 41.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI Blog.
  42. 42.Joshua David Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka. 2020. Contrastive learning with hard negative samples. In International Conference on Learning Representations.
  43. 43.Abhinav Shrivastava, Abhinav Gupta, and Ross Girshick. 2016. Training region-based object detectors with online hard example mining. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 761–769.
  44. 44.Vincent Sitzmann, Eric Chan, Richard Tucker, Noah Snavely, and Gordon Wetzstein. 2020. MetaSDF: Meta-learning signed distance functions. Advances in Neural Information Processing Systems, 33:10136–10147.
  45. 45.Yumin Suh, Bohyung Han, Wonsik Kim, and Kyoung Mu Lee. 2019. Stochastic class-based hard example mining for deep metric learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7251–7259.
  46. 46.Vinay Kumar Verma, Gundeep Arora, Ashish Mishra, and Piyush Rai. 2018. Generalized zero-shot learning via synthesized examples. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4281–4289.
  47. 47.Jue Wang and Wei Lu. 2020. Two are better than one: Joint entity and relation extraction with table-sequence encoders. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1706–1721.
  48. 48.Sean Welleck, Jason Weston, Arthur Szlam, and Kyunghyun Cho. 2019. Dialogue natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3731–3741.
  49. 49.Sean Welleck, Jason Weston, Arthur Szlam, and Kyunghyun Cho. 2020. Dialogue natural language inference. In 57th Annual Meeting of the Association for Computational Linguistics, ACL 2019, pages 3731–3741. Association for Computational Linguistics (ACL).
  50. 50.Svante Wold, Kim Esbensen, and Paul Geladi. 1987. Principal component analysis. Chemometrics and Intelligent Laboratory Systems, 2(1-3):37–52.
  51. 51.Chien-Sheng Wu, Andrea Madotto, Zhaojiang Lin, Peng Xu, and Pascale Fung. 2020. Getting to know you: User attribute extraction from dialogues. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 581–589.
  52. 52.Yongqin Xian, Christoph H Lampert, Bernt Schiele, and Zeynep Akata. 2018. Zero-shot learning - a comprehensive evaluation of the good, the bad and the ugly. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(9):2251–2265.
  53. 53.Hongbin Ye, Ningyu Zhang, Shumin Deng, Mosha Chen, Chuanqi Tan, Fei Huang, and Huajun Chen. 2021. Contrastive triple extraction with generative transformer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 14257–14265.
  54. 54.Y Yuan, X Zhou, S Pan, Q Zhu, Z Song, and L Guo. 2021a. A relation-specific attention network for joint entity and relation extraction. In International Joint Conference on Artificial Intelligence. International Joint Conference on Artificial Intelligence.
  55. 55.Yue Yuan, Xiaofei Zhou, Shirui Pan, Qiannan Zhu, Zeliang Song, and Li Guo. 2021b. A relation-specific attention network for joint entity and relation extraction. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence, pages 4054–4060.
  56. 56.Qingkai Zeng, Jinfeng Lin, Wenhao Yu, Jane Cleland-Huang, and Meng Jiang. 2021. Enhancing taxonomy completion with concept generation via fusing relational representations. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pages 2104–2113.
  57. 57.Meishan Zhang, Yue Zhang, and Guohong Fu. 2017. End-to-end neural relation extraction with global optimization. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 1730–1740.
  58. 58.Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing dialogue agents: I have a dog, do you have pets too? In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2204–2213.
  59. 59.Yinhe Zheng, Rongsheng Zhang, Minlie Huang, and Xiaoxi Mao. 2020. A pre-training based personalized dialogue generation model with persona-sparse data. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 9693–9700.
  60. 60.Luyao Zhu, Wei Li, Yong Shi, and Kun Guo. 2020. SentiVec: learning sentiment-context vector via kernel optimization function for sentiment analysis. IEEE Transactions on Neural Networks and Learning Systems, 32(6):2561–2572.

Citation

MLA
Zhu, L., et al. “PAED: Zero-Shot Persona Attribute Extraction in Dialogues”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 9771–87, https://doi.org/10.18653/v1/2023.acl-long.544.
APA
Zhu, L., Li, W., Mao, R., Pandelea, V., & Cambria, E. (2023). PAED: Zero-Shot Persona Attribute Extraction in Dialogues. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9771–9787. https://doi.org/10.18653/v1/2023.acl-long.544
Chicago
Zhu, L., W. Li, R. Mao, V. Pandelea, and E. Cambria. 2023. “PAED: Zero-Shot Persona Attribute Extraction in Dialogues”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9771–87. https://doi.org/10.18653/v1/2023.acl-long.544.
Harvard
Zhu, L. et al. (2023) “PAED: Zero-Shot Persona Attribute Extraction in Dialogues”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 9771–9787. Available at: https://doi.org/10.18653/v1/2023.acl-long.544.
Vancouver
1. Zhu L, Li W, Mao R, Pandelea V, Cambria E (2023) PAED: Zero-Shot Persona Attribute Extraction in Dialogues. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 9771–9787

BibTeX

@inproceedings{zhu-etal-2023-paed,
    title = "{PAED}: Zero-Shot Persona Attribute Extraction in Dialogues",
    author = "Zhu, Luyao  and
      Li, Wei  and
      Mao, Rui  and
      Pandelea, Vlad  and
      Cambria, Erik",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.544/",
    doi = "10.18653/v1/2023.acl-long.544",
    pages = "9771--9787"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/