keyword
PAED dataset
The PAED dataset, or Persona Attribute Extraction in Dialogues dataset, is a natural language processing benchmark designed for identifying and extracting structured personal attributes from conversational text. Developed to advance personalized human-computer interaction and conversational artificial intelligence, the dataset annotates user characteristics, preferences, and traits in the form of structured subject-relation-object triplets derived from dialogue utterances. It provides standardized, high-quality annotations that resolve ambiguities and inconsistencies found in earlier conversational persona corpora by applying precise text-to-label matching criteria. The dataset is primarily utilized to train, evaluate, and benchmark information extraction models, especially in zero-shot learning scenarios where systems must detect previously unseen persona attributes across multi-turn dialogues.
1 item

