The AboutMe dataset is a collection of web-scraped self-descriptions and personal biographical profiles used to study demographic representation and the social impacts of data curation in natural language processing. By associating online text with creator-level metadata such as geographic affiliations, social roles, and topical interests, the dataset enables researchers to trace web content to its authors. It primarily serves as an analytical resource for evaluating how automated quality filters, language identification algorithms, and selection heuristics systematically alter the representation of diverse populations and demographic groups during the creation of pretraining corpora for large language models.