ArnetMiner: extraction and mining of academic social networks
Jie TangJing ZhangLimin YaoJuan-Zi LiLi ZhangZhong Su
Presents ArnetMiner, an academic social network mining system that integrates conditional random fields for automated researcher profiling, hidden Markov random fields for publication name disambiguation, and joint author-conference-topic models to enable effective expertise and association search.
The article describes the development of the ArnetMiner system to extract researcher profiles from the web, integrate publication records from digital libraries, model academic social networks in a unified way, and deliver search services such as expertise and association search. These tasks address the growing need for semantics-rich tools that go beyond simple keyword matching in large scholarly databases, where incomplete profiles and name ambiguities hinder effective discovery.
The article set out to evaluate methods for automatic profile extraction, name disambiguation during data integration, and simultaneous topic modeling of papers, authors, and venues. It also aimed to demonstrate how the resulting models support practical search applications.
The authors built the system through a sequence of experiments on real web data and publication collections. They used conditional random fields for unified tagging of researcher homepages, a hidden Markov random field framework incorporating publication relationships for disambiguation, and three variants of an author-conference-topic model for joint topic estimation. Performance was measured on nearly 900 annotated homepages and a disambiguation dataset of 14 names, with search evaluations drawn from frequent user queries against pooled judgments from similar systems.
The unified extraction approach achieved an average F1 score of 83.37 percent, substantially above rule-based and classification baselines. The disambiguation framework reached 91.39 percent average F1, improving over hierarchical clustering by more than 10 points when relationships among papers were included. The author-conference-topic models outperformed language-model, LDA, and author-topic baselines on expertise search, with the first variant delivering the highest mean average precision of 71 percent across papers, authors, and conferences. Roughly 448,000 researcher profiles and over one million papers were integrated into the live system.
These results show that holistic modeling captures dependencies across data types that separate methods miss, leading to more accurate profiles, cleaner author identities, and more relevant search results. For decision makers, the gains translate into reduced manual curation effort, lower risk of missed expertise, and faster discovery of collaborators or venues.
The article recommends extending the topic models with citation and temporal links, automating the choice of the number of distinct persons in disambiguation, and refining extraction rules for greater coverage. Further work on these fronts is needed before the system can scale without manual intervention to all author names.
The main limitations are reliance on manually supplied person counts for disambiguation, evaluations confined to computer science data, and fixed hyperparameter settings in the topic models. Confidence is high for the reported extraction and disambiguation tasks on the tested collections, but lower for broader domains or fully automatic operation.
- Paper: The Author-Topic Model for Authors and Documents, Michal Rosen-Zvi et al. (2004). Its author-topic model is a direct methodological precursor to ArnetMiner’s joint modeling of authors and research topics.
- Paper: Unsupervised Learning by Probabilistic Latent Semantic Analysis, Thomas Hofmann (2001). Its probabilistic latent-topic framework supplies foundational context for the topic-modeling component ArnetMiner adapts to authors and venues.
- Paper: Early results for Named Entity Recognition with Conditional Random Fields, Feature Induction and Web-Enhanced Lexicons, A. McCallum et al. (2003). Its CRF-based sequence-labeling work provides a relevant methodological foundation for ArnetMiner’s CRF extraction of researcher profiles.
- Paper: PathSim, Yizhou Sun et al. (2011). It carries academic-network analysis toward typed, path-based similarity search, extending the kind of researcher and publication relationships ArnetMiner brings together.
