ArnetMiner: extraction and mining of academic social networks

Jie TangJing ZhangLimin YaoJuan-Zi LiLi ZhangZhong Su

article2008KDD2,586 citationsTest of Time Award for Applied Science

Presents ArnetMiner, an academic social network mining system that integrates conditional random fields for automated researcher profiling, hidden Markov random fields for publication name disambiguation, and joint author-conference-topic models to enable effective expertise and association search.

Listen

The article describes the development of the ArnetMiner system to extract researcher profiles from the web, integrate publication records from digital libraries, model academic social networks in a unified way, and deliver search services such as expertise and association search. These tasks address the growing need for semantics-rich tools that go beyond simple keyword matching in large scholarly databases, where incomplete profiles and name ambiguities hinder effective discovery.

The article set out to evaluate methods for automatic profile extraction, name disambiguation during data integration, and simultaneous topic modeling of papers, authors, and venues. It also aimed to demonstrate how the resulting models support practical search applications.

The authors built the system through a sequence of experiments on real web data and publication collections. They used conditional random fields for unified tagging of researcher homepages, a hidden Markov random field framework incorporating publication relationships for disambiguation, and three variants of an author-conference-topic model for joint topic estimation. Performance was measured on nearly 900 annotated homepages and a disambiguation dataset of 14 names, with search evaluations drawn from frequent user queries against pooled judgments from similar systems.

The unified extraction approach achieved an average F1 score of 83.37 percent, substantially above rule-based and classification baselines. The disambiguation framework reached 91.39 percent average F1, improving over hierarchical clustering by more than 10 points when relationships among papers were included. The author-conference-topic models outperformed language-model, LDA, and author-topic baselines on expertise search, with the first variant delivering the highest mean average precision of 71 percent across papers, authors, and conferences. Roughly 448,000 researcher profiles and over one million papers were integrated into the live system.

These results show that holistic modeling captures dependencies across data types that separate methods miss, leading to more accurate profiles, cleaner author identities, and more relevant search results. For decision makers, the gains translate into reduced manual curation effort, lower risk of missed expertise, and faster discovery of collaborators or venues.

The article recommends extending the topic models with citation and temporal links, automating the choice of the number of distinct persons in disambiguation, and refining extraction rules for greater coverage. Further work on these fronts is needed before the system can scale without manual intervention to all author names.

The main limitations are reliance on manually supplied person counts for disambiguation, evaluations confined to computer science data, and fixed hyperparameter settings in the topic models. Confidence is high for the reported extraction and disambiguation tasks on the tested collections, but lower for broader domains or fully automatic operation.

  • Paper: PathSim, Yizhou Sun et al. (2011). It carries academic-network analysis toward typed, path-based similarity search, extending the kind of researcher and publication relationships ArnetMiner brings together.
Cover for ArnetMiner: extraction and mining of academic social networks

Abstract

This paper addresses several key issues in the ArnetMiner system, which aims at extracting and mining academic social networks. Specifically, the system focuses on: 1) Extracting researcher profiles automatically from the Web; 2) Integrating the publication data into the network from existing digital libraries; 3) Modeling the entire academic network; and 4) Providing search services for the academic network. So far, 448,470 researcher profiles have been extracted using a unified tagging approach. We integrate publications from online Web databases and propose a probabilistic framework to deal with the name ambiguity problem. Furthermore, we propose a unified modeling approach to simultaneously model topical aspects of papers, authors, and publication venues. Search services such as expertise search and people association search have been provided based on the modeling results. In this paper, we describe the architecture and main features of the system. We also present the empirical evaluation of the proposed methods.

Table of Contents

  • 1. INTRODUCTION
  • 2. RELATED WORK
  • 2.1 Person Profile Extraction
  • 2.2 Name Disambiguation
  • 2.3 Topic Modeling
  • 2.4 Academic Search
  • 3. OVERVIEW OF ARNETMINER
  • 4. RESEARCHERPROFILEEXTRACTION
  • 4.1 Problem Definition
  • 4.2 A Unified Approach to Profiling
  • 4.2.1 Process
  • 4.2.2 CRF model and Features
  • 4.3 Profile Extraction Performance
  • 5. NAMEDISAMBIGUATION
  • 5.1 Problem Definition
  • 5.2 A Unified Probabilistic Framework
  • 5.2.1 Formalization using HMRF
  • 5.2.2 EM framework
  • 5.3 Name Disambiguation Performance
  • 6. MODELING ACADEMIC NETWORK
  • 6.1 Our Proposed Topic Models
  • 6.2 ACT Model 1
  • 6.3 ACT Model 2
  • 6.4 ACT Model 3
  • 7. ACADEMIC SEARCH SERVICES
  • 7.1 Applying ACT Models to Expertise Search
  • 7.1.1 Process
  • 7.1.2 Expertise Search Performance
  • 7.2 Applying ACT Models to Association Search
  • 7.3 Other Applications
  • 8. CONCLUSION
  • 9. ACKNOWLEDGMENTS
  • 10. REFERENCES

Knowls

  1. Knowl 1 — Author-Conference-Topic Model 1 (ACT1)

    model/method

    The Author-Conference-Topic Model 1 (ACT1) is a generative probabilistic topic model that simultaneously models the topical distributions of papers, authors, and publication venues (conferences, journals, or books). In ACT1, each author is associated with a multinomial distribution over topics, and each word token in a paper alongside the publication venue stamp is generated conditioned on a sampled latent topic.

    Given a collection of DD papers where paper d∈{1,…,D}d \in \{1, \dots, D\} has word vector wd=(wd1,…,wdNd)\mathbf{w}_d = (w_{d1}, \dots, w_{d N_d}) drawn from vocabulary VV, author subset ad⊆{1,…,A}\mathbf{a}_d \subseteq \{1, \dots, A\}, and publication venue cd∈{1,…,C}c_d \in \{1, \dots, C\}:

    1. For each topic z∈{1,…,T}z \in \{1, \dots, T\}, draw word distribution ϕz∼Dirichlet(β)\boldsymbol{\phi}_z \sim \text{Dirichlet}(\boldsymbol{\beta}) and venue distribution ψz∼Dirichlet(μ)\boldsymbol{\psi}_z \sim \text{Dirichlet}(\boldsymbol{\mu}).
    2. For each author x∈{1,…,A}x \in \{1, \dots, A\}, draw topic distribution θx∼Dirichlet(α)\boldsymbol{\theta}_x \sim \text{Dirichlet}(\boldsymbol{\alpha}).
    3. For each word wdiw_{di} (i∈{1,…,Nd}i \in \{1, \dots, N_d\}) in paper dd:
      • Draw author xdi∼Uniform(ad)x_{di} \sim \text{Uniform}(\mathbf{a}_d);
      • Draw topic zdi∼Multinomial(θxdi)z_{di} \sim \text{Multinomial}(\boldsymbol{\theta}_{x_{di}});
      • Draw word wdi∼Multinomial(ϕzdi)w_{di} \sim \text{Multinomial}(\boldsymbol{\phi}_{z_{di}});
      • Draw venue stamp cdi∼Multinomial(ψzdi)c_{di} \sim \text{Multinomial}(\boldsymbol{\psi}_{z_{di}}).

    In Gibbs sampling inference, the joint posterior distribution of assigning the ii-th token in paper dd to author xdix_{di} and topic zdiz_{di}, conditioned on all other variables, is given by:

    P(zdi,xdi∣z−di,x−di,w,c,α,β,μ)∝mxdizdi−di+αzdi∑z=1T(mxdiz−di+αz)⋅nzdiwdi−di+βwdi∑v=1V(nzdiv−di+βv)⋅nzdicd−d+μcd∑c=1C(nzdic−d+μc)P(z_{di}, x_{di} \mid \mathbf{z}_{-di}, \mathbf{x}_{-di}, \mathbf{w}, \mathbf{c}, \boldsymbol{\alpha}, \boldsymbol{\beta}, \boldsymbol{\mu}) \propto \frac{m_{x_{di} z_{di}}^{-di} + \alpha_{z_{di}}}{\sum_{z=1}^T (m_{x_{di} z}^{-di} + \alpha_z)} \cdot \frac{n_{z_{di} w_{di}}^{-di} + \beta_{w_{di}}}{\sum_{v=1}^V (n_{z_{di} v}^{-di} + \beta_v)} \cdot \frac{n_{z_{di} c_d}^{-d} + \mu_{c_d}}{\sum_{c=1}^C (n_{z_{di} c}^{-d} + \mu_c)}

    where superscript −di-di denotes counts excluding the current token, mxzm_{xz} is the number of times topic zz is assigned to author xx, nzwn_{zw} is the number of times word ww is generated by topic zz, and nzcn_{zc} is the number of times venue cc is generated by topic zz. Hyperparameters are configured as α=50/T\alpha = 50/T, β=0.01\beta = 0.01, and μ=0.1\mu = 0.1.

    Parameter estimates are obtained as:

    ϕzw=nzw+βw∑v=1V(nzv+βv),ψzc=nzc+μc∑c′=1C(nzc′+μc′),θxz=mxz+αz∑z′=1T(mxz′+αz′)\phi_{zw} = \frac{n_{zw} + \beta_w}{\sum_{v=1}^V (n_{zv} + \beta_v)}, \quad \psi_{zc} = \frac{n_{zc} + \mu_c}{\sum_{c'=1}^C (n_{zc'} + \mu_{c'})}, \quad \theta_{xz} = \frac{m_{xz} + \alpha_z}{\sum_{z'=1}^T (m_{xz'} + \alpha_{z'})}
  2. Knowl 2 — Author-Conference-Topic Models 2 and 3 (ACT2 and ACT3)

    model/method

    ACT2 and ACT3 are alternative formulations for simultaneously modeling text, authorship, and venues:

    ACT2 (Author-Venue Pair Conditioned Model): Assumes authors select a venue prior to writing and generate topics conditioned on the author-venue pair (xdi,cd)(x_{di}, c_d). For each word token wdiw_{di} in paper dd with venue cdc_d:

    1. Draw author-venue pair (xdi,cd)(x_{di}, c_d) uniformly from {(a,cd)∣a∈ad}\{(a, c_d) \mid a \in \mathbf{a}_d\}.
    2. Draw topic zdi∼Multinomial(θ(xdi,cd))z_{di} \sim \text{Multinomial}(\boldsymbol{\theta}_{(x_{di}, c_d)}), with θ(x,c)∼Dirichlet(α)\boldsymbol{\theta}_{(x, c)} \sim \text{Dirichlet}(\boldsymbol{\alpha}).
    3. Draw word wdi∼Multinomial(ϕzdi)w_{di} \sim \text{Multinomial}(\boldsymbol{\phi}_{z_{di}}).

    The Gibbs sampling update is:

    P(zdi,(xc)di∣z−di,x−di,c−d,w,α,β)∝m(xc)dizdi−di+αzdi∑z=1T(m(xc)diz−di+αz)⋅nzdiwdi−di+βwdi∑v=1V(nzdiv−di+βv)P(z_{di}, (xc)_{di} \mid \mathbf{z}_{-di}, \mathbf{x}_{-di}, \mathbf{c}_{-d}, \mathbf{w}, \boldsymbol{\alpha}, \boldsymbol{\beta}) \propto \frac{m_{(xc)_{di} z_{di}}^{-di} + \alpha_{z_{di}}}{\sum_{z=1}^T (m_{(xc)_{di} z}^{-di} + \alpha_z)} \cdot \frac{n_{z_{di} w_{di}}^{-di} + \beta_{w_{di}}}{\sum_{v=1}^V (n_{z_{di} v}^{-di} + \beta_v)}

    ACT3 (Continuous Venue Normal Linear Model): Treats the venue index cdc_d as a continuous response generated after all paper words are sampled, drawn from a normal linear model cd∼N(η⊤τd,σ2)c_d \sim \mathcal{N}(\boldsymbol{\eta}^\top \boldsymbol{\tau}_d, \sigma^2), where τd∈RT\boldsymbol{\tau}_d \in \mathbb{R}^T is the normalized topic frequency vector of paper dd (with τd[k]=1Nd∑i=1NdI[zdi=k]\tau_d[k] = \frac{1}{N_d} \sum_{i=1}^{N_d} I[z_{di} = k]).

    Inference uses a Gibbs EM algorithm. In the E-step, the posterior topic assignment is:

    P(zdi,xdi∣z−di,x−di,cd,w,α,β,η,σ2)∝12πσ2exp⁡(−(cd−η⊤τd)22σ2)⋅mxdizdi−di+αzdi∑z(mxdiz−di+αz)⋅nzdiwdi−di+βwdi∑v(nzdiv−di+βv)P(z_{di}, x_{di} \mid \mathbf{z}_{-di}, \mathbf{x}_{-di}, c_d, \mathbf{w}, \boldsymbol{\alpha}, \boldsymbol{\beta}, \boldsymbol{\eta}, \sigma^2) \propto \frac{1}{\sqrt{2\pi\sigma^2}} \exp\left(-\frac{(c_d - \boldsymbol{\eta}^\top \boldsymbol{\tau}_d)^2}{2\sigma^2}\right) \cdot \frac{m_{x_{di} z_{di}}^{-di} + \alpha_{z_{di}}}{\sum_z (m_{x_{di} z}^{-di} + \alpha_z)} \cdot \frac{n_{z_{di} w_{di}}^{-di} + \beta_{w_{di}}}{\sum_v (n_{z_{di} v}^{-di} + \beta_v)}

    In the M-step, regression parameters are updated as:

    ηnew=(E[A⊤A])−1E[A]⊤c,σnew2=1D(c⊤c−c⊤E[A](E[A⊤A])−1E[A]⊤c)\boldsymbol{\eta}_{\text{new}} = (\mathbb{E}[A^\top A])^{-1} \mathbb{E}[A]^\top \mathbf{c}, \quad \sigma^2_{\text{new}} = \frac{1}{D} \left(\mathbf{c}^\top \mathbf{c} - \mathbf{c}^\top \mathbb{E}[A] (\mathbb{E}[A^\top A])^{-1} \mathbb{E}[A]^\top \mathbf{c}\right)

    where AA is a D×TD \times T matrix whose dd-th row is E[τd]=1Nd∑i=1Ndϕdi\mathbb{E}[\boldsymbol{\tau}_d] = \frac{1}{N_d} \sum_{i=1}^{N_d} \boldsymbol{\phi}_{di}, and E[A⊤A]=∑d=1D1Nd2(∑i=1Nd∑j≠iϕdiϕdj⊤+∑i=1Nddiag{ϕdi})\mathbb{E}[A^\top A] = \sum_{d=1}^D \frac{1}{N_d^2} \left(\sum_{i=1}^{N_d} \sum_{j \neq i} \boldsymbol{\phi}_{di} \boldsymbol{\phi}_{dj}^\top + \sum_{i=1}^{N_d} \text{diag}\{\boldsymbol{\phi}_{di}\}\right) with ϕdi\boldsymbol{\phi}_{di} denoting the topic probability vector generating word wdiw_{di}.

  3. Knowl 3 — Hidden Markov Random Field Formulation for Name Disambiguation

    model/method

    Author name disambiguation is formulated as assigning each publication pi∈P={p1,…,pn}p_i \in P = \{p_1, \dots, p_n\} sharing an ambiguous author name aa to one of kk true underlying researcher entities yh∈{1,…,k}y_h \in \{1, \dots, k\}. Each paper pip_i is represented by a feature vector xi\mathbf{x}_i of word counts from its title, venue, abstract, coauthors, and references.

    Dependencies and relational constraints between publications are modeled using a Hidden Markov Random Field (HMRF) maximizing the a-posteriori conditional probability:

    P(y∣x)=1Zexp⁡(−∑i=1nD(xi,yh)−∑i=1n∑j≠i(D(xi,xj)∑k=15wkrk(xi,xj)))P(\mathbf{y} \mid \mathbf{x}) = \frac{1}{Z} \exp\left( -\sum_{i=1}^n D(\mathbf{x}_i, y_h) - \sum_{i=1}^n \sum_{j \neq i} \left( D(\mathbf{x}_i, \mathbf{x}_j) \sum_{k=1}^5 w_k r_k(\mathbf{x}_i, \mathbf{x}_j) \right) \right)

    where ZZ is a normalization partition factor, D(xi,yh)D(\mathbf{x}_i, y_h) is the distance from paper vector xi\mathbf{x}_i to the centroid of researcher entity yhy_h, D(xi,xj)D(\mathbf{x}_i, \mathbf{x}_j) is the pairwise distance between papers, and rk(xi,xj)∈{0,1}r_k(\mathbf{x}_i, \mathbf{x}_j) \in \{0, 1\} indicates the presence of relationship type kk with assigned weight wkw_k.

    Five relationship types are defined between papers pip_i and pjp_j:

    1. CoPubvenue (r1r_1, w1=0.2w_1 = 0.2): Both papers are published in the same venue (pi.pubvenue=pj.pubvenuep_i.\text{pubvenue} = p_j.\text{pubvenue}).
    2. CoAuthor (r2r_2, w2=0.7w_2 = 0.7): Papers share at least one secondary coauthor (ai(r)=aj(s)a_i^{(r)} = a_j^{(s)} for r,s>0r, s > 0).
    3. Citation (r3r_3, w3=0.3w_3 = 0.3): pip_i cites pjp_j or pjp_j cites pip_i.
    4. Constraints (r4r_4, w4=1.0w_4 = 1.0): Explicit constraints supplied via user feedback.
    5. τ\tau-CoAuthor (r5r_5, w5=0.7τw_5 = 0.7^\tau for τ>1\tau > 1): τ\tau-step extended co-authorship chain (e.g., a secondary author of pip_i coauthors a separate paper with a secondary author of pjp_j).
  4. Knowl 4 — Expectation-Maximization Algorithm for HMRF Name Disambiguation

    algorithm

    An Expectation-Maximization (EM) algorithm performs iterative cluster reassignment, centroid updating, and distance metric parameter learning for HMRF-based name disambiguation.

    The parameterized distance between feature vectors xi\mathbf{x}_i and xj\mathbf{x}_j employs a learned diagonal weight matrix A=diag(a11,…,amm,… )A = \text{diag}(a_{11}, \dots, a_{mm}, \dots):

    D(xi,xj)=1−xi⊤Axj∥xi∥A∥xj∥A,where ∥x∥A=x⊤AxD(\mathbf{x}_i, \mathbf{x}_j) = 1 - \frac{\mathbf{x}_i^\top A \mathbf{x}_j}{\|\mathbf{x}_i\|_A \|\mathbf{x}_j\|_A}, \quad \text{where } \|\mathbf{x}\|_A = \sqrt{\mathbf{x}^\top A \mathbf{x}}
    Input: Paper vectors x1,…,xn\mathbf{x}_1, \dots, \mathbf{x}_n, number of true persons kk, relationship weights wkw_k, learning rate η\eta
    Output: Researcher assignments y1,…,yny_1, \dots, y_n, centroids y1,…,yky_1, \dots, y_k, metric matrix AA
    Initialize centroids y1,…,yky_1, \dots, y_k and metric matrix AA
    repeat
        // E-step: Paper Reassignment
        repeat
            for each paper i=1i = 1 to nn do
                Assign paper xi\mathbf{x}_i to entity yhy_h (h∈{1,…,k}h \in \{1, \dots, k\}) minimizing:
                f(yh,xi)=D(xi,yh)+∑j≠i(D(xi,xj)∑kwkrk(xi,xj))f(y_h, \mathbf{x}_i) = D(\mathbf{x}_i, y_h) + \sum_{j \neq i} \left( D(\mathbf{x}_i, \mathbf{x}_j) \sum_{k} w_k r_k(\mathbf{x}_i, \mathbf{x}_j) \right)
            end for
        until no paper changes assignment
        // M-step: Representative and Metric Parameter Update
        for each entity h=1h = 1 to kk do
            yh←∑i:yi=hxi∥∑i:yi=hxi∥Ay_h \leftarrow \frac{\sum_{i: y_i = h} \mathbf{x}_i}{\|\sum_{i: y_i = h} \mathbf{x}_i\|_A}
        end for
        for each diagonal element amma_{mm} in AA do
            amm←amm+η∑i∂f(yh,xi)∂amma_{mm} \leftarrow a_{mm} + \eta \sum_i \frac{\partial f(y_h, \mathbf{x}_i)}{\partial a_{mm}}
        end for
    until convergence

    The gradient with respect to diagonal parameter amma_{mm} is:

    ∂D(xi,xj)∂amm=ximxjm∥xi∥A∥xj∥A−xi⊤Axjxim2∥xi∥A2+xjm2∥xj∥A22∥xi∥A∥xj∥A∥xi∥A2∥xj∥A2\frac{\partial D(\mathbf{x}_i, \mathbf{x}_j)}{\partial a_{mm}} = \frac{x_{im} x_{jm} \|\mathbf{x}_i\|_A \|\mathbf{x}_j\|_A - \mathbf{x}_i^\top A \mathbf{x}_j \frac{x_{im}^2 \|\mathbf{x}_i\|_A^2 + x_{jm}^2 \|\mathbf{x}_j\|_A^2}{2 \|\mathbf{x}_i\|_A \|\mathbf{x}_j\|_A}}{\|\mathbf{x}_i\|_A^2 \|\mathbf{x}_j\|_A^2}
  5. Knowl 5 — Combined Language Model and Topic Model Scoring for Expertise Search

    model/method

    To combine query term matching with broad topical semantics, academic expertise search scores the relevance of an academic object (paper dd, author aa, or conference cc) to a query qq by multiplying a Dirichlet-smoothed unigram language model (LM) probability by an Author-Conference-Topic (ACT) model probability.

    For a paper dd, the probability of generating word ww is defined as:

    P(w∣d)=PLM(w∣d)×PACT(w∣d)P(w \mid d) = P_{\text{LM}}(w \mid d) \times P_{\text{ACT}}(w \mid d)

    where PLM(w∣d)P_{\text{LM}}(w \mid d) is the Dirichlet-smoothed language model:

    PLM(w∣d)=NdNd+λtf(w,d)Nd+(1−NdNd+λ)tf(w,D)NDP_{\text{LM}}(w \mid d) = \frac{N_d}{N_d + \lambda} \frac{\text{tf}(w, d)}{N_d} + \left(1 - \frac{N_d}{N_d + \lambda}\right) \frac{\text{tf}(w, D)}{N_D}

    with NdN_d the length of document dd, tf(w,d)\text{tf}(w, d) the frequency of word ww in dd, NDN_D the total word count in corpus DD, tf(w,D)\text{tf}(w, D) the corpus frequency of ww, and λ\lambda the Dirichlet prior smoothing parameter set based on average paper length.

    Under ACT1, the document word generation probability is computed as:

    PACT1(w∣d,θ,ϕ)=∑z=1T∑x=1AdP(w∣z,ϕz)P(z∣x,θx)P(x∣d)P_{\text{ACT1}}(w \mid d, \boldsymbol{\theta}, \boldsymbol{\phi}) = \sum_{z=1}^T \sum_{x=1}^{A_d} P(w \mid z, \boldsymbol{\phi}_z) P(z \mid x, \boldsymbol{\theta}_x) P(x \mid d)

    with P(x∣d)=1/∣ad∣P(x \mid d) = 1 / |\mathbf{a}_d|.

    Given a multi-term query qq, the overall relevance score for a document is P(q∣d)=∏w∈qP(w∣d)P(q \mid d) = \prod_{w \in q} P(w \mid d). Analogous likelihoods P(q∣a)P(q \mid a) and P(q∣c)P(q \mid c) are computed for authors and conferences by aggregating their published papers.

  6. Knowl 6 — Social Network Association Search via Topic-Distribution KL Divergence

    algorithm

    Given an academic social network G=(V,E)G = (V, E) and an association query (ai,aj)(a_i, a_j) between source author aia_i and target author aja_j, association search discovers and ranks referral paths connecting the two individuals.

    The directed distance from author aia_i to author aja_j is defined by the Kullback-Leibler (KL) divergence between their estimated ACT author-topic distributions θai\boldsymbol{\theta}_{a_i} and θaj\boldsymbol{\theta}_{a_j}:

    KL(ai,aj)=∑z=1Tθaizlog⁡θaizθajz\text{KL}(a_i, a_j) = \sum_{z=1}^T \theta_{a_i z} \log \frac{\theta_{a_i z}}{\theta_{a_j z}}

    The score of a referral chain is the accumulated sum of KL divergences along its constituent edges.

    Input: Graph G=(V,E)G = (V, E), topic distributions θa\boldsymbol{\theta}_a for all a∈Va \in V, query pair (ai,aj)(a_i, a_j), relaxation factor γ\gamma, maximum path length MM
    Output: Ranked list of near-shortest referral chains from aia_i to aja_j
    // Stage 1: Shortest Association Finding
    Compute edge weights w(u,v)=KL(u,v)w(u, v) = \text{KL}(u, v) for all (u,v)∈E(u, v) \in E
    Run Dijkstra's algorithm with a min-heap to compute the minimum path score Lmin⁡L_{\min} from aia_i to aja_j
    // Stage 2: Near-Shortest Associations Finding
    Initialize empty path collection P\mathcal{P}
    Execute depth-first search (DFS) starting from aia_i towards aja_j:
        Prune any branch whose accumulated path score exceeds (1+γ)Lmin⁡(1 + \gamma) L_{\min}
        Prune any branch whose hop count exceeds MM
        If path reaches aja_j with score L≤(1+γ)Lmin⁡L \le (1 + \gamma) L_{\min}, add path to P\mathcal{P}
    Sort paths in P\mathcal{P} in ascending order of their accumulated scores
    return P\mathcal{P}
  7. Knowl 7 — Unified Researcher Profile Extraction via Conditional Random Fields

    model/method

    Automated extraction of structured researcher profiles from Web pages is formulated as a unified sequence tagging task using Conditional Random Fields (CRFs) based on an extended FOAF ontology schema capturing 19 profile properties (including Name, Affiliation, Position, Email, Phone, Address, Photo, Degree institutions, Dates, Majors, and Research Interests).

    The extraction pipeline consists of three phases:

    1. Relevant Page Identification: Given a researcher name, candidate Web pages retrieved via the Google API are classified into homepages/introducing pages versus irrelevant pages using a Support Vector Machine (SVM) binary classifier (achieving 92.39% F1-measure).
    2. Preprocessing and Tag Assignment: Web page contents are segmented into five unit types: standard unigram words, special words (URLs, emails, dates, numbers, special symbols identified via regular expressions), <image> tags, base noun phrase terms, and punctuation marks. Unit-specific candidate property tags are assigned to restrict the tag space.
    3. CRF Sequence Labeling: A linear-chain CRF is trained via Maximum Likelihood Estimation to find the optimal tag sequence Y∗=arg⁡max⁡YP(Y∣X)\mathbf{Y}^* = \arg\max_{\mathbf{Y}} P(\mathbf{Y} \mid \mathbf{X}) using 108,409 Boolean feature functions spanning three categories:
      • Content features: Token text, morphology, image dimensions, aspect ratio, color count, file format, filename tokens, ALT text, and face detection outputs.
      • Pattern features: Predefined positive/negative keywords, regex patterns, researcher name matches, and preceding line break counts.
      • Term features: Base noun phrases and vocabulary dictionary matches.
  8. Knowl 8 — Empirical Performance of Researcher Profile Extraction

    empirical result

    The unified CRF-based profile extraction framework was evaluated on 898 annotated researcher Web pages (obtained by filtering an initial random sample of 1,000 researcher names from ArnetMiner, labeled by 7 annotators with majority voting) and compared against rule induction via Amilcare (LP2LP^2) and an SVM classifier.

    Could not parse LaTeX table

    The unified CRF model achieved an average F1-score of 83.37%, outperforming Amilcare by +29.93% and SVM by +9.80%. Removing label transition features from the CRF caused performance to drop by 11.28% in F1-score (to 72.09%), confirming that modeling sequential dependencies between adjacent profile properties is critical. Feature ablation demonstrated that combining content, term, and pattern features yielded the highest performance (~83.4% F1), compared to content features alone (~71% F1), content+term (~75% F1), or content+pattern (~79% F1).

  9. Knowl 9 — Empirical Evaluation of HMRF Name Disambiguation and Relationship Types

    empirical result

    The HMRF name disambiguation method was evaluated on a benchmark dataset of 14 ambiguous person names spanning 861 publications and 161 actual distinct individuals, labeled by 5 human annotators using author homepages, affiliations, and emails. Performance was compared against a hierarchical clustering baseline.

    Could not parse LaTeX table

    The HMRF approach achieved an average F1-score of 91.39%, outperforming the baseline by +10.75%. Incremental ablation of the relationship graph showed:

    • Removing all relational edges dropped disambiguation F1 by 44.72% (to ~46.7%).
    • Adding CoPubvenue, Citation, CoAuthor, and τ\tau-CoAuthor sequentially improved F1 at each step, with the CoAuthor relationship providing the largest individual gain (+24.38% F1).
  10. Knowl 10 — Comparative Retrieval Performance for Academic Expertise Search

    empirical result

    Expertise search across papers, authors, and conferences was evaluated on a subset of ArnetMiner data (14,134 persons, 10,716 papers, 1,434 conferences) using frequent queries from system search logs, pooled top-30 results from ArnetMiner, Libra, and Rexa, and 4-grade human relevance judgments (T=80T = 80 topics for all topic models).

    Could not parse LaTeX table

    ACT1 achieved the highest performance across all evaluation metrics, reaching an overall Mean Average Precision (MAP) of 71.0% (compared to 61.0% for LM, 63.7% for Author-Topic, 63.9% for ACT2, and 63.8% for ACT3). For author search, ACT1 achieved 89.6% MAP and 91.4% P@5. Benchmark search engines Libra and Rexa achieved average MAP scores of 48.3% and 45.0% on the same query set.

  11. Knowl 11 — Limitations in Cluster Count Determination and Graph/Temporal Modeling

    limitation

    The ArnetMiner system architecture and underlying models have three main limitations identified by the authors:

    1. Preset Cluster Count in Name Disambiguation: The HMRF name disambiguation framework assumes that the true number of individuals kk associated with an ambiguous name is known and supplied empirically in advance, which cannot scale automatically to large-scale uncurated digital libraries.
    2. Absence of Link and Network Structure in Topic Models: The ACT probabilistic topic models do not incorporate explicit network structures, such as citation links or co-authorship graphs, directly into the generative topical process.
    3. Static Temporal Assumption: The topic models treat publication records statically, ignoring temporal dynamics such as the evolution of author research interests or the drifting themes of conferences over time.

Coverage note — None was omitted; all key contributed models (CRF extraction, HMRF name disambiguation, ACT1/ACT2/ACT3 topic models, LM-ACT expertise search, KL-divergence association search), algorithms, empirical evaluations, and stated limitations are fully represented.

References

  1. 1.L. A. Adamic and E. Adar. How to search a social network. Social Networks, 27:187–203, 2005.
  2. 2.C. Andrieu, N. de Freitas, A. Doucet, and M. I. Jordan. An introduction to mcmc for machine learning. Machine Learning, 50:5–43, 2003.
  3. 3.R. Baeza-Yates and B. Ribeiro-Neto. Modern Information Retrieval. ACM Press, 1999.
  4. 4.K. Balog, L. Azzopardi, and M. de Rijke. Formal models for expert finding in enterprise corpora. In Proc. of SIGIR’06, pages 43–55, 2006.
  5. 5.S. Basu, M. Bilenko, and R. J. Mooney. A probabilistic framework for semi-supervised clustering. In Proc. of KDD’04, pages 59–68, 2004.
  6. 6.R. Bekkerman and A. McCallum. Disambiguating web appearances of people in a social network. In Proc. of WWW’05, pages 463–470, 2005.
  7. 7.D. M. Blei and J. D. McAuliffe. Supervised topic models. In Proc. of NIPS’07, 2007.
  8. 8.D. M. Blei, A. Y. Ng, and M. I. Jordan. Latent dirichlet allocation. Journal of Machine Learning Research, 3:993–1022, 2003.
  9. 9.D. Brickley and L. Miller. Foaf vocabulary specification. In Namespace Document, http://xmlns.com/foaf/0.1/, September 2004.
  10. 10.C. Buckley and E. M. Voorhees. Retrieval evaluation with incomplete information. In Proc. of SIGIR’04, pages 25–32, 2004.
  11. 11.F. Ciravegna. An adaptive algorithm for information extraction from web-related texts. In Proc. of IJCAI’01 Workshop, August 2001.
  12. 12.C. Cortes and V. Vapnikn. Support-vector networks. Machine Learning, 20:273–297, 1995.
  13. 13.N. Craswell, A. P. de Vries, and I. Soboroff. Overview of the trec-2005 enterprise track. In TREC’05, pages 199–205, 2005.
  14. 14.H. Han, L. Giles, H. Zha, C. Li, and K. Tsioutsiouliklis. Two supervised learning approaches for name disambiguation in author citations. In Proc. of JCDL’04, pages 296–305, 2004.
  15. 15.H. Han, H. Zha, and C. L. Giles. Name disambiguation in author citations using a k-way spectral clustering method. In Proc. of JCDL’05, pages 334–343, 2005.
  16. 16.T. Hofmann. Collaborative filerting via gaussian probabilistic latent semantic analysis. In Proc.of SIGIR’03, pages 259–266, 1999.
  17. 17.T. Hofmann. Probabilistic latent semantic indexing. In Proc.of SIGIR’99, pages 50–57, 1999.
  18. 18.H. Kautz, B. Selman, and M. Shah. Referral web: Combining social networks and collaborative filtering. Communications of the ACM, 40(3):63–65, 1997.
  19. 19.T. Kristjansson, A. Culotta, P. Viola, and A. McCallum. Interactive information extraction with constrained conditional random fields. In Proc. of AAAI’04, 2004.
  20. 20.J. Lafferty, A. McCallum, and F. Pereira. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In Proc. of ICML’01, 2001.
  21. 21.A. McCallum. Multi-label text classification with a mixture model trained by em. In Proc. of AAAI’99 Workshop, 1999.
  22. 22.D. Mimno and A. McCallum. Expertise modeling for matching papers with reviewers. In Proc. of KDD’07, pages 500–509, 2007.
  23. 23.T. Minka. Estimating a dirichlet distribution. In Technique Report, http://research.microsoft.com/ minka/papers/dirichlet/, 2003.
  24. 24.Z. Nie, Y. Ma, S. Shi, J.-R. Wen, and W.-Y. Ma. Web object retrieval. In Proc. of WWW’07, pages 81–90, 2007.
  25. 25.M. Rosen-Zvi, T. Griffiths, M. Steyvers, and P. Smyth. The author-topic model for authors and documents. In Proc. of UAI’04, 2004.
  26. 26.M. Steyvers, P. Smyth, and T. Griffiths. Probabilistic author-topic models for information discovery. In Proc. of SIGKDD’04, 2004.
  27. 27.Y. F. Tan, M.-Y. Kan, and D. Lee. Search engine driven author disambiguation. In Proc. of JCDL’06, pages 314–315, 2006.
  28. 28.J. Tang, D. Zhang, and L. Yao. Social network extraction of academic researchers. In Proc. of ICDM’07, pages 292–301, 2007.
  29. 29.X. Wei and W. B. Croft. Lda-based document models for ad-hoc retrieval. In Proc. of SIGIR’06, pages 178–185, 2006.
  30. 30.E. Xun, C. Huang, and M. Zhou. A unified statistical model for the identification of english basenp. In Proc. of ACL’00, 2000.
  31. 31.X. Yin, J. Han, and P. Yu. Object distinction: Distinguishing objects with identical names. In Proc. of ICDE’2007, pages 1242–1246, 2007.
  32. 32.K. Yu, G. Guan, and M. Zhou. Resume information extraction with cascaded hybrid model. In Proc. of ACL’05, pages 499–506, 2005.

Citation

MLA
Tang, J., et al. “ArnetMiner”. Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2008, pp. 990–98, https://doi.org/10.1145/1401890.1402008.
APA
Tang, J., Zhang, J., Yao, L., Li, J., Zhang, L., & Su, Z. (2008). ArnetMiner. Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 990–998. https://doi.org/10.1145/1401890.1402008
Chicago
Tang, J., J. Zhang, L. Yao, J. Li, L. Zhang, and Z. Su. 2008. “ArnetMiner”. Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 990–98. https://doi.org/10.1145/1401890.1402008.
Harvard
Tang, J. et al. (2008) “ArnetMiner”, Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining. ACM, pp. 990–998. Available at: https://doi.org/10.1145/1401890.1402008.
Vancouver
1. Tang J, Zhang J, Yao L, Li J, Zhang L, Su Z (2008) ArnetMiner. In: Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining. ACM, pp 990–998

BibTeX

@inproceedings{Tang_2008, series={KDD08}, title={ArnetMiner: extraction and mining of academic social networks}, url={http://dx.doi.org/10.1145/1401890.1402008}, DOI={10.1145/1401890.1402008}, booktitle={Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining}, publisher={ACM}, author={Tang, Jie and Zhang, Jing and Yao, Limin and Li, Juanzi and Zhang, Li and Su, Zhong}, year={2008}, month=Aug, pages={990–998}, collection={KDD08} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF