Challenges of Big Data Analysis

Jianqing FanFang HanHan Liu

article2013National Science Review1,437 citationsNational Science Review 2015 Best Paper Award

Analyzes critical statistical and computational pitfalls in high-dimensional big data, demonstrating how issues like noise accumulation, spurious correlations, and incidental endogeneity invalidate standard inference methods.

Listen

Modern organizations and researchers face an unprecedented expansion of high-dimensional datasets across genomics, biomedical imaging, finance, and online networks. While these vast datasets offer valuable potential for detecting subtle subpopulation patterns and broad commonalities, conventional analytical tools fail when applied at this scale. The article evaluates the fundamental computational and statistical challenges created by massive sample sizes and high dimensionality, demonstrating why standard modeling assumptions break down and what new analytical frameworks are required.

The review synthesizes theoretical and applied developments across statistics, optimization, and distributed systems, evaluating empirical demonstrations in areas such as gene expression profiling and large-scale simulation. The analysis shows that high dimensionality severely degrades classical models by accumulating estimation noise until predictive power is lost, and it introduces high spurious correlations that produce false discoveries and underestimate error variances. Furthermore, the article demonstrates that incidental endogeneity—where unrelated variables unintentionally correlate with residual noise—naturally arises in wide datasets, invalidating the standard exogeneity assumptions on which most regularized models depend and leading to inconsistent variable selection.

These findings mean that applying legacy statistical models or unadjusted modern selectors to large-scale data creates substantial risks of false scientific conclusions, poor risk forecasting, and misallocated operational investments. To maintain statistical validity and computational tractability, practitioners must adopt specialized workflows. The article highlights regularized estimation techniques, independence screening to rapidly filter out irrelevant features, and generalized moment-based methods that directly account for endogeneity. Computationally, distributed architectures and randomized dimensionality reduction methods are necessary to scale processing linearly rather than exponentially.

Organizations should modernize their analytical pipelines by integrating two-stage screening and selection procedures, adopting algorithms that accommodate incidental correlations, and utilizing distributed infrastructure for large-scale storage. However, analysts must remain cautious regarding the persistent challenges of data heterogeneity, measurement biases from aggregated data sources, and temporal dependencies, which remain active areas requiring continued methodological development.

Cover for Challenges of Big Data Analysis

Abstract

Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data introduce unique computational and statistical challenges, including scalability and storage bottleneck, noise accumulation, spurious correlation, incidental endogeneity, and measurement errors. These challenges are distinguished and require new computational and statistical paradigm. This article give overviews on the salient features of Big Data and how these features impact on paradigm change on statistical and computational methods as well as computing architectures. We also provide various new perspectives on the Big Data analysis and computation. In particular, we emphasis on the viability of the sparsest solution in high-confidence set and point out that exogeneous assumptions in most statistical methods for Big Data can not be validated due to incidental endogeneity. They can lead to wrong statistical inferences and consequently wrong scientific conclusions.

Table of Contents

  • INTRODUCTION
  • BACKGROUND
  • GOALS AND CHALLENGES OF ANALYZING BIG DATA
  • PARADIGM SHIFTS
  • ORGANIZATION OF THIS PAPER
  • RISE OF BIG DATA
  • Genomics
  • Neuroscience
  • Economics and finance
  • Other applications
  • SALIENT FEATURES OF BIG DATA
  • Heterogeneity
  • Noise accumulation
  • Spurious correlation
  • Incidental endogeneity
  • IMPACT ON STATISTICAL THINKING
  • Penalized quasi-likelihood
  • Sparsest solution in high confidence set
  • Independence screening
  • Dealing with incidental endogeneity
  • IMPACT ON COMPUTING INFRASTRUCTURE
  • Hadoop distributed file system
  • MapReduce
  • Cloud computing
  • IMPACT ON COMPUTATIONAL METHODS
  • First-order methods for nonsmooth optimization
  • Dimension reduction and random projection
  • CONCLUSIONS AND FUTURE PERSPECTIVES
  • ACKNOWLEDGEMENTS
  • FUNDING
  • REFERENCES

Knowls

  1. Knowl 1 — Sparsest Solution in High Confidence Set Formulation

    model/method

    The sparsest solution in high confidence set (SSHCS) framework separates the empirical data information (captured by score or gradient bounds) from the structural sparsity assumption on model parameters.

    Let ℓn(β)\ell_n(\beta) denote an empirical loss, negative quasi-likelihood, or goodness-of-fit function for a parameter vector β∈Rd\beta \in \mathbb{R}^d, with expected loss ℓ(β)=E[ℓn(β)]\ell(\beta) = \mathbb{E}[\ell_n(\beta)] whose population parameter vector β0\beta_0 satisfies ℓ′(β0)=0\ell'(\beta_0) = \mathbf{0}. A high-confidence set Cn⊂RdC_n \subset \mathbb{R}^d for β0\beta_0 is defined via the L∞L_\infty-norm of the gradient vector ℓn′(β)\ell'_n(\beta):

    Cn={β∈Rd:∥ℓn′(β)∥∞≤γn}C_n = \left\{ \beta \in \mathbb{R}^d : \|\ell'_n(\beta)\|_\infty \le \gamma_n \right\}

    where γn>0\gamma_n > 0 is a threshold chosen to guarantee a confidence level of at least 1−δn1 - \delta_n with δn→0\delta_n \to 0 as sample size n→∞n \to \infty, such that:

    P(β0∈Cn)=P(∥ℓn′(β0)∥∞≤γn)≥1−δn.\mathbb{P}(\beta_0 \in C_n) = \mathbb{P}\left( \|\ell'_n(\beta_0)\|_\infty \le \gamma_n \right) \ge 1 - \delta_n.

    Assuming the true parameter β0\beta_0 is sparse, the estimator is defined as the sparsest vector residing in the high-confidence set:

    β^=arg⁡min⁡β∈Cn∥β∥1=arg⁡min⁡∥ℓn′(β)∥∞≤γn∥β∥1.\hat{\beta} = \arg\min_{\beta \in C_n} \|\beta\|_1 = \arg\min_{\|\ell'_n(\beta)\|_\infty \le \gamma_n} \|\beta\|_1.

    When ℓn(β)\ell_n(\beta) is convex, CnC_n is a convex set and the estimation problem is a convex program. For a linear model y=Xβ+εy = X\beta + \varepsilon with quadratic loss ℓn(β)=12∥y−Xβ∥22\ell_n(\beta) = \frac{1}{2}\|y - X\beta\|_2^2, the high-confidence set is Cn={β:∥XT(y−Xβ)∥∞≤γn}C_n = \{\beta : \|X^T(y - X\beta)\|_\infty \le \gamma_n\} with γn≥∥XTε∥∞\gamma_n \ge \|X^T\varepsilon\|_\infty, which recovers the Dantzig selector. This principle also unifies constrained L1L_1-minimization approaches for precision matrix estimation (CLIME) and linear programming discriminant analysis (LPDA), and extends to problems with measurement errors and endogeneity by modifying the defining estimating equations in CnC_n.

  2. Knowl 2 — Incidental Endogeneity and Inconsistency of Regularized Estimators

    theoretical result

    In high-dimensional linear regression Y=∑j=1dβjXj+εY = \sum_{j=1}^d \beta_j X_j + \varepsilon with true sparse support S={j:βj≠0}S = \{j : \beta_j \neq 0\}, standard penalized estimators (such as Lasso, SCAD, and MCP) rely on the strict exogeneity assumption:

    E(εXj)=0for all j=1,…,d.\mathbb{E}(\varepsilon X_j) = 0 \quad \text{for all } j = 1, \dots, d.

    This condition is a necessary prerequisite for any penalized least-squares estimator to achieve variable selection consistency.

    In ultra-high dimensions (d≫nd \gg n), incidental endogeneity inevitably arises: even when the true active predictors {Xj}j∈S\{X_j\}_{j \in S} are uncorrelated with the disturbance ε\varepsilon, some scientifically unrelated covariates XkX_k for k∉Sk \notin S will exhibit genuine, non-zero population correlations with ε\varepsilon due to omitted confounders, systematic measurement errors, or latent selection biases introduced during data collection across heterogeneous sources. When E(εXk)≠0\mathbb{E}(\varepsilon X_k) \neq 0 for some k∉Sk \notin S, the full exogeneity condition fails, causing classical penalized least-squares estimators (Lasso, SCAD, MCP) to lose variable selection consistency and produce biased parameter estimates.

  3. Knowl 3 — Focused Generalized Method of Moments under Over-Identification

    model/method

    The Focused Generalized Method of Moments (FGMM) is an estimation framework designed to resolve incidental endogeneity in high-dimensional linear regression Y=∑j∈SβjXj+εY = \sum_{j \in S} \beta_j X_j + \varepsilon. Rather than requiring full exogeneity across all dd features (E[εXj]=0\mathbb{E}[\varepsilon X_j] = 0 for all j=1,…,dj = 1,\dots,d), FGMM requires only that the conditional mean restriction holds on the true active support S={j:βj≠0}S = \{j : \beta_j \neq 0\}:

    E(ε∣{Xj}j∈S)=E(Y−∑j∈SβjXj  ∣  {Xj}j∈S)=0.\mathbb{E}\left(\varepsilon \mid \{X_j\}_{j \in S}\right) = \mathbb{E}\left(Y - \sum_{j \in S} \beta_j X_j \;\Bigg|\; \{X_j\}_{j \in S}\right) = 0.

    This implies the focused over-identification moment conditions:

    E(εXj)=0andE(εXj2)=0for all j∈S.\mathbb{E}(\varepsilon X_j) = 0 \quad \text{and} \quad \mathbb{E}(\varepsilon X_j^2) = 0 \quad \text{for all } j \in S.

    By optimizing moment conditions targeted specifically to the active subset SS rather than penalizing least-squares across all covariates, FGMM consistently identifies the true active set of variables SS and produces valid parameter estimates, even when unselected covariates XkX_k (k∉Sk \notin S) are correlated with the residual noise ε\varepsilon.

  4. Knowl 4 — Spurious Correlation and Variance Underestimation in High Dimensions

    theoretical result

    In high dimensions where dimension dd is large relative to sample size nn, independent random variables exhibit large sample correlations. For independent Gaussian vectors X=(X1,…,Xd)T∼Nd(0,Id)X = (X_1, \dots, X_d)^T \sim \mathcal{N}_d(\mathbf{0}, I_d), the maximum absolute sample correlation between a designated variable X1X_1 and the remaining variables:

    r^=max⁡j≥2∣Corr^(X1,Xj)∣\hat{r} = \max_{j \ge 2} |\widehat{\text{Corr}}(X_1, X_j)|

    as well as the maximum multiple correlation between X1X_1 and linear combinations of any size-kk subset S⊂{2,…,d}S \subset \{2, \dots, d\}:

    R^=max⁡∣S∣=kmax⁡{βj}j=1k∣Corr^(X1,∑j∈SβjXj)∣\hat{R} = \max_{|S|=k} \max_{\{\beta_j\}_{j=1}^k} \left| \widehat{\text{Corr}}\left(X_1, \sum_{j \in S} \beta_j X_j\right) \right|

    increase toward 1 as dd grows for a fixed nn.

    In linear regression y=Xβ+εy = X\beta + \varepsilon with residual variance Var(ε)=σ2In\text{Var}(\varepsilon) = \sigma^2 I_n, when an active feature subset S^\hat{S} is selected by data-driven algorithms in the presence of spurious correlations, the standard post-selection residual variance estimator:

    σ^2=yT(In−PS^)yn−∣S^∣\hat{\sigma}^2 = \frac{y^T (I_n - P_{\hat{S}}) y}{n - |\hat{S}|}

    (where PS^=XS^(XS^TXS^)−1XS^TP_{\hat{S}} = X_{\hat{S}} (X_{\hat{S}}^T X_{\hat{S}})^{-1} X_{\hat{S}}^T is the projection matrix onto the columns of XS^X_{\hat{S}}) severely underestimates the true variance σ2\sigma^2. This bias leads to invalid hypothesis tests, overly optimistic confidence intervals, and false scientific discoveries unless corrected by sample-splitting procedures such as refitted cross-validation.

  5. Knowl 5 — Local Linear Approximation Algorithm for Non-Convex Penalties

    algorithm

    The Local Linear Approximation (LLA) algorithm minimizes folded-concave penalized objective functions of the form:

    min⁡β∈Rd{ℓn(β)+∑j=1dPλ,γ(∣βj∣)}\min_{\beta \in \mathbb{R}^d} \left\{ \ell_n(\beta) + \sum_{j=1}^d P_{\lambda,\gamma}(|\beta_j|) \right\}

    where ℓn(β)\ell_n(\beta) is a convex empirical loss and Pλ,γ(⋅)P_{\lambda,\gamma}(\cdot) is a non-convex penalty (such as SCAD or MCP) with tuning parameter λ>0\lambda > 0 and concavity parameter γ>0\gamma > 0. At step k+1k+1, LLA replaces Pλ,γ(∣βj∣)P_{\lambda,\gamma}(|\beta_j|) with its first-order Taylor expansion around the current magnitude ∣β^j(k)∣|\hat{\beta}_j^{(k)}|, reducing the non-convex optimization problem to a sequence of weighted L1L_1-penalized convex programs.

    Input: Convex loss function ℓn(⋅)\ell_n(\cdot), penalty derivative Pλ,γ′(⋅)P'_{\lambda,\gamma}(\cdot), initial estimate β^(0)=0\hat{\beta}^{(0)} = \mathbf{0}, tolerance τ>0\tau > 0, max iterations KmaxK_{max}
    Output: Sparse parameter estimate β^\hat{\beta}
    k←0k \leftarrow 0
    repeat
        for j←1j \leftarrow 1 to dd do
            wk,j←Pλ,γ′(∣β^j(k)∣)w_{k,j} \leftarrow P'_{\lambda,\gamma}(|\hat{\beta}_j^{(k)}|)
        end for
        β^(k+1)←arg⁡min⁡β∈Rd{ℓn(β)+∑j=1dwk,j∣βj∣}\hat{\beta}^{(k+1)} \leftarrow \arg\min_{\beta \in \mathbb{R}^d} \left\{ \ell_n(\beta) + \sum_{j=1}^d w_{k,j} |\beta_j| \right\}
        k←k+1k \leftarrow k + 1
    until ∥β^(k)−β^(k−1)∥2≤τ\|\hat{\beta}^{(k)} - \hat{\beta}^{(k-1)}\|_2 \le \tau or k≥Kmaxk \ge K_{max}
    return β^(k)\hat{\beta}^{(k)}
  6. Knowl 6 — Sure Independence Screening for Ultra-High Dimensional Feature Space

    model/method

    Sure Independence Screening (SIS) is a fast, marginal dimension reduction technique designed for ultra-high-dimensional regimes (d≫nd \gg n, with log⁡d=O(na)\log d = O(n^a) for a∈(0,1)a \in (0,1)). It scales down the feature space from ultra-high dimension dd to a manageable dimension d0<nd_0 < n, allowing subsequent joint regularized regression methods to be applied.

    For standardized covariates Xj∈RnX_j \in \mathbb{R}^n (j=1,…,dj = 1, \dots, d) and response vector y∈Rny \in \mathbb{R}^n, SIS estimates the univariate marginal regression coefficients:

    β^jM=arg⁡min⁡βj12n∥y−Xjβj∥22=XjTy∥Xj∥22,j=1,…,d.\hat{\beta}_j^M = \arg\min_{\beta_j} \frac{1}{2n}\|y - X_j \beta_j\|_2^2 = \frac{X_j^T y}{\|X_j\|_2^2}, \quad j = 1, \dots, d.

    Covariates are ranked according to their marginal magnitude ∣β^jM∣|\hat{\beta}_j^M| (or equivalently, marginal sample correlation with yy or marginal deviance reduction). The screened subset of surviving features is:

    S^={j∈{1,…,d}:∣β^jM∣≥δ}\hat{S} = \left\{ j \in \{1,\dots,d\} : |\hat{\beta}_j^M| \ge \delta \right\}

    for a predefined threshold δ>0\delta > 0, or by retaining the top d0=⌊n/log⁡n⌋d_0 = \lfloor n / \log n \rfloor ranked covariates.

    The computational complexity of SIS is O(nd)O(nd), scaling linearly with sample size and feature dimension. Under conditions on marginal signal strength and covariance bounds, SIS satisfies the sure screening property: P(S0⊆S^)→1\mathbb{P}(S_0 \subseteq \hat{S}) \to 1 as n→∞n \to \infty, where S0={j:β0,j≠0}S_0 = \{j : \beta_{0,j} \neq 0\} is the true active support set.

  7. Knowl 7 — Noise Accumulation in High-Dimensional Classification

    theoretical result

    In high-dimensional classification, estimating a large number of parameters corresponding to weak or non-informative features leads to noise accumulation that degrades decision boundaries and can completely overwhelm true signals.

    Consider two Gaussian classes in dimension dd with equal covariance:

    X1,…,Xn∼Nd(μ1,Id)andY1,…,Yn∼Nd(μ2,Id)X_1, \dots, X_n \sim \mathcal{N}_d(\mu_1, I_d) \quad \text{and} \quad Y_1, \dots, Y_n \sim \mathcal{N}_d(\mu_2, I_d)

    where μ1=0\mu_1 = \mathbf{0} and μ2\mu_2 is ss-sparse (e.g., only the first s=10s = 10 coordinates have value c>0c > 0, and the remaining d−sd - s coordinates are zero). When all dd features are used to build a classification or projection rule (such as principal component projection):

    1. For small feature subsets m≈sm \approx s, the classifier achieves high discriminative power.
    2. For intermediate m>sm > s, accumulated signal from the ss informative features compensates for the noise of the m−sm - s null features.
    3. When m≫sm \gg s (e.g., m=1000m = 1000), the accumulated estimation noise from the null features dominates the true signal, causing conventional classification rules using all features to perform no better than random guessing.

    This behavior demonstrates why sparse modeling and feature selection are required in high-dimensional classification and regression.

  8. Knowl 8 — Random Projection for High-Dimensional Linear Dimension Reduction

    model/method

    Random Projection (RP) is a computationally efficient linear dimension reduction technique that maps an n×dn \times d data matrix DD to a lower-dimensional n×kn \times k matrix D~RP\tilde{D}^{RP} with k≪dk \ll d while approximately preserving pairwise Euclidean distances between all nn data points.

    Given a random matrix R∈Rd×kR \in \mathbb{R}^{d \times k} whose columns have unit Euclidean norms (or entries drawn i.i.d. from Gaussian or scaled discrete distributions ±1\pm 1), the compressed representation is computed via matrix multiplication:

    D~RP=DR.\tilde{D}^{RP} = D R.

    Theoretical justifications rely on the Johnson-Lindenstrauss lemma and the geometric property of high-dimensional spaces that any finite set of kk random vectors in Rd\mathbb{R}^d are almost mutually orthogonal when dd is large, ensuring RTR≈IkR^T R \approx I_k without requiring the computationally expensive Gram-Schmidt orthogonalization.

    Whereas Principal Component Analysis (PCA) computes the leading kk eigenvectors U^k∈Rd×k\hat{U}_k \in \mathbb{R}^{d \times k} of the sample covariance matrix with computational complexity O(d2n+d3)O(d^2 n + d^3), Random Projection has a computational complexity of O(ndk)O(ndk), scaling strictly linearly in nn, dd, and kk. In high dimensions, RP provides equal or superior pairwise distance preservation compared to PCA at a fraction of the computational cost.

  9. Knowl 9 — Subpopulation Heterogeneity via High-Dimensional Mixture Models

    model/method

    When Big Data are formed by aggregating datasets from diverse subpopulations, the overall population is characterized by substantial heterogeneity modeled as a finite mixture of conditional distributions:

    p(y∣x)=∑j=1mλjpj(y; θj(x))p(y \mid x) = \sum_{j=1}^m \lambda_j p_j\left(y;\, \theta_j(x)\right)

    where mm is the number of subpopulations, λj≥0\lambda_j \ge 0 (with ∑j=1mλj=1\sum_{j=1}^m \lambda_j = 1) is the mixing proportion of the jj-th subpopulation, and pj(y; θj(x))p_j(y;\, \theta_j(x)) is the conditional probability distribution of the response yy given covariates x∈Rdx \in \mathbb{R}^d with subpopulation-specific parameter vector θj(x)\theta_j(x).

    In moderate sample sizes nn, rare subpopulations with small proportions λj≪1\lambda_j \ll 1 yield small effective sub-samples nλjn\lambda_j, causing their parameters to be unidentifiable or discarded as outliers. In Big Data regimes where nn is massive, nλjn\lambda_j can be moderately large even for very small λj\lambda_j, enabling precise statistical inference for rare subpopulations. In high-dimensional feature spaces (d≫nλjd \gg n\lambda_j), estimating this mixture model requires regularized estimation across each subpopulation component to prevent overfitting and noise accumulation.

Coverage note — General descriptions of distributed infrastructure frameworks (HDFS architecture, MapReduce programming stages, and cloud computing principles) were omitted as standalone knowls because they represent standard computer science platforms reviewed by the paper rather than its statistical and methodological contributions.

References

  1. 1.Stein, L. The case for cloud computing in genome informatics. Genome Biol 2010; 11: 207.
  2. 2.Donoho, D. High-dimensional data analysis: the curses and blessings of dimensionality. In: The American Mathematical Society Conference, Los Angeles, CA, United States, 7–12 August 2000.
  3. 3.Bickel, P. Discussion on the paper ‘Sure independence screening for ultrahigh dimensional feature space’ by Fan and Lv. J Roy Stat Soc B 2008; 70: 883–4.
  4. 4.Fan, J and Fan, Y. High dimensional classification using features annealed independence rules. Ann Stat 2008: 36: 2605–37.
  5. 5.Pittelkow, PH and Ghosh, M. Theoretical measures of relative performance of classifiers for high dimensional data with small sample sizes. J Roy Stat Soc B 2008; 70: 159–73.
  6. 6.Tibshirani, R. Regression shrinkage and selection via the lasso. J Roy Stat Soc B 1996; 58: 267–88.
  7. 7.Chen, S, Donoho, D and Saunders, M. Atomic decomposition by basis pursuit. SIAM J Sci Comput 1998; 20: 33–61.
  8. 8.Fan, J and Li, R. Variable selection via nonconcave penalized likelihood and its oracle properties. J Am Stat Assoc 2001; 96: 1348–60.
  9. 9.Candes, E and Tao, T. The Dantzig selector: statistical estimation when p is much larger than n. Ann Stat 2007; 35: 2313–51.
  10. 10.Zhang, C-H. Nearly unbiased variable selection under minimax concave penalty. Ann Stat 2010; 38: 894–942.
  11. 11.Fan, J and Lv, J. Sure independence screening for ultrahigh dimensional feature space (with discussion). J Roy Stat Soc B 2008; 70: 849–911.
  12. 12.Hall, P and Miller, H. Using generalized correlation to effect variable selection in very high dimensional problems. J Comput Graph Stat 2009; 18: 533–50.
  13. 13.Genovese, C, Jin, J and Wasserman, L et al. A comparison of the lasso and marginal regression. J Mach Learn Res 2012; 13: 2107–43.
  14. 14.Fan, J, Guo, S and Hao, N. Variance estimation using refitted cross-validation in ultrahigh dimensional regression. J Roy Stat Soc B 2012; 74: 37–65.
  15. 15.Liao, Y and Jiang, W. Posterior consistency of nonparametric conditional moment restricted models. Ann Stat 2011; 39: 3003–31.
  16. 16.Fan, J and Liao, Y. Endogeneity in ultrahigh dimension. Technical report. Princeton University, 2012.
  17. 17.Fan, J. Features of big data and sparsest solution in high confidence set. Technical report. Princeton University, 2013.
  18. 18.Donoho, D and Elad, M. Optimally sparse representation in general (nonorthogonal) dictionaries via L1L_1 minimization. Proc Natl Acad Sci USA 2003; 100: 2197–202.
  19. 19.Efron, B, Hastie, T and Johnstone, I et al. Least angle regression. Ann Stat 2004; 32: 407–99.
  20. 20.Friedman, J and Popescu, B. Gradient directed regularization for linear regression and classification. Technical report. Stanford University, 2003.
  21. 21.Fu, WJ. Penalized regressions: the bridge versus the lasso. J Comput Graph Stat 1998; 7: 397–416.
  22. 22.Wu, T and Lange, K. Coordinate descent algorithms for lasso penalized regression. Ann Appl Stat 2008; 2: 224–44.
  23. 23.Daubechies, I, Defrise, M and De Mol, C. An iterative thresholding algorithm for linear inverse problems with a sparsity constraint. Commun Pur Appl Math 2004; 57: 1413–57.
  24. 24.Beck, A and Teboulle, M. A fast iterative shrinkage-thresholding algorithm for linear inverse problems. SIAM J Imaging Sciences 2009; 2: 183–202.
  25. 25.Lange, K, Hunter, D and Yang, I. Optimization transfer using surrogate objective functions. J Comput Graph Stat 2000; 9: 1–20.
  26. 26.Hunter, D and Li, R. Variable selection using MM algorithms. Ann Stat 2005; 33: 1617–42.
  27. 27.Zou, H and Li, R. One-step sparse estimates in nonconcave penalized likelihood models. Ann Stat 2008; 36: 1509–33.
  28. 28.Fan, J, Samworth, R and Wu, Y. Ultrahigh dimensional feature selection: beyond the linear model. J Mach Learn Res 2009; 10: 2013–38.
  29. 29.Boyd, S, Parikh, N and Chu, E et al. Distributed optimization and statistical learning via the alternating direction method of multipliers. Found Trends Mach Learn 2011; 3: 1–122.
  30. 30.Bradley, J, Kyrola, A and Bickson, D et al. Parallel coordinate descent for L1L_1-regularized loss minimization. arXiv:1105.5379, 2011.
  31. 31.Low, Y, Bickson, and Dand Gonzalez, J et al. Distributed graphlab: a framework for machine learning and data mining in the cloud. Proc Int Conf VLDB Endowment 2012; 5: 716–27.
  32. 32.Worthey, E, Mayer, A and Syverson, G et al. Making a definitive diagnosis: successful clinical application of whole exome sequencing in a child with intractable inflammatory bowel disease. Genet Med 2010; 13: 255–62.
  33. 33.Chen, R, Mias, G and Li-Pook-Than, J et al. Personal omics profiling reveals dynamic molecular and medical phenotypes. Cell 2012; 148: 1293–307.
  34. 34.Cohen, J, Kiss, R and Pertsemlidis, A et al. Multiple rare alleles contribute to low plasma levels of HDL cholesterol. Science 2004; 305: 869–72.
  35. 35.Han, F and Pan, W. A data-adaptive sum test for disease association with multiple common or rare variants. Hum Hered 2010; 70: 42–54.
  36. 36.Bickel, P, Brown, J and Huang, H et al. An overview of recent developments in genomics and associated statistical methods. Philos T R Soc A 2009; 367: 4313–37.
  37. 37.Leek, J and Storey, J. Capturing heterogeneity in gene expression studies by surrogate variable analysis. PLoS Genet 2007; 3: e161.
  38. 38.Benjamini, Y and Hochberg, Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J Roy Stat Soc B 1995; 57: 289–300.
  39. 39.Storey, J. The positive false discovery rate: a Bayesian interpretation and the q-value. Ann Stat 2003; 31: 2013–35.
  40. 40.Schwartzman, A, Dougherty, R and Lee, J et al. Empirical null and false discovery rate analysis in neuroimaging. Neuroimage 2009; 44: 71–82.
  41. 41.Efron, B. Correlated z-values and the accuracy of large-scale statistical estimates. J Am Stat Assoc 2010; 105: 1042–55.
  42. 42.Fan, J, Han, X and Gu, W. Control of the false discovery rate under arbitrary covariance dependence. J Am Stat Assoc 2012; 107: 1019–45.
  43. 43.Edgar, R, Domrachev, M and Lash, AE. Gene expression omnibus: NCBI gene expression and hybridization array data repository. Nucleic Acids Res 2002; 30: 207–10.
  44. 44.Jonides, J, Nee, D and Berman, M. What has functional neuroimaging told us about the mind? So many examples little space. Cortex 2006; 42: 414–7.
  45. 45.Visscher, K and Weissman, D. Would the field of cognitive neuroscience be advanced by sharing functional MRI data? BMC Med 2011; 9: 34.
  46. 46.Milham, M, Mennes, M and Gutman, D et al. The International Neuroimaging Data-sharing Initiative (INDI) and the Functional Connectomes Project. 17th Annual Meeting of the Organization for Human Brain Mapping, Quebec City, 2011.
  47. 47.Di Martino, A., Yan, CG and Li, Q et al. The autism brain imaging data exchange: Towards a large-scale evaluation of the intrinsic brain architecture in autism. Mol Psychiatry 2013, doi:10.1038/mp.2013.78.
  48. 48.The ADHD-200 Consortium. The ADHD-200 consortium: a model to advance the translational potential of neuroimaging in clinical neuroscience. Front Syst Neurosci 2012; 6: 62.
  49. 49.Fritsch, V, Varoquaux, G and Thyreau, B et al. Detecting outliers in high-dimensional neuroimaging datasets with robust covariance estimators. Med Image Anal 2012; 16: 1359–70.
  50. 50.Song, S and Bickel, P. Large vector auto regressions. arXiv:1106.3915, 2011.
  51. 51.Han, F and Liu, H. Transition matrix estimation in high dimensional time series. In: The 30th International Conference on Machine Learning, Atlanda, GA, USA, 16–21 June, 2013.
  52. 52.Cochrane, J. Asset Pricing. Princeton, NJ: Princeton University Press, 2001.
  53. 53.Dempster, M. Risk Management: Value at Risk and Beyond. Cambridge: Cambridge University Press, 2002.
  54. 54.Stock, J and Watson, M. Forecasting using principal components from a large number of predictors. J Am Stat Assoc 2002; 97: 1167–79.
  55. 55.Bai, J and Ng, S. Determining the number of factors in approximate factor models. Econometrica 2002; 70: 191–221.
  56. 56.Bai, J. Inferential theory for factor models of large dimensions. Econometrica 2003; 71: 135–71.
  57. 57.Forni, M, Hallin, M and Lippi, M et al. The generalized dynamic factor model: one-sided estimation and forecasting. J Am Stat Assoc 2005; 100: 830–40.
  58. 58.Fan, J, Fan, Y and Lv, J. High dimensional covariance matrix estimation using a factor model. J. Econometrics 2008; 147: 186–97.
  59. 59.Bickel, P and Levina, E. Covariance regularization by thresholding. Ann Stat 2008; 36: 2577–604.
  60. 60.Cai, T and Liu, W. Adaptive thresholding for sparse covariance matrix estimation. J Am Stat Assoc 2011; 106: 672–84.
  61. 61.Agarwal, A, Negahban, S and Wainwright, M. Noisy matrix decomposition via convex relaxation: optimal rates in high dimensions. Ann Stat 2012; 40: 1171–97.
  62. 62.Liu, H, Han, F and Yuan, M et al. High-dimensional semiparametric Gaussian copula graphical models. Ann Stat 2012; 40: 2293–326.
  63. 63.Xue, L and Zou, H. Regularized rank-based estimation of high-dimensional nonparanormal graphical models. Ann Stat 2012; 40: 2541–71.
  64. 64.Liu, H, Han, F and Zhang, C-H. Transelliptical graphical models. In: The 25th Conference in Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA, 3–8 December, 2012.
  65. 65.Fan, J, Liao, Y and Mincheva, M. Large covariance estimation by thresholding principal orthogonal complements. J Roy Stat Soc B 2013; 75: 603–80.
  66. 66.Pourahmadi, M. Modern Methods to Covariance Estimation: With High-Dimensional Data. New York: Wiley, 2013.
  67. 67.Aramaki, E, Maskawa, S and Morita, M. Twitter catches the flu: detecting influenza epidemics using twitter. In: The Conference on Empirical Methods in Natural Language Processing, Edinburgh, UK, 27–29 July, 2011.
  68. 68.Bollen, J, Mao, H and Zeng, X. Twitter mood predicts the stock market. J Comput Sci 2011; 2: 1–8.
  69. 69.Asur, S and Huberman, B. Predicting the future with social media. In: The IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT), Toronto, Canada, 31 August–3 September, 2010.
  70. 70.Khalili, A and Chen, J. Variable selection in finite mixture of regression models. J Am Stat Assoc 2007; 102: 1025–38.
  71. 71.Städler, N, Bühlmann, P and van de Geer, S. ℓ1\ell_1-penalization for mixture regression models. Test 2010; 19: 209–56.
  72. 72.Hastie, T, Tibshirani, R and Friedman, J. The Elements of Statistical Learning. Berlin: Springer, 2009.
  73. 73.Bühlmann, P and van de Geer, S. Statistics for High-Dimensional Data: Methods, Theory and Applications. Berlin: Springer, 2011.
  74. 74.Cai, T and Jiang, T. Phase transition in limiting distributions of coherence of high-dimensional random matrices. J Multivariate Anal 2012; 107: 24–39.
  75. 75.Engle, R, Hendry, D and Richard, J-F. Exogeneity. Econometrica 1983; 51: 277–304.
  76. 76.Brazma, A, Parkinson, H and Sarkans, U et al. ArrayExpress—a public repository for microarray gene expression data at the EBI. Nucleic Acids Res 2003; 31: 68–71.
  77. 77.Valiathan, R, Marco, M and Leitinger, B et al. Discoidin domain receptor tyrosine kinases: new players in cancer progression. Cancer Metastasis Rev 2012; 31: 295–321.
  78. 78.Akaike, H. A new look at the statistical model identification. IEEE Trans Automat Control 1974; 19: 716–23.
  79. 79.Barron, A, Birgé, L and Massart, P. Risk bounds for model selection via penalization. Probab Theory Related Fields 1999; 113: 301–413.
  80. 80.Antoniadis, A. Wavelets in statistics: a review. J Ital Stat Soc 1997; 6: 97–130.
  81. 81.Antoniadis, A and Fan, J. Regularization of wavelet approximations. J Am Stat Assoc 2001; 96: 939–55.
  82. 82.Donoho, D and Johnstone, J. Ideal spatial adaptation by wavelet shrinkage. Biometrika 1994; 81: 425–55.
  83. 83.Liang, K-Y and Zeger, S. Longitudinal data analysis using generalized linear models. Biometrika 1986; 73: 13–22.
  84. 84.Cai, T, Liu, W and Luo, X. A constrained L1L_1 minimization approach to sparse precision matrix estimation. J Am Stat Assoc 2011; 106: 594–607.
  85. 85.Cai, T and Liu, W. A direct estimation approach to sparse linear discriminant analysis. J Am Stat Assoc 2011; 106: 1566–77.
  86. 86.Bickel, P, Ritov, Y and Tsybakov, A. Simultaneous analysis of lasso and Dantzig selector. Ann Stat 2009; 37: 1705–32.
  87. 87.Gautier, E and Tsybakov, A. High-dimensional instrumental variables regression and confidence sets. arXiv:1105.2454, 2011.
  88. 88.Fan, J and Song, R. Sure independence screening in generalized linear models with NP-dimensionality. Ann Stat 2010; 38: 3567–604.
  89. 89.Fan, J, Feng, Y and Song, R. Nonparametric independence screening in sparse ultra-high dimensional additive models. J Am Stat Assoc 2011; 106: 544–57.
  90. 90.Zhao, S and Li, Y. Principled sure independence screening for Cox models with ultra-high-dimensional covariates. J Multivariate Anal 2012; 105: 397–411.
  91. 91.Li, R, Zhong, W and Zhu, L. Feature screening via distance correlation learning. J Am Stat Assoc 2012; 107: 1129–39.
  92. 92.Li, G, Peng, H and Zhang, J et al. Robust rank correlation based screening. Ann Stat 2012; 40: 1846–77.
  93. 93.Ke, T, Jin, J and Fan, J. Covariance assisted screening and estimation. arXiv:1205.4645, 2012.
  94. 94.Boyd, S and Vandenberghe, L. Convex Optimization. Cambridge: Cambridge University Press, 2004.
  95. 95.Fodor, I. A survey of dimension reduction techniques. Technical report. US Department of Energy, 2002.
  96. 96.Avriel, M. Nonlinear Programming: Analysis and Methods. New York: Courier Dover, 2003.
  97. 97.Friedman, J, Hastie, T and Höfling, H et al. Pathwise coordinate optimization. Ann Appl Stat 2007; 1: 302–32.
  98. 98.Nesterov, Y. Efficiency of coordinate descent methods on huge-scale optimization problems. SIAM J Optim 2012; 22: 341–62.
  99. 99.Candes, E, Wakin, M and Boyd, S. Enhancing sparsity by reweighted L1L_1 minimization. J Fourier Anal Appl 2008; 14: 877–905.
  100. 100.Wang, Z, Liu, H and Zhang, T. Optimal computational and statistical rates of convergence for sparse nonconvex learning problems. arXiv:1306.4960, 2013.
  101. 101.Agarwal, A, Negahban, S and Wainwright, M. Fast global convergence of gradient methods for high-dimensional statistical recovery. Ann Stat 2012; 40: 2452–82.
  102. 102.Loh, P-L and Wainwright, M. Regularized M-estimators with nonconvexity: statistical and algorithmic theory for local optima. arXiv:1305.2436, 2013.
  103. 103.Golub, G and Van Loan, C. Matrix Computations. Baltimore, MD: The Johns Hopkins University Press, 2012.
  104. 104.Johnson, W and Lindenstrauss, J. Extensions of Lipschitz mappings into a Hilbert space. Contemp Math 1984; 26: 189–206.
  105. 105.Donoho, D. Compressed sensing. IEEE Trans Inform Theory 2006; 52: 1289–306.
  106. 106.Tsaig, Y and Donoho, D. Extensions of compressed sensing. Signal Process 2006; 86: 549–71.
  107. 107.Lustig, M, Donoho, D and Pauly, J. Sparse MRI: the application of compressed sensing for rapid MR imaging. Magn Reson Med 2007; 58: 1182–95.
  108. 108.Figueiredo, M, Nowak, R and Wright, S. Gradient projection for sparse reconstruction: application to compressed sensing and other inverse problems. IEEE J Sel Top Signal Process 2007; 1: 586–97.
  109. 109.Candes, E and Wakin, M. An introduction to compressive sampling. Signal Process Magazine 2008; 25: 21–30.
  110. 110.Marks, R and Zurada, J. Computational Intelligence: Imitating Life. Piscataway, NJ: IEEE, 1994.
  111. 111.Achlioptas, D. Database-friendly random projections. In: The 20th ACM SIGMOD-SIGACT-SIGART Symposium on Principles of Database Systems, Dallas, TX, USA, 16–18 May, 2001.
  112. 112.Deerwester, S, Dumais, S and Furnas, GR. Indexing by latent semantic analysis. J Assn Inf Sci 1990; 41: 391–407.
  113. 113.Rao, K, Yip, P and Britanak, V. Discrete Cosine Transform: Algorithms, Advantages, Applications. New York: Academic, 2007.
  114. 114.Mahoney, M and Drineas, P. CUR matrix decompositions for improved data analysis. Proc Natl Acad Sci USA 2009; 106: 697–702.
  115. 115.Owen, J and Rabinovitch, R. On the class of elliptical distributions and their applications to the theory of portfolio choice. J Finance 1983; 38: 745–52.
  116. 116.Blanchard, G, Kawanabe, M and Sugiyama, M et al. In search of non-Gaussian components of a high-dimensional distribution. J Mach Learn Res 2006; 7: 247–82.
  117. 117.Han, F and Liu, H. Scale-Invariant Sparse PCA on High Dimensional Meta-elliptical Data. J Am Stat Assoc, doi:10.1080/01621459.2013.844699.
  118. 118.Candes, E, Li, X and Ma, Y et al. Robust principal component analysis? J. ACM 2011; 58: 11: 1–37.
  119. 119.Loh, P-L and Wainwright, M. High-dimensional regression with noisy and missing data: provable guarantees with nonconvexity. Ann Stat 2012; 40: 1637–64.
  120. 120.Lam, C and Yao, Q. Factor modeling for high-dimensional time series: inference for the number of factors. Ann Stat 2012; 40: 694–726.
  121. 121.Han, F and Liu, H. Principal component analysis on non-Gaussian dependent data. In: The 30th International Conference on Machine Learning, Atlanda, GA, USA, 16–21 June, 2013.
  122. 122.Huang, J, Sun, T and Ying, Z et al. Oracle inequalities for the lasso in the Cox model. Ann Stat 2013; 41: 1142–65.

Citation

MLA
Fan, J., et al. “Challenges of Big Data Analysis”. National Science Review, vol. 1, no. 2, 2014, pp. 293–314, https://doi.org/10.1093/nsr/nwt032.
APA
Fan, J., Han, F., & Liu, H. (2014). Challenges of Big Data analysis. National Science Review, 1(2), 293–314. https://doi.org/10.1093/nsr/nwt032
Chicago
Fan, J., F. Han, and H. Liu. 2014. “Challenges of Big Data Analysis”. National Science Review 1 (2): 293–314. https://doi.org/10.1093/nsr/nwt032.
Harvard
Fan, J., Han, F. and Liu, H. (2014) “Challenges of Big Data analysis”, National Science Review, 1(2), pp. 293–314. Available at: https://doi.org/10.1093/nsr/nwt032.
Vancouver
1. Fan J, Han F, Liu H (2014) Challenges of Big Data analysis. National Science Review 1:293–314

BibTeX

@article{Fan_2014, title={Challenges of Big Data analysis}, volume={1}, ISSN={2095-5138}, url={http://dx.doi.org/10.1093/nsr/nwt032}, DOI={10.1093/nsr/nwt032}, number={2}, journal={National Science Review}, publisher={Oxford University Press (OUP)}, author={Fan, Jianqing and Han, Fang and Liu, Han}, year={2014}, month=Feb, pages={293–314} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF