Learning Patterns of Activity Using Real-Time Tracking

Chris StaufferW. Eric L. Grimson

article2000TPAMI3,737 citations

Introduces an adaptive Gaussian mixture model for real-time background subtraction and pairs it with a co-occurrence clustering method to automatically classify object silhouettes and scene activities without manual supervision.

Listen

This paper describes a real-time visual monitoring system that tracks moving objects across extended outdoor sites and automatically learns typical activity patterns from those tracks. The work addresses the practical need for surveillance tools that can operate continuously without manual setup, identify normal traffic flows or pedestrian routes, and flag deviations such as unusual volumes, paths, or interactions.

The authors first developed an adaptive background-subtraction tracker that models each pixel as a mixture of Gaussians and updates the model online. Foreground pixels are grouped into connected regions and followed across frames with a multiple-hypothesis Kalman tracker. The resulting sequences of position, velocity, size, and binary silhouettes are then processed by an unsupervised classification pipeline: vector quantization builds a codebook of representative prototypes, joint co-occurrence counts are accumulated across entire tracks, and these statistics are used to construct a hierarchical binary tree that separates the prototypes into increasingly specific activity classes.

Tests on scenes monitored continuously since 1997 showed that the tracker processed 11–13 frames per second, remained stable through lighting shifts, weather, and repetitive clutter, and recorded more than ten million objects. The learned hierarchy cleanly separated opposing traffic directions, road versus path movement, cars from trucks, individual pedestrians from groups, and vehicle silhouettes from human silhouettes, with daily activity histograms matching expected rush-hour and lunchtime patterns. Preliminary anomaly detection compared instantaneous states and sequence co-occurrences against the accumulated model to highlight rare events.

These capabilities allow sites to build statistical descriptions of normal behavior and raise alerts for outliers without predefined rules or labeled training data. Because the same two parameters govern both tracking and classification, the approach transfers across indoor and outdoor cameras with little retuning. The main limitations are reduced performance when objects frequently overlap for long periods and the requirement that observed sequences connect related activities; isolated roads or very sparse data can leave some classes disconnected in the hierarchy. Further work on context cycles, richer local features, and combined prototype-plus-co-occurrence outlier scoring would strengthen anomaly detection before large-scale deployment.

  • Paper: SAM 2: Segment Anything in Images and Videos, Nikhila Ravi et al. (2025). This contemporary segment-anything framework extends the early focus on real-time visual tracking and activity recognition into modern promptable video segmentation.
Cover for Learning Patterns of Activity Using Real-Time Tracking

Abstract

Abstract—Our goal is to develop a visual monitoring system that passively observes moving objects in a site and learns patterns of activity from those observations. For extended sites, the system will require multiple cameras. Thus, key elements of the system are motion tracking, camera coordination, activity classification, and event detection. In this paper, we focus on motion tracking and show how one can use observed motion to learn patterns of activity in a site. Motion segmentation is based on an adaptive background subtraction method that models each pixel as a mixture of Gaussians and uses an on-line approximation to update the model. The Gaussian distributions are then evaluated to determine which are most likely to result from a background process. This yields a stable, real-time outdoor tracker that reliably deals with lighting changes, repetitive motions from clutter, and long-term scene changes. While a tracking system is unaware of the identity of any object it tracks, the identity remains the same for the entire tracking sequence. Our system leverages this information by accumulating joint co-occurrences of the representations within a sequence. These joint co-occurrence statistics are then used to create a hierarchical binary-tree classification of the representations. This method is useful for classifying sequences, as well as individual instances of activities in a site.

Table of Contents

  • Learning Patterns of Activity Using Real-Time Tracking
  • 1 INTRODUCTION
  • 2 BUILDING A ROBUST MOTION TRACKER
  • 2.1 Previous Work and Current Shortcomings of Motion Tracking
  • 2.2 Our Approach to Motion Tracking
  • 3 ADAPTIVE BACKGROUNDING FOR MOTION TRACKING
  • 3.1 Online Mixture Model
  • 3.2 Background Model Estimation
  • 3.3 Connected Components
  • 3.4 Multiple Hypothesis Tracking
  • 4 PERFORMANCE OF THE TRACKER
  • 5 IMPROVING THE TRACKER
  • 6 INTERPRETING THE MOTION TRACKS
  • 6.1 Previous Work in Classification
  • 7 THE CLASSIFICATION METHOD
  • 7.1 Codebook Generation
  • 7.2 Accumulating Co-Occurrence Statistics
  • 7.3 Hierarchical Classification
  • 7.4 Classifying a Sequence
  • 7.5 A Simple Example
  • 8 RESULTS
  • 8.1 Classifying Activities
  • 8.2 Classifying Motion Silhouettes
  • 9 DETECTING UNUSUAL EVENTS
  • 10 CLASSIFICATION SHORTCOMINGS AND FUTURE WORK
  • 11 CONCLUSIONS
  • ACKNOWLEDGMENTS
  • REFERENCES

Knowls

  1. Knowl 1 — Adaptive Mixture of Gaussians Pixel Model

    model/method

    In video background subtraction, the time series of values observed at a specific pixel location over time is modeled as a mixture of KK Gaussian distributions (with KK typically chosen between 3 and 5). For a color or grayscale measurement Xt∈RnX_t \in \mathbb{R}^n at frame tt, the probability density of observing XtX_t is:

    P(Xt)=∑i=1Kωi,t η(Xt;μi,t,Σi,t)P(X_t) = \sum_{i=1}^K \omega_{i,t} \, \eta(X_t; \mu_{i,t}, \Sigma_{i,t})

    where ωi,t\omega_{i,t} is the estimated weight (prior probability) of the ii-th Gaussian component at time tt satisfying ∑i=1Kωi,t=1\sum_{i=1}^K \omega_{i,t} = 1, μi,t∈Rn\mu_{i,t} \in \mathbb{R}^n is its mean vector, and Σi,t\Sigma_{i,t} is its covariance matrix. The Gaussian density η\eta in nn dimensions is given by:

    η(Xt;μ,Σ)=1(2π)n/2∣Σ∣1/2exp⁡(−12(Xt−μ)TΣ−1(Xt−μ))\eta(X_t; \mu, \Sigma) = \frac{1}{(2\pi)^{n/2} |\Sigma|^{1/2}} \exp\left(-\frac{1}{2}(X_t - \mu)^T \Sigma^{-1} (X_t - \mu)\right)

    To eliminate the computational burden of inverting full covariance matrices in real time, the color channels (e.g., red, green, blue) are assumed to be conditionally independent and to share equal variance, reducing the covariance matrix to an isotropic form:

    Σk,t=σk,t2I\Sigma_{k,t} = \sigma_{k,t}^2 I

    where σk,t2\sigma_{k,t}^2 is a scalar variance and II is the n×nn \times n identity matrix.

  2. Knowl 2 — Online Update Procedure for Pixel Mixture of Gaussians

    algorithm

    To update the mixture of Gaussians model for each pixel in real time without batch Expectation-Maximization, an online KK-means-like approximation is used at every frame tt.

    Input: Current pixel value XtX_t, learning rate α\alpha, current mixture parameters {ωk,t−1,μk,t−1,σk,t−12\omega_{k,t-1}, \mu_{k,t-1}, \sigma_{k,t-1}^2}k=1K_{k=1}^K
    Output: Updated mixture parameters {ωk,t,μk,t,σk,t2\omega_{k,t}, \mu_{k,t}, \sigma_{k,t}^2}k=1K_{k=1}^K, and match status
    match_found = false
    for k=1k = 1 to KK do
        if ∥Xt−μk,t−1∥≤2.5σk,t−1\|X_t - \mu_{k,t-1}\| \le 2.5 \sigma_{k,t-1} then
            match_found = true
            matched_index = kk
            break
        end if
    end for
    if match_found then
        for k=1k = 1 to KK do
            if k==matchedindexk == matched_index then
                Mk,t=1M_{k,t} = 1
                ρ=α η(Xt;μk,t−1,σk,t−12I)\rho = \alpha \, \eta(X_t; \mu_{k,t-1}, \sigma_{k,t-1}^2 I)
                μk,t=(1−ρ)μk,t−1+ρXt\mu_{k,t} = (1 - \rho) \mu_{k,t-1} + \rho X_t
                σk,t2=(1−ρ)σk,t−12+ρ(Xt−μk,t)T(Xt−μk,t)\sigma_{k,t}^2 = (1 - \rho) \sigma_{k,t-1}^2 + \rho (X_t - \mu_{k,t})^T (X_t - \mu_{k,t})
            else
                Mk,t=0M_{k,t} = 0
                μk,t=μk,t−1\mu_{k,t} = \mu_{k,t-1}
                σk,t2=σk,t−12\sigma_{k,t}^2 = \sigma_{k,t-1}^2
            end if
            ωk,t=(1−α)ωk,t−1+αMk,t\omega_{k,t} = (1 - \alpha) \omega_{k,t-1} + \alpha M_{k,t}
        end for
    else
        find component rr with the lowest prior weight ωr,t−1\omega_{r,t-1}
        μr,t=Xt\mu_{r,t} = X_t
        σr,t2=σinit2\sigma_{r,t}^2 = \sigma_{\text{init}}^2 // high initial variance
        ωr,t=ωinit\omega_{r,t} = \omega_{\text{init}} // low initial prior weight
        for k=1k = 1 to KK and k≠rk \neq r do
            ωk,t=(1−α)ωk,t−1\omega_{k,t} = (1 - \alpha) \omega_{k,t-1}
            μk,t=μk,t−1\mu_{k,t} = \mu_{k,t-1}
            σk,t2=σk,t−12\sigma_{k,t}^2 = \sigma_{k,t-1}^2
        end for
    end if
    normalize weights such that ∑k=1Kωk,t=1\sum_{k=1}^K \omega_{k,t} = 1
    return {ωk,t,μk,t,σk,t2\omega_{k,t}, \mu_{k,t}, \sigma_{k,t}^2}k=1K_{k=1}^K, match_found

    The parameter α\alpha is the global learning rate defining the adaptation time constant 1/α1/\alpha. Unmatched components keep their means and variances intact while their weights decay, preserving previous background distributions if a foreground object becomes stationary and subsequently departs.

  3. Knowl 3 — Background Model Identification via Fitness-to-Variance Ranking

    model/method

    In an adaptive Gaussian mixture background model, persistent static background surfaces exhibit high evidence (large mixture weight ωk\omega_k) and low variation (small standard deviation σk\sigma_k), whereas transient moving objects generate low weights or larger variances. To identify which Gaussian components belong to the background process at time tt:

    1. The KK Gaussian components of a pixel are sorted in descending order of the ratio ωk/σk\omega_k / \sigma_k.
    2. The first BB distributions in this sorted list are designated as the background model, where BB is the minimum number of distributions needed to account for a predefined fraction TT of the observed background data:

    B=arg⁡min⁡b(∑k=1bωk>T)B = \arg\min_b \left( \sum_{k=1}^b \omega_k > T \right)

    Here, T∈(0,1]T \in (0, 1] represents the minimum expected portion of recent observations produced by the background.

    If the current pixel value XtX_t matches any of the first BB distributions (i.e., falls within 2.52.5 standard deviations of its mean), the pixel is classified as background; otherwise, it is labeled as foreground. Higher values of TT allow the background model to become multimodal, supporting background elements with repetitive motions such as swaying foliage, flashing lights, or water ripples.

  4. Knowl 4 — Co-Occurrence Matrix Estimation from Tracked Feature Multisets

    algorithm

    To learn patterns of activity without requiring rigid temporal sequence alignment, tracked sequences of feature vectors are converted into discrete codebook prototype labels and aggregated into a pairwise co-occurrence matrix.

    Given a codebook of KcodeK_{\text{code}} prototypes obtained through online vector quantization:

    Input: Set of ZZ tracked sequences {S(1),S(2),…,S(Z)}\{S^{(1)}, S^{(2)}, \dots, S^{(Z)}\}, where each sequence S=(S1,S2,…,S∣S∣)S = (S_1, S_2, \dots, S_{|S|}) contains prototype labels Sm∈{1,…,Kcode}S_m \in \{1, \dots, K_{\text{code}}\}
    Output: Normalized co-occurrence matrix C∈RKcode×KcodeC \in \mathbb{R}^{K_{\text{code}} \times K_{\text{code}}}
    initialize Ci,jtotal=0C^{\text{total}}_{i,j} = 0 for all i,j∈{1,…,Kcode}i, j \in \{1, \dots, K_{\text{code}}\}
    for each sequence SS in {S(1),…,S(Z)}\{S^{(1)}, \dots, S^{(Z)}\} do
        NS=∣S∣N_S = |S|
        P=NS2−NSP = N_S^2 - N_S // number of valid ordered pairs (Sm,Sn)(S_m, S_n) with m≠nm \neq n
        if P>0P > 0 then
            for m=1m = 1 to NSN_S do
                for n=1n = 1 to NSN_S do
                    if m≠nm \neq n then
                        i=Smi = S_m
                        j=Snj = S_n
                        Ci,jtotal=Ci,jtotal+1PC^{\text{total}}_{i,j} = C^{\text{total}}_{i,j} + \frac{1}{P}
                    end if
                end for
            end for
        end if
    end for
    for i=1i = 1 to KcodeK_{\text{code}} do
        for j=1j = 1 to KcodeK_{\text{code}} do
            Ci,j=Ci,jtotalZC_{i,j} = \frac{C^{\text{total}}_{i,j}}{Z}
        end for
    end for
    return CC

    Treating each sequence as an unordered multiset where all pairwise cross-associations are weighted inversely by P=∣S∣2−∣S∣P = |S|^2 - |S| ensures that long and short tracking sequences contribute equally to the empirical co-occurrence statistics.

  5. Knowl 5 — Hierarchical Binary-Tree Activity Clustering via Co-Occurrence Decomposition

    algorithm

    Given a co-occurrence matrix C∈RK×KC \in \mathbb{R}^{K \times K} over KK prototype states, a hierarchical binary classification tree is constructed by recursively factoring probability mass functions (pmfs) over the prototypes to explain the empirical co-occurrence matrix.

    At each node of the tree, the observed co-occurrence is modeled as a mixture of two sub-classes (c∈{0,1}c \in \{0, 1\}) with class priors π0,π1\pi_0, \pi_1 and class-conditional pmfs p0,p1p_0, p_1: C^i,j=∑c∈{0,1}πc pc(i) pc(j)\hat{C}_{i,j} = \sum_{c \in \{0,1\}} \pi_c \, p_c(i) \, p_c(j)

    The parameters are iteratively optimized to minimize the sum-of-squared errors E=∑i,j(Ci,j−C^i,j)2E = \sum_{i,j} (C_{i,j} - \hat{C}_{i,j})^2:

    Input: Co-occurrence matrix C∈RK×KC \in \mathbb{R}^{K \times K}, learning rates απ\alpha_\pi and αp\alpha_p (with απ>αp\alpha_\pi > \alpha_p), maximum iterations MiterM_{\text{iter}}
    Output: Split parameters (π0,p0,C0)(\pi_0, p_0, C^0) and (π1,p1,C1)(\pi_1, p_1, C^1)
    initialize π0,π1\pi_0, \pi_1 and positive vectors p0,p1p_0, p_1 randomly such that ∑ipc(i)=1\sum_i p_c(i) = 1 and π0+π1=1\pi_0 + \pi_1 = 1
    for iter = 1 to MiterM_{\text{iter}} do
        compute C^i,j=π0p0(i)p0(j)+π1p1(i)p1(j)\hat{C}_{i,j} = \pi_0 p_0(i) p_0(j) + \pi_1 p_1(i) p_1(j) for all i,ji,j
        for c∈{0,1}c \in \{0, 1\} do
            Δπc=∑i,j(Ci,j−C^i,j)pc(i)pc(j)\Delta \pi_c = \sum_{i,j} (C_{i,j} - \hat{C}_{i,j}) p_c(i) p_c(j)
            πc=(1−απ)πc+απΔπc\pi_c = (1 - \alpha_\pi) \pi_c + \alpha_\pi \Delta \pi_c
            
            for i=1i = 1 to KK do
                Δpc(i)=∑j(Ci,j−C^i,j)pc(j)\Delta p_c(i) = \sum_j (C_{i,j} - \hat{C}_{i,j}) p_c(j)
                pc(i)=(1−αp)pc(i)+αpΔpc(i)p_c(i) = (1 - \alpha_p) p_c(i) + \alpha_p \Delta p_c(i)
            end for
            re-normalize pcp_c such that ∑ipc(i)=1\sum_i p_c(i) = 1
        end for
        re-normalize priors such that π0+π1=1\pi_0 + \pi_1 = 1
    end for
    for i=1i = 1 to KK and j=1j = 1 to KK do
        Ci,j0=Ci,j p0(i) p0(j)C^0_{i,j} = C_{i,j} \, p_0(i) \, p_0(j)
        Ci,j1=Ci,j p1(i) p1(j)C^1_{i,j} = C_{i,j} \, p_1(i) \, p_1(j)
    end for
    return (π0,p0,C0),(π1,p1,C1)(\pi_0, p_0, C^0), (\pi_1, p_1, C^1)

    The restricted co-occurrence matrices C0C^0 and C1C^1 are passed to the child branches for recursive subdivision until the similarity between child distributions exceeds a stopping threshold, defining a leaf classifier.

  6. Knowl 6 — Instance and Sequence Classification via Prototype Log-Likelihoods

    model/method

    In the co-occurrence classification framework, each leaf node in the pruned hierarchical binary tree corresponds to an underlying class cc characterized by an assigned prior probability πc\pi_c and a probability mass function pcp_c over the KK codebook prototypes.

    Given an observation sequence S=(X1,X2,…,XM)S = (X_1, X_2, \dots, X_M) of an object (where M≥1M \ge 1, supporting both full trajectories and single-frame instances), each observation XmX_m is quantized to its nearest codebook prototype index s(Xm)∈{1,…,K}s(X_m) \in \{1, \dots, K\}. The sequence is summarized by a prototype frequency histogram h∈NKh \in \mathbb{N}^K, where hkh_k is the count of observations mapped to prototype kk.

    Assuming conditional independence of prototype observations in a sequence given class cc, the posterior class log-likelihood score is computed as:

    log⁡P(S∣c)+log⁡πc=log⁡πc+∑m=1Mlog⁡pc(s(Xm))=log⁡πc+∑k=1Khklog⁡pc(k)\log P(S \mid c) + \log \pi_c = \log \pi_c + \sum_{m=1}^M \log p_c(s(X_m)) = \log \pi_c + \sum_{k=1}^K h_k \log p_c(k)

    Classification assigns the sequence (or isolated instance) to the optimal class c∗c^*:

    c∗=arg⁡max⁡c[log⁡πc+h⋅log⁡pc]c^* = \arg\max_c \left[ \log \pi_c + h \cdot \log p_c \right]

    Because the model operates on discrete prototype counts, single observations (such as an instantaneous bounding box feature vector or a single binary silhouette) can be classified directly without requiring a complete multi-frame trajectory.

  7. Knowl 7 — Multiple Hypothesis Object Tracking with Interaction Breakage

    model/method

    Foreground regions detected via Gaussian mixture background subtraction are grouped into connected components and tracked across consecutive video frames using a linearly predictive multiple hypothesis tracking (MHT) scheme based on Kalman filters that incorporate both spatial position (x,y)(x, y) and bounding box size.

    Tracking and correspondence management operate according to the following rules:

    1. Model Matching: Existing Kalman track models are probabilistically matched against available foreground connected components larger than 1-2 pixels. Matches with low prediction error update the corresponding Kalman filter.
    2. Track Propagation and Pruning: When an active model lacks a matching connected component, a null match is hypothesized: the track state is propagated forward by the Kalman motion model, and its fitness score (the inverse of the prediction error variance) is decayed by a constant factor. If fitness falls below a threshold, the track is deleted; if the object reappears within its predicted uncertainty window before deletion, the track resumes.
    3. Model Initialization: Unmatched connected components across consecutive frames are paired to hypothesize new Kalman track models, which are instantiated into the active track pool if corroborated in subsequent frames.
    4. Interaction Policy: When two or more moving objects visually overlap or interact, the tracking system intentionally breaks the tracks rather than attempting to guess true correspondences across ambiguous occlusions. This guarantees that accumulated sequences remain clean single-object multisets for downstream co-occurrence learning.
  8. Knowl 8 — Performance and Environmental Robustness of the Gaussian Mixture Tracker

    empirical result

    The adaptive Gaussian mixture background subtraction tracker was evaluated on continuous real-world outdoor video streams:

    • Real-Time Frame Rate: On an SGI O2 workstation equipped with an R10000 processor, the tracker achieved continuous processing speeds of 11 to 13 frames per second on 160×120160 \times 120 pixel video streams, with variations depending on the percentage of foreground pixels in the scene.
    • Long-Term Autonomous Operation: The system continuously monitored five distinct outdoor scenes without manual parameter adjustment from October 1997 through 2000, tracking in excess of 10 million moving objects.
    • Environmental Invariance: Stable tracking persisted across weather conditions including rain, snow, sleet, hail, sunny, overcast, and foggy conditions, as well as day/night diurnal illumination cycles and repetitive scene clutter (e.g., swaying branches and water specularities).
    • Recovery from Illumination Transients: Following rapid illumination shifts (e.g., fast cloud transitions or lighting changes relative to the learning rate α\alpha), the background mixture models stabilized and restored normal tracking within 10 to 20 seconds without requiring manual reinitialization.
  9. Knowl 9 — Unsupervised Activity and Silhouette Clustering Results

    empirical result

    The co-occurrence hierarchical clustering method was validated on 24-hour tracking datasets using two distinct input feature spaces:

    1. Kinematic Activity Classification (5-Tuple Representation): Objects were parameterized by position, velocity, and size (x,y,dx,dy,size)(x, y, dx, dy, \text{size}) and quantized into 400 prototypes. The resulting 4-level binary tree automatically partitioned the scene activities:
      • Level 1 Split: Separated traffic moving around the building into eastbound versus westbound directions.
      • Level 2 Split: Separated roadway vehicle motion from pedestrian footpath motion.
      • Leaf Level (Level 4): Formed specific activity clusters, including directional pedestrians on walkways, lawn mowers and pedestrians on grass lawns, loading-dock activities, passenger cars, and delivery trucks. Hourly event distributions showed distinct traffic profiles corresponding to morning rush hour, midday pedestrian peaks, and evening rush hour.
    2. Shape Classification (1,024-Tuple Silhouette Representation): Binary motion silhouettes cropped to 32×3232 \times 32 pixels were quantized into 400 prototypes. The co-occurrence hierarchy produced a clean initial bifurcation separating vehicles from pedestrians (with only blurry, low-resolution silhouettes remaining ambiguous and shared between classes), followed by sub-branches isolating individual pedestrians, groups of pedestrians, and scene clutter (such as wind-blown tree foliage, debris, and dusk lighting effects on building walls).
  10. Knowl 10 — Limitations of Trajectory Co-Occurrence Graph Topology

    limitation

    The unsupervised co-occurrence classification framework exhibits two primary structural limitations:

    1. Disconnected Path Topology: Because similarity between prototype states is derived strictly from empirical co-occurrences within continuous single-object tracks, classes that are topologically disjoint in the scene cannot be linked. For example, if vehicles traverse two parallel, non-intersecting roadways with no trajectories moving between them, the mutual co-occurrence between prototypes on the two paths is zero. Consequently, the hierarchy cannot discover that the objects on both roads belong to the same category without multi-camera correspondence or post-training supervisory labels.
    2. Codebook Discretization Scaling: Partitioning continuous input spaces into discrete prototype codebooks requires codebook sizes (KK) that grow rapidly with feature dimensionality. Because the co-occurrence matrix scales as O(K2)O(K^2), substantial observation volumes are required to sufficiently populate transition frequencies and avoid zero-count sparsity.

Coverage note — Camera self-calibration and ground plane estimation were deliberately omitted as they are cited as external contributions covered in companion work (Lee et al., 2000).

References

  1. 1.D. Beymer, P. McLauchlan, B. Coifman, and J. Malik, ªA Real-Time Computer Vision System for Measuring Traffic Parameters,º Proc. Computer Vision and Pattern Recognition, June 1997.
  2. 2.R. Collins, A. Lipton, and T. Kanade, ªA System for Video Surveillance and Monitoring,º Proc. Am. Nuclear Soc. (ANS) Eighth Int'l Topical Meeting Robotic and Remote Systems, Apr. 1999.
  3. 3.L. Davis et al., ªVisual Surveillance of Human Activity,º Proc. Asian Conf. Computer Vision (ACCV98), Jan. 1998.
  4. 4.A Dempster, N. Laird, and D. Rubin, ªMaximum Likelihood from Incomplete Data via the EM Algorithm,º J. Royal Statistical Soc., vol. 39 (Series B), pp. 1-38, 1977.
  5. 5.N. Friedman and S. Russell, ªImage Segmentation in Video Sequences: A Probabilistic Approach,º Proc. 13th Conf. Uncertainty in Artificial Intelligence (UAI), Aug. 1997.
  6. 6.A. Gersho and R.M. Gray, Vector Quantization and Signal Compression. Kluwer Academic, 1991.
  7. 7.W.E.L. Grimson, C. Stauffer, R. Romano, and L. Lee, ªUsing Adaptive Tracking to Classify and Monitor Activities in a Site,º Computer Vision and Pattern Recognition (CVPR 98), June 1998.
  8. 8.B.K.P. Horn, Robot Vision, pp. 66-69, 299-333. MIT Press, 1986.
  9. 9.I. Haritaoglu, D. Harwood, and L.S. Davis, ªW4: Who? When? Where? What? A Real Time System for Detecting and Tracking People,º Proc. Third Int'l Conf. Automatic Face and Gesture Recognition, Apr. 1998.
  10. 10.I. Horswill and M. Yamamoto, ªA $1000 Active Stereo Vision System,º Proc. IEEE/IAP Workshop Visual Behaviors, Aug. 1994.
  11. 11.I. Horswill, ªVisual Routines and Visual Search: A Real-Time Implementation and Automata-Theoretic Analysis,º Proc. Int'l Joint Conf. AI, 1995.
  12. 12.Y. Ivanov, A. Bobick, and J. Liu, ªFast Lighting Independent Background Subtraction,º Technical Report no. 437, MIT Media Laboratory, 1997.
  13. 13.N. Johnson and D.C. Hogg, ªLearning the Distribution of Object Trajectories for Event Recognition,º Proc. British Machine Vision Conf., D. Pycock, ed., pp. 583-592, Sept. 1995.
  14. 14.N. Johnson and D.C. Hogg, ªLearning the Distribution of Object Trajectories for Event Recognition,º Image and Vision Computing, vol. 14, no. 8, pp. 609-615, Aug. 1996.
  15. 15.T. Kanade, ªA Stereo Machine for Video-Rate Dense Depth Mapping and Its New Applications,º Proc. Image Understanding Workshop, pp. 805-811, Feb. 1995.
  16. 16.D. Koller, J. Weber, T. Huang, J. Malik, G. Ogasawara, B. Rao, and S. Russel, ªTowards Robust Automatic Traffic Scene Analysis in Real-Time,º Proc. Int'l Conf. Pattern Recognition, Nov. 1994.
  17. 17.K. Konolige, ªSmall Vision Systems: Hardware and Implementation,º Proc. Eighth Int'l Symp. Robotics Research, Oct. 1997.
  18. 18.L. Lee, R. Romano, and G. Stein, ªMonitoring Activities from Multiple Video Streams: Establishing a Common Coordinate Frame,º IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 22, no. 8, pp. 758-767, Aug. 2000.
  19. 19.A. Lipton, H. Fujiyoshi, and R.S. Patil, ªMoving Target Classification and Tracking from Real-Time Video,º Proc. IEEE Workshop Applications of Computer Vision (WACV), pp. 8-14, Oct. 1998.
  20. 20.N. Oliver, B. Rosario, and A. Pentland, ªA Bayesian Computer Vision System for Modeling Human Interactions,º Proc. Int'l Conf. Vision Systems '99, Jan. 1999.
  21. 21.C. Ridder, O. Munkelt, and H. Kirchner, ªAdaptive Background Estimation and Foreground Detection Using Kalman-Filtering,º Proc. Int'l Conf. Recent Advances in Mechatronics, ICRAM '95, pp. 193-199, 1995.
  22. 22.J. Shi and J. Malik, ªNormalized Cuts and Image Segmentation,º Proc. IEEE Conf. Computer Vision and Pattern Recognition, June 1997.
  23. 23.C. Stauffer and W.E.L. Grimson, ªAdaptive Background Mixture Models for Real-Time Tracking,º Proc. Computer Vision and Pattern Recognition 1999 (CVPR '99), June 1999.
  24. 24.C.R. Wren, A. Azarbayejani, T. Darrell, and A. Pentland, ªPfinder: Real-Time Tracking of the Human Body,º IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 19, no. 7, pp. 780-785, July 1997.

Citation

MLA
Stauffer, C., and W. E. L. Grimson. “Learning Patterns of Activity Using Real-time Tracking”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 22, no. 8, 2000, pp. 747–57, https://doi.org/10.1109/34.868677.
APA
Stauffer, C., & Grimson, W. E. L. (2000). Learning patterns of activity using real-time tracking. IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(8), 747–757. https://doi.org/10.1109/34.868677
Chicago
Stauffer, C., and W. E. L. Grimson. 2000. “Learning Patterns of Activity Using Real-time Tracking”. IEEE Transactions on Pattern Analysis and Machine Intelligence 22 (8): 747–57. https://doi.org/10.1109/34.868677.
Harvard
Stauffer, C. and Grimson, W.E.L. (2000) “Learning patterns of activity using real-time tracking”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(8), pp. 747–757. Available at: https://doi.org/10.1109/34.868677.
Vancouver
1. Stauffer C, Grimson WEL (2000) Learning patterns of activity using real-time tracking. IEEE Transactions on Pattern Analysis and Machine Intelligence 22:747–757

BibTeX

@article{Stauffer_2000, title={Learning patterns of activity using real-time tracking}, volume={22}, ISSN={0162-8828}, url={http://dx.doi.org/10.1109/34.868677}, DOI={10.1109/34.868677}, number={8}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Stauffer, C. and Grimson, W.E.L.}, year={2000}, pages={747–757} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF