You'll never walk alone: Modeling social behavior for multi-target tracking

S. PellegriniAndreas EssKonrad SchindlerLuc Van Gool

article2009ICCV1,840 citations

Proposes a social interaction motion model, Linear Trajectory Avoidance, that predicts collision-free pedestrian paths to maintain reliable tracking and data association in crowded scenes during prolonged occlusions.

Listen

Multi-target visual tracking in crowded public environments is a critical capability for autonomous vehicles, robotics, and intelligent surveillance. Traditional tracking systems rely on simple linear or independent motion models that predict each individual's path in isolation. These conventional approaches ignore fundamental human behaviors, such as steering toward destinations and proactively adjusting paths to avoid collisions with other people and static obstacles. As a result, standard systems frequently lose track of targets or swap personal identities when individuals pass near one another or become temporarily occluded.

The article introduces and evaluates the Linear Trajectory Avoidance (LTA) dynamic model. The objective is to demonstrate that incorporating social interactions and destination-driven path planning into a unified, metric-space motion model significantly enhances multi-person tracking accuracy and trajectory prediction.

The researchers designed an energy-minimization framework where each pedestrian optimizes their velocity by anticipating the point of closest approach with others while maintaining a desired speed and heading. Six core behavioral parameters were trained using 25 minutes of annotated overhead video footage encompassing 650 pedestrian trajectories. The model was then evaluated across three settings: short-term trajectory prediction on an oblique shopping street video, a simple patch-based tracking experiment at a low frame rate (2.5 frames per second), and a state-of-the-art pedestrian tracker mounted on a moving vehicle using real-world street footage.

The evaluation produced four key findings. First, in trajectory prediction benchmarks, LTA reduced average prediction errors by 24% compared to standard linear extrapolation and by 6% compared to the traditional social force baseline. At a one-meter error tolerance threshold, LTA successfully predicted roughly 70% of trajectories, outperforming linear extrapolation (approximately 50%) and destination-only modeling (approximately 63%). Second, in low-frame-rate tracking, LTA consistently maintained target tracks through close encounters where linear models failed due to overshooting and path drift. Third, in vehicle-mounted mobile tracking, LTA substantially lowered tracking errors during occlusions, reducing recoverable identity switches by 44% (from 18 to 10 switches in one sequence) while keeping false positives and missed detections steady. Fourth, these performance gains were achieved at a negligible computational cost of under 10 milliseconds per frame for 15 individuals.

These findings show that tracking systems do not require complex, brittle data-association heuristics if their underlying motion models reflect basic human behavioral dynamics. Even approximate or coarsely estimated destination vectors substantially improve path prediction during sensor dropouts and visual occlusions. This provides immediate operational benefits for autonomous systems by enhancing safety, lowering identity confusion, and reducing computational overhead without requiring specialized hardware.

Based on these results, engineering teams developing autonomous navigation and surveillance systems should integrate destination-aware, social dynamic models into their tracking pipelines. For mobile observers where long-term destinations are unknown, rough heuristics—such as assuming forward motion parallel to the roadway—should be used as an effective substitute. Future development should incorporate a stochastic formulation to resolve ambiguous collision-avoidance directions, and extend the energy formulations to account for coordinated group behaviors, such as people walking together.

The primary limitation of the current model is its deterministic formulation, which occasionally causes large error spikes when it chooses the wrong side to pass an oncoming person. Additionally, evaluating performance in moderately crowded scenes with mostly straight paths may understate the model's advantages in denser crowds. Despite these boundaries, confidence in the findings is high, supported by consistent validation across multiple trackers, camera viewpoints, and standard performance metrics.

No sufficiently relevant recommendations were found.

Cover for You'll never walk alone: Modeling social behavior for multi-target tracking

Abstract

Object tracking typically relies on a dynamic model to predict the object’s location from its past trajectory. In crowded scenarios a strong dynamic model is particularly important, because more accurate predictions allow for smaller search regions, which greatly simplifies data association. Traditional dynamic models predict the location for each target solely based on its own history, without taking into account the remaining scene objects. Collisions are resolved only when they happen. Such an approach ignores important aspects of human behavior: people are driven by their future destination, take into account their environment, anticipate collisions, and adjust their trajectories at an early stage in order to avoid them. In this work, we introduce a model of dynamic social behavior, inspired by models developed for crowd simulation. The model is trained with videos recorded from birds-eye view at busy locations, and applied as a motion model for multi-people tracking from a vehicle-mounted camera. Experiments on real sequences show that accounting for social interactions and scene knowledge improves tracking performance, especially during occlusions.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Modeling Social Behavior
  • 3.1. Static Obstacles
  • 3.2. Application of the Model
  • 4. Training
  • 5. Results
  • 5.1. Prediction
  • 5.2. Patch-based Tracking
  • 5.3. Tracking with a Moving Observer
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Linear Trajectory Avoidance Total Energy Objective Function

    model/method

    The Linear Trajectory Avoidance (LTA) dynamic model determines the preferred candidate ground-plane velocity v~i∈R2\tilde{\mathbf{v}}_i \in \mathbb{R}^2 for subject sis_i at current 2D ground position pi∈R2\mathbf{p}_i \in \mathbb{R}^2 by minimizing a total objective energy functional Ei(v~i)E_i(\tilde{\mathbf{v}}_i):

    Ei(v~i)=Ii(v~i)+λ1Si(v~i)+λ2Di(v~i)E_i(\tilde{\mathbf{v}}_i) = I_i(\tilde{\mathbf{v}}_i) + \lambda_1 S_i(\tilde{\mathbf{v}}_i) + \lambda_2 D_i(\tilde{\mathbf{v}}_i)

    where:

    • Ii(v~i)I_i(\tilde{\mathbf{v}}_i) is the interaction energy capturing collision avoidance across surrounding subjects.

    • Si(v~i)=(ui−∥v~i∥)2S_i(\tilde{\mathbf{v}}_i) = (u_i - \|\tilde{\mathbf{v}}_i\|)^2 is the speed regularizer penalizing deviation of the candidate speed from the pedestrian's desired speed ui∈R+u_i \in \mathbb{R}^+.

    • Di(v~i)=−(zi−pi)⋅v~i∥zi−pi∥∥v~i∥D_i(\tilde{\mathbf{v}}_i) = -\frac{(\mathbf{z}_i - \mathbf{p}_i) \cdot \tilde{\mathbf{v}}_i}{\|\mathbf{z}_i - \mathbf{p}_i\| \|\tilde{\mathbf{v}}_i\|} is the directional regularizer rewarding velocity vectors aligned with the vector pointing from current position pi\mathbf{p}_i toward destination goal point zi∈R2\mathbf{z}_i \in \mathbb{R}^2.

    • λ1,λ2≥0\lambda_1, \lambda_2 \ge 0 are positive weighting hyperparameters balancing interaction avoidance against speed and direction maintenance.

  2. Knowl 2 — Expected Point of Closest Approach and Distance Formulation

    equation

    Under the assumption that subject sis_i expects subject sjs_j to continue with constant velocity vj∈R2\mathbf{v}_j \in \mathbb{R}^2 from position pj∈R2\mathbf{p}_j \in \mathbb{R}^2, the relative offset is defined as k=pi−pj\mathbf{k} = \mathbf{p}_i - \mathbf{p}_j and the relative velocity under candidate velocity v~i\tilde{\mathbf{v}}_i is defined as q=v~i−vj\mathbf{q} = \tilde{\mathbf{v}}_i - \mathbf{v}_j. The squared distance between sis_i and sjs_j at future time t>0t > 0 is:

    dij2(t,v~i)=∥k+tq∥2d_{ij}^2(t, \tilde{\mathbf{v}}_i) = \|\mathbf{k} + t\mathbf{q}\|^2

    The unconstrained time of closest approach is found at t∗=−k⋅q∥q∥2t^* = -\frac{\mathbf{k} \cdot \mathbf{q}}{\|\mathbf{q}\|^2}. Constraining the collision anticipation strictly to future time steps (t>0t > 0), if t∗≤0t^* \le 0 the minimum distance occurs at t=0t = 0 (yielding minimum distance ∥k∥2\|\mathbf{k}\|^2). When t∗>0t^* > 0, substituting t∗t^* yields the minimum squared distance at closest approach:

    dij∗2(v~i)=∥k−k⋅q∥q∥2q∥2d_{ij}^{*2}(\tilde{\mathbf{v}}_i) = \left\|\mathbf{k} - \frac{\mathbf{k} \cdot \mathbf{q}}{\|\mathbf{q}\|^2}\mathbf{q}\right\|^2

  3. Knowl 3 — Social Interaction Potential with Visual Field Weighting

    equation

    The pairwise interaction potential Eij(v~i)E_{ij}(\tilde{\mathbf{v}}_i) between subject sis_i and subject sjs_j is modeled as a Gaussian over the minimum distance at closest approach dij∗2(v~i)d_{ij}^{*2}(\tilde{\mathbf{v}}_i):

    Eij(v~i)=exp⁡(−dij∗2(v~i)2σd2)E_{ij}(\tilde{\mathbf{v}}_i) = \exp\left(-\frac{d_{ij}^{*2}(\tilde{\mathbf{v}}_i)}{2\sigma_d^2}\right)

    where σd\sigma_d is the standard deviation representing the personal comfort distance threshold. The aggregate interaction energy Ii(v~i)I_i(\tilde{\mathbf{v}}_i) experienced by subject sis_i from all other subjects srs_r (r≠ir \neq i) is computed as:

    Ii(v~i)=∑r≠iwr(i)Eir(v~i)I_i(\tilde{\mathbf{v}}_i) = \sum_{r \neq i} w_r(i) E_{ir}(\tilde{\mathbf{v}}_i)

    with composite weight wr(i)=wrd(i)wrϕ(i)w_r(i) = w_r^d(i) w_r^\phi(i) factoring distance and field of view:

    wrd(i)=exp⁡(−∥pi−pr∥22σw2)w_r^d(i) = \exp\left(-\frac{\|\mathbf{p}_i - \mathbf{p}_r\|^2}{2\sigma_w^2}\right)

    wrϕ(i)={(1+cos⁡(ϕ)2)βif ∣ϕ∣≤π20otherwisew_r^\phi(i) = \begin{cases} \left(\frac{1 + \cos(\phi)}{2}\right)^\beta & \text{if } |\phi| \le \frac{\pi}{2} \\ 0 & \text{otherwise} \end{cases}

    where σw\sigma_w defines the spatial radius of influence of other agents, ϕ\phi is the angular displacement of subject srs_r relative to the heading of sis_i, and β\beta controls the directional peakiness of the pedestrian's field-of-view attention.

  4. Knowl 4 — Inertial Velocity Transition and Ground-Plane Position Update Rule

    model/method

    After computing the optimal desired velocity v~i∗=arg⁡min⁡v~iEi(v~i)\tilde{\mathbf{v}}_i^* = \arg\min_{\tilde{\mathbf{v}}_i} E_i(\tilde{\mathbf{v}}_i) via gradient descent with line search, physical inertia prevents instantaneous velocity changes. The position of subject sis_i at prediction step tNt_N is updated using a geometric velocity mixing rule:

    pitN=pi+(αNvi+(1−αN)v~i∗)tN\mathbf{p}_i^{t_N} = \mathbf{p}_i + \left( \alpha^N \mathbf{v}_i + (1 - \alpha^N)\tilde{\mathbf{v}}_i^* \right) t_N

    where pi\mathbf{p}_i and vi\mathbf{v}_i are the subject's ground position and velocity at the current time step, NN denotes the discrete prediction interval scaling the equation to varying video framerates, and α∈[0,1]\alpha \in [0, 1] is the velocity mixture coefficient controlling physical inertia.

  5. Knowl 5 — Learned Parameter Set for Linear Trajectory Avoidance

    data/table

    The six free parameters of the Linear Trajectory Avoidance (LTA) motion model were optimized on overhead birds-eye video data (650 pedestrian trajectories recorded over 25 minutes) using a genetic algorithm minimizing the sum of squared distance errors to ground truth over 4.8-second simulation windows initialized every 1.2 seconds:

    σd\sigma_d σw\sigma_w λ1\lambda_1 λ2\lambda_2 β\beta α\alpha
    0.361 2.088 2.33 2.073 1.462 0.730

    These optimal parameters indicate that:

    • σd=0.361\sigma_d = 0.361 m sets comfortable personal space tolerance to approximately 1 m (3σd3\sigma_d).

    • σw=2.088\sigma_w = 2.088 m limits significant interaction influence to neighbors within approximately 6 m (3σw3\sigma_w).

    • β=1.462\beta = 1.462 focuses path planning attention predominantly in the frontal field of view.

    • α=0.730\alpha = 0.730 provides momentum smoothing across time frames.

  6. Knowl 6 — Static Obstacle Integration in Linear Trajectory Avoidance

    model/method

    Static obstacles in the scene are incorporated into the Linear Trajectory Avoidance (LTA) motion framework by treating them as virtual dynamic subjects with zero velocity (v=0\mathbf{v} = \mathbf{0}). At each time step, the obstacle's virtual position relative to pedestrian sis_i is approximated by the single point on the obstacle boundary closest to pi\mathbf{p}_i. For moving-observer stereo setups, static obstacles are extracted by projecting stereo depth maps onto a polar occupancy grid map on the ground plane, enabling dynamic subjects to steer around non-pedestrian barriers using the same energy formulation.

  7. Knowl 7 — Parallel Hypothesis Extension in Multi-Person Tracking

    model/method

    In visual tracking-by-detection from a moving camera, pedestrian detections are projected into 3D world coordinates using visual odometry and ground-plane constraints. The Linear Trajectory Avoidance (LTA) model is incorporated into trajectory hypothesis generation by replacing independent linear trajectory extrapolation with a parallel hypothesis extension step:

    1. At each frame, all currently active trajectory hypotheses are extended concurrently by evaluating the LTA energy functional for each target against all other co-existing hypotheses and static occupancy map obstacles.

    2. During occlusion events where visual detections are unavailable, LTA forecasts the occluded target's path based on surrounding pedestrian trajectories and destination direction rather than linear extrapolation.

    3. Hypotheses compete for incoming visual detections in data association, maintaining track identity through occlusions and reducing identity switches when targets emerge.

  8. Knowl 8 — Trajectory Prediction Accuracy of LTA Against Baselines

    empirical result

    On an oblique-view shopping street benchmark (86 trajectories, evaluated across ~300 simulations of 4.8 seconds each at 2.5 FPS), the LTA model was evaluated against three alternative motion models:

    • Linear extrapolation (LIN)
    • Social Force model with elliptical potentials (SF)
    • Destination-only model without social interactions (DEST)

    LTA achieves a 6% reduction in average Euclidean prediction error compared to both SF and DEST, and a 24% reduction compared to LIN. Evaluated as the percentage of trajectories whose prediction error remains within a spatial threshold T=1.0T = 1.0 meter throughout the simulation window, LIN achieves ≈50%\approx 50\%, DEST achieves ≈63%\approx 63\%, SF performs slightly above DEST, and LTA reaches ≈70%\approx 70\% accuracy.

  9. Knowl 9 — Tracking Performance Under Varying Association Thresholds (CLEAR MOT)

    data/table

    Multi-target tracking performance of the Linear Trajectory Avoidance (LTA) model compared against a linear constant-velocity model (LIN) integrated into a mobile tracking-by-detection system across different Mahalanobis distance thresholds dd on two video sequences:

    ID switches Misses False positives
    Threshold dd 1.5 2.0 2.5 3.0 1.5 2.0 2.5 3.0 1.5 2.0 2.5 3.0
    Seq#1
    LIN 55 55 51 48 0.29 0.28 0.28 0.28 0.19 0.19 0.19 0.19
    LTA 48 42 45 41 0.28 0.28 0.28 0.28 0.19 0.19 0.19 0.19
    Seq#2
    LIN 35 33 31 31 0.21 0.21 0.21 0.21 0.08 0.08 0.09 0.09
    LTA 31 30 26 25 0.21 0.21 0.20 0.20 0.08 0.09 0.08 0.09

    LTA consistently reduces the number of identity switches across all association thresholds while maintaining comparable miss and false positive rates. When manually accounting for recoverable tracks at d=3.0d = 3.0 by excluding targets that permanently leave and re-enter the camera field of view, ID switches in Seq#2 decrease from 18 (LIN) to 10 (LTA), representing a 44% reduction.

  10. Knowl 10 — Limitations of Deterministic Avoidance and Group Dynamics

    limitation

    The Linear Trajectory Avoidance (LTA) formulation has two primary limitations:

    1. Deterministic Local Optimization Artifacts: When oncoming pedestrians encounter each other head-on, the energy landscape presents symmetric local minima (passing on the left vs. right). Because LTA is deterministic, choosing the opposite side from the ground truth path results in large Euclidean distance penalties in trajectory evaluation, generating a long tail of large errors.

    2. Absence of Group and Non-Convex Modeling: The energy formulation does not model social groupings (multiple people walking together who maintain proximity rather than repelling each other). In addition, static obstacles are approximated by a single closest point, which degrades for highly non-convex physical obstacles.

Coverage note — Standard underlying detection pipelines (HOG detector details) and standard visual odometry/homography baseline formulas were omitted as they represent existing prior work rather than novel contributions of this paper.

References

  1. 1.S. Ali and M. Shah. Floor fields for tracking in high density crowd scenes. In ECCV, 2008.
  2. 2.M. Andriluka, S. Roth, and B. Schiele. People-tracking-by-detection and people-detection-by-tracking. In CVPR’08.
  3. 3.G. Antonini, S. V. Martinez, M. Bierlaire, and J. Thiran. Behavioral priors for detection and tracking of pedestrians in video sequences. IJCV, 69:159–180, 2006.
  4. 4.K. Bernardin and R. Stiefelhagen. Evaluating multiple object tracking performance: The CLEAR MOT metrics. EURASIP Journal on Image and Video Processing, 2008.
  5. 5.M. D. Breitenstein, F. Reichlin, B. Leibe, E. Koller-Meier, and L. V. Gool. Robust tracking-by-detection using a detector confidence particle filter. In ICCV, 2009.
  6. 6.G. Brostow and R. Cipolla. Unsupervised bayesian detection of independent motion in crowds. In CVPR, 2006.
  7. 7.N. Dalal and B. Triggs. Histograms of oriented gradients for human detection. In CVPR, 2005.
  8. 8.A. Ess, B. Leibe, K. Schindler, and L. van Gool. A mobile vision system for robust multi-person tracking. In CVPR’08.
  9. 9.P. Felzenszwalb, D. McAllester, and D. Ramanan. A discriminatively trained, multiscale, deformable part model. In CVPR, 2008.
  10. 10.H. Grabner and H. Bischof. On-line boosting and vision. In CVPR, 2006.
  11. 11.D. Helbing and P. Molnár. Social force model for pedestrian dynamics. Physical Review E, 51(5):4282–4286, 1995.
  12. 12.C. Huang, B. Wu, and R. Nevatia. Robust object tracking by hierarchical association of detection responses. In ECCV’08.
  13. 13.A. Johansson, D. Helbing, and P. K. Shukla. Specification of a microscopic pedestrian model by evolutionary adjustment to video tracking data. Advances in Complex Systems, 10(2):271–288, 2007.
  14. 14.R. Kaucic, A. G. Perera, G. Brooksby, J. Kaufhold, and A. Hoogs. A unified framework for tracking through occlusions and across sensor gaps. In CVPR, 2005.
  15. 15.F. Klügl and G. Rindsfüser. Large-scale agent-based pedestrian simulation. In MATES ’07: Proc. of the 5th German Conference on Multiagent Systems Technology, 2007.
  16. 16.B. Leibe, K. Schindler, N. Cornelis, and L. Van Gool. Coupled detection and tracking from static cameras and moving vehicles. IEEE TPAMI, 30(10):1683–1698, 2008.
  17. 17.A. Lerner, Y. Chrysanthou, and D. Lischinski. Crowds by example. In EUROGRAPHICS, 2007.
  18. 18.R. Mehran, A. Oyama, and M. Shah. Abnormal crowd behavior detection using social force model. In CVPR, 2009.
  19. 19.K. Okuma, A. Taleghani, N. de Freitas, J. Little, and D. Lowe. A boosted particle filter: Multitarget detection and tracking. In ECCV, 2004.
  20. 20.A. Penn and A. Turner. Space syntax based agent simulation. In PED, 2002.
  21. 21.A. Schadschneider. Cellular automaton approach to pedestrian dynamics—theory. In PED. 2001.
  22. 22.B. Wu and R. Nevatia. Detection and tracking of multiple, partially occluded humans by bayesian combination of edgelet part detectors. IJCV, 75(2):247–266, 2007.
  23. 23.L. Zhang, Y. Li, and R. Nevatia. Global data association for multi-object tracking using network flows. In CVPR, 2008.
  24. 24.T. Zhao and R. Nevatia. Tracking multiple humans in crowded environment. In CVPR, 2004.

Citation

MLA
Pellegrini, S., et al. “You'll Never Walk Alone: Modeling Social Behavior for Multi-target Tracking”. 2009 IEEE 12th International Conference on Computer Vision, 2009, pp. 261–68, https://doi.org/10.1109/ICCV.2009.5459260.
APA
Pellegrini, S., Ess, A., Schindler, K., & van Gool, L. (2009). You'll never walk alone: Modeling social behavior for multi-target tracking. 2009 IEEE 12th International Conference on Computer Vision, 261–268. https://doi.org/10.1109/ICCV.2009.5459260
Chicago
Pellegrini, S., A. Ess, K. Schindler, and L. van Gool. 2009. “You'll Never Walk Alone: Modeling Social Behavior for Multi-target Tracking”. 2009 IEEE 12th International Conference on Computer Vision, 261–68. https://doi.org/10.1109/ICCV.2009.5459260.
Harvard
Pellegrini, S. et al. (2009) “You'll never walk alone: Modeling social behavior for multi-target tracking”, 2009 IEEE 12th International Conference on Computer Vision. IEEE, pp. 261–268. Available at: https://doi.org/10.1109/ICCV.2009.5459260.
Vancouver
1. Pellegrini S, Ess A, Schindler K, van Gool L (2009) You'll never walk alone: Modeling social behavior for multi-target tracking. In: 2009 IEEE 12th International Conference on Computer Vision. IEEE, pp 261–268

BibTeX

@inproceedings{Pellegrini_2009, title={You’ll never walk alone: Modeling social behavior for multi-target tracking}, url={http://dx.doi.org/10.1109/ICCV.2009.5459260}, DOI={10.1109/iccv.2009.5459260}, booktitle={2009 IEEE 12th International Conference on Computer Vision}, publisher={IEEE}, author={Pellegrini, S and Ess, A and Schindler, K and van Gool, L}, year={2009}, month=Sept, pages={261–268} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE