Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks

Agrim GuptaJustin JohnsonLi Fei-FeiSilvio SavareseAlexandre Alahi

article2018CVPR2,489 citations

Proposes a generative adversarial framework that combines a novel pooling mechanism with a variety loss to predict diverse, socially compliant pedestrian trajectories with high computational efficiency.

Listen

Autonomous platforms such as self-driving cars and social robots must predict how pedestrians will move through crowded spaces. Human motion depends on nearby people, follows unspoken social rules, and admits many equally valid futures given the same past observations. Existing methods either ignore distant interactions, produce only average paths, or run too slowly for real-time use.

The article evaluates a generative adversarial network called Social GAN that addresses these issues. An encoder-decoder recurrent generator produces candidate trajectories while a recurrent discriminator scores their social plausibility. A new global pooling module shares information across all people in a scene, and a variety loss encourages the generator to output diverse yet realistic futures. The model was trained and tested on the ETH and UCY pedestrian datasets using a leave-one-out protocol, with trajectories observed for 3.2 seconds and predicted for 3.2 or 4.8 seconds.

Social GAN produces lower average and final displacement errors than prior recurrent baselines, with the largest gains on longer horizons. It generates multiple distinct, collision-free paths that respect social conventions such as yielding or group cohesion. The pooling module improves qualitative realism while the variety loss raises accuracy by roughly one-third when many samples are drawn. Computation is sixteen times faster than the leading prior method because pooling occurs only once rather than at every time step.

These results indicate that generative models can supply the multiple plausible futures required by downstream planners without sacrificing speed or safety margins. Real-time deployment on vehicles or robots becomes more feasible, and the learned social norms reduce the risk of awkward or dangerous maneuvers.

Further validation on larger and more diverse scenes, including integration with vehicle dynamics and perception noise, would strengthen before widespread adoption. The main limitations are reliance on perfect position data at the pooling step and training without synthetic augmentation, which may limit generalization to rare events.

  • Paper: Social LSTM: Human Trajectory Prediction in Crowded Spaces, Alexandre Alahi et al. (2016). Social GAN directly extends Social LSTM’s recurrent pedestrian predictor and social-pooling mechanism with adversarial training, global interaction pooling, and multimodal trajectory generation.
  • Paper: Generative Adversarial Networks, Ian J. Goodfellow et al. (2014). This foundational GAN paper supplies the generator–discriminator framework that Social GAN adapts to socially plausible trajectory prediction.
  • Paper: Improved Techniques for Training GANs, Tim Salimans et al. (2016). Its practical GAN-stabilization techniques provide important context for understanding how Social GAN can train a recurrent generator and discriminator effectively.
  • Paper: Improved Training of Wasserstein GANs, Ishaan Gulrajani et al. (2017). WGAN-GP’s treatment of adversarial-training instability clarifies the broader optimization problems that motivate and contextualize Social GAN’s GAN design.
  • Paper: GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium, Martin Heusel et al. (2017). The two-time-scale analysis explains how separately paced generator and discriminator updates can support stable adversarial learning, a useful lens for Social GAN’s training procedure.
Cover for Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks

Abstract

Understanding human motion behavior is critical for autonomous moving platforms (like self-driving cars and social robots) if they are to navigate human-centric environments. This is challenging because human motion is inherently multimodal: given a history of human motion paths, there are many socially plausible ways that people could move in the future. We tackle this problem by combining tools from sequence prediction and generative adversarial networks: a recurrent sequence-to-sequence model observes motion histories and predicts future behavior, using a novel pooling mechanism to aggregate information across people. We predict socially plausible futures by training adversarially against a recurrent discriminator, and encourage diverse predictions with a novel variety loss. Through experiments on several datasets we demonstrate that our approach outperforms prior work in terms of accuracy, variety, collision avoidance, and computational complexity.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Problem Definition
  • 3.2 Generative Adversarial Networks
  • 3.3 Socially-Aware GAN
  • 3.4 Pooling Module
  • 3.5 Encouraging Diverse Sample Generation
  • 3.6 Implementation Details
  • 4 Experiments
  • 4.1 Quantitative Evaluation
  • 4.2 Qualitative Evaluation
  • 4.2.1 Pooling Vs No-Pooling
  • 4.2.2 Pooling in Action
  • 4.3 Structure in Latent Space
  • 5 Conclusion
  • 6 Acknowledgment
  • References

Knowls

  1. Knowl 1 — Social GAN (SGAN) Framework for Multi-Agent Trajectory Prediction

    model/method

    Social GAN (SGAN) is a conditional generative adversarial framework designed to forecast socially plausible, multimodal future trajectories for multiple interacting agents in crowded environments.

    The system consists of three core elements:

    1. Generator (GG): A recurrent sequence-to-sequence network that encodes past observed trajectory coordinates for each pedestrian, aggregates social interactions among all pedestrians using a pooling module, and samples future trajectory coordinates conditioned on a latent noise vector z∼N(0,I)z \sim \mathcal{N}(0, \mathbf{I}).
    2. Pooling Module (extPM ext{PM}): A compact module that collects information across all agent encoders by computing relative spatial displacements and applying a shared multi-layer perceptron followed by a symmetric elementwise max-pooling operation.
    3. Discriminator (DD): A recurrent encoder that receives the entire trajectory—combining the observed past and either the ground-truth future (Treal=[Xi,Yi]T_{\text{real}} = [X_i, Y_i]) or the predicted future (Tfake=[Xi,Y^i]T_{\text{fake}} = [X_i, \hat{Y}_i])—and outputs a classification score determining whether the trajectory is socially acceptable and realistic.
  2. Knowl 2 — Variety Loss for Multimodal Trajectory Generation

    equation

    To prevent mode collapse and avoid generating trajectories that merely reflect the average of multiple plausible paths, the generative network is trained using a variety loss. For a given observed pedestrian trajectory, the generator draws kk random latent vectors z(1),…,z(k)z^{(1)}, \dots, z^{(k)} from a standard normal distribution N(0,I)\mathcal{N}(0, \mathbf{I}) to produce kk candidate future paths Y^i(1),…,Y^i(k)\hat{Y}_i^{(1)}, \dots, \hat{Y}_i^{(k)}, and computes the loss strictly with respect to the candidate closest in L2L_2 distance to the ground truth future trajectory YiY_i:

    Lvariety=min⁡k∥Yi−Y^i(k)∥2\mathcal{L}_{\text{variety}} = \min_{k} \| Y_i - \hat{Y}_i^{(k)} \|_2

    where Yi=(xit,yit)t=tobs+1tpredY_i = (x_i^t, y_i^t)_{t=t_{\text{obs}}+1}^{t_{\text{pred}}} represents the ground truth coordinates of agent ii across future prediction time steps, Y^i(k)\hat{Y}_i^{(k)} is the kk-th generated prediction, and kk is a predefined integer hyperparameter.

    By computing gradients only through the best predicted sample among the kk draws, this loss encourages the generative distribution to cover the multi-modal space of valid physical and social paths without penalizing alternative valid candidate paths.

  3. Knowl 3 — Global Relative Social Pooling Module

    model/method

    To capture human-human interactions across arbitrary numbers of agents without relying on fixed spatial grid boundaries, the pooling module computes a global context vector for each agent ii by processing relative coordinate displacements together with hidden states.

    For each person ii at time step tt, the relative position (xjt−xit,yjt−yit)(x_j^t - x_i^t, y_j^t - y_i^t) with respect to every other person jj in the scene is concatenated with person jj's hidden state hjth_j^t. This combined vector is processed through an independent Multi-Layer Perceptron (MLP) with parameter weights shared across all pairs. An elementwise max-pooling operator then aggregates the transformed vectors across all agents jj in the scene:

    Pi=max⁡j≠iMLP([hjt, (xjt−xit,yjt−yit)])P_i = \max_{j \neq i} \text{MLP}\left( [ h_j^t, \, (x_j^t - x_i^t, y_j^t - y_i^t) ] \right)

    where PiP_i is the pooled social context vector for pedestrian ii, and max⁡\max denotes the symmetric elementwise maximum operation over all other agents jj. This formulation makes the pooling operation invariant to the ordering and count of agents in the scene, avoids spatial grid discretization artifacts, and captures interactions across the entire scene rather than restricting context to an immediate local neighborhood.

  4. Knowl 4 — Recurrent Generator Architecture and Direct Coordinate Decoding

    equation

    The generator encodes past positions and decodes predicted future positions through recurrent neural units using direct coordinate regression.

    For each person ii at time step t≤tobst \le t_{\text{obs}}, spatial coordinates (xit,yit)(x_i^t, y_i^t) are embedded into a fixed-length vector eite_i^t via an MLP embedding function ϕ(⋅;Wee)\phi(\cdot; W_{ee}) with ReLU activations. The encoder LSTM updates its hidden state via:

    eit=ϕ(xit,yit;Wee)e_i^t = \phi(x_i^t, y_i^t; W_{ee})

    heit=LSTM(heit−1,eit;Wencoder)h_{ei}^t = \text{LSTM}(h_{ei}^{t-1}, e_i^t; W_{\text{encoder}})

    After observation time tobst_{\text{obs}}, the pooled interaction context vector PiP_i and the final encoder hidden state heitobsh_{ei}^{t_{\text{obs}}} are embedded via an MLP γ(⋅;Wc)\gamma(\cdot; W_c) and concatenated with a latent noise vector z∼N(0,I)z \sim \mathcal{N}(0, \mathbf{I}) to initialize the decoder hidden state hditobsh_{di}^{t_{\text{obs}}}:

    citobs=γ(Pi,heitobs;Wc)c_i^{t_{\text{obs}}} = \gamma(P_i, h_{ei}^{t_{\text{obs}}}; W_c)

    hditobs=[citobs,z]h_{di}^{t_{\text{obs}}} = [c_i^{t_{\text{obs}}}, z]

    For prediction time steps t=tobs+1,…,tpredt = t_{\text{obs}}+1, \dots, t_{\text{pred}}, coordinates (x^it,y^it)(\hat{x}_i^t, \hat{y}_i^t) are regressed directly via an MLP γ\gamma from the decoder hidden states rather than predicting bivariate Gaussian distribution parameters:

    eit=ϕ(x^it−1,y^it−1;Wed)e_i^t = \phi(\hat{x}_i^{t-1}, \hat{y}_i^{t-1}; W_{ed})

    hdit=LSTM(γ(Pi,hdit−1),eit;Wdecoder)h_{di}^t = \text{LSTM}(\gamma(P_i, h_{di}^{t-1}), e_i^t; W_{\text{decoder}})

    (x^it,y^it)=γ(hdit)(\hat{x}_i^t, \hat{y}_i^t) = \gamma(h_{di}^t)

  5. Knowl 5 — Recurrent Discriminator for Social Etiquette Verification

    model/method

    The discriminator evaluates whether a joint trajectory sequence conforms to natural human social interaction norms.

    It consists of an LSTM encoder that processes the concatenation of the observed past trajectory Xi=(xit,yit)t=1tobsX_i = (x_i^t, y_i^t)_{t=1}^{t_{\text{obs}}} and the corresponding future trajectory over t=tobs+1,…,tpredt = t_{\text{obs}}+1, \dots, t_{\text{pred}}, taking either the real ground truth sequence Treal=[Xi,Yi]T_{\text{real}} = [X_i, Y_i] or the generated candidate sequence Tfake=[Xi,Y^i]T_{\text{fake}} = [X_i, \hat{Y}_i].

    The final hidden state of the discriminator encoder is passed into a Multi-Layer Perceptron (MLP) to produce a scalar real-versus-fake classification score. By discriminating over complete trajectories across all agents, the discriminator learns to penalize trajectories that violate social norms—such as navigating too close to others, failing to yield right-of-way, or exhibiting unnatural sudden path deviations—without requiring hand-crafted energy potentials.

  6. Knowl 6 — Trajectory Forecasting Benchmark Protocol and Metrics

    experimental setup

    Evaluation is conducted on two public pedestrian datasets, ETH and UCY, which collectively contain 5 distinct subsets across 4 scenes comprising 1,536 pedestrian trajectories:

    • ETH: Subsets ETH and HOTEL.
    • UCY: Subsets UNIV, ZARA1, and ZARA2.

    All coordinates are converted to world coordinates and sampled at 0.4-second intervals (2.5 Hz). Models observe 8 time steps (3.2 seconds) and predict trajectories over two horizons: 8 time steps (3.2 seconds) and 12 time steps (4.8 seconds). A leave-one-out cross-validation scheme is employed, where training occurs on 4 subsets and testing on the remaining held-out subset.

    Performance is measured via two standard error metrics (reported in meters, lower is better):

    1. Average Displacement Error (ADE): The average L2L_2 distance between ground truth trajectory coordinates and predicted coordinates over all predicted time steps:

    ADE=1tpred−tobs∑t=tobs+1tpred∥(xit,yit)−(x^it,y^it)∥2\text{ADE} = \frac{1}{t_{\text{pred}} - t_{\text{obs}}} \sum_{t=t_{\text{obs}}+1}^{t_{\text{pred}}} \| (x_i^t, y_i^t) - (\hat{x}_i^t, \hat{y}_i^t) \|_2

    1. Final Displacement Error (FDE): The Euclidean distance between the predicted destination and the true destination at the final predicted step tpredt_{\text{pred}}:

    FDE=∥(xitpred,yitpred)−(x^itpred,y^itpred)∥2\text{FDE} = \| (x_i^{t_{\text{pred}}}, y_i^{t_{\text{pred}}}) - (\hat{x}_i^{t_{\text{pred}}}, \hat{y}_i^{t_{\text{pred}}}) \|_2

    For stochastic models, NN candidate trajectories are sampled at test time, and the candidate with the lowest L2L_2 error relative to ground truth is selected for quantitative evaluation.

  7. Knowl 7 — Pedestrian Trajectory Prediction Accuracy on ETH and UCY Benchmarks

    data/table

    The table below compares Average Displacement Error (ADE) and Final Displacement Error (FDE) in meters across five dataset splits for prediction horizons tpred=8t_{\text{pred}} = 8 and tpred=12t_{\text{pred}} = 12 time steps (8/128 / 12). In model identifiers, kVk\text{V} indicates training with variety loss using kk samples, P\text{P} indicates inclusion of the pooling module, and NN indicates drawing NN test samples.

    Metric Dataset Linear LSTM S-LSTM SGAN-1V-1 SGAN-20V-20 SGAN-20VP-20
    ADE ETH 0.84 / 1.33 0.70 / 1.09 0.73 / 1.09 0.79 / 1.13 0.61 / 0.81 0.60 / 0.87
    HOTEL 0.35 / 0.39 0.55 / 0.86 0.49 / 0.79 0.71 / 1.01 0.48 / 0.72 0.52 / 0.67
    UNIV 0.56 / 0.82 0.36 / 0.61 0.41 / 0.67 0.37 / 0.60 0.36 / 0.60 0.44 / 0.76
    ZARA1 0.41 / 0.62 0.25 / 0.41 0.27 / 0.47 0.25 / 0.42 0.21 / 0.34 0.22 / 0.35
    ZARA2 0.53 / 0.77 0.31 / 0.52 0.33 / 0.56 0.32 / 0.52 0.27 / 0.42 0.29 / 0.42
    AVG 0.54 / 0.79 0.43 / 0.70 0.45 / 0.72 0.49 / 0.74 0.39 / 0.58 0.41 / 0.61
    FDE ETH 1.60 / 2.94 1.45 / 2.41 1.48 / 2.35 1.61 / 2.21 1.22 / 1.52 1.19 / 1.62
    HOTEL 0.60 / 0.72 1.17 / 1.91 1.01 / 1.76 1.44 / 2.18 0.95 / 1.61 1.02 / 1.37
    UNIV 1.01 / 1.59 0.77 / 1.31 0.84 / 1.40 0.75 / 1.28 0.75 / 1.26 0.84 / 1.52
    ZARA1 0.74 / 1.21 0.53 / 0.88 0.56 / 1.00 0.53 / 0.91 0.42 / 0.69 0.43 / 0.68
    ZARA2 0.95 / 1.48 0.65 / 1.11 0.70 / 1.17 0.66 / 1.11 0.54 / 0.84 0.58 / 0.84
    AVG 0.98 / 1.59 0.91 / 1.52 0.91 / 1.54 1.00 / 1.54 0.78 / 1.18 0.81 / 1.21

    SGAN models trained with variety loss (k=20k=20) and evaluated by sampling N=20N=20 predictions achieve the lowest displacement errors overall, lowering average ADE from 0.70m (LSTM) and 0.72m (S-LSTM) down to 0.58m at tpred=12t_{\text{pred}}=12, and average FDE from 1.52m down to 1.18m.

  8. Knowl 8 — Inference Latency and Computational Speedup of Social GAN

    data/table

    Inference execution times benchmarked on an NVIDIA Tesla P100 GPU demonstrate significant efficiency gains for SGAN over Social LSTM (S-LSTM):

    Prediction Steps (tpredt_{\text{pred}}) LSTM S-LSTM SGAN SGAN-P
    8 steps (3.2s) 0.02 s 1.79 s 0.04 s 0.12 s
    12 steps (4.8s) 0.03 s 2.61 s 0.05 s 0.15 s
    Speed-Up relative to S-LSTM 82×\times 1×\times 49×\times 16×\times

    SGAN with pooling (SGAN-P) is 16×16\times faster than S-LSTM (0.15 s vs 2.61 s for 12 prediction steps), allowing the model to generate 16 candidate trajectory samples in the time S-LSTM takes to generate a single prediction. This acceleration stems from: (1) passing the pooled context once during decoder initialization rather than pooling at every decoding recurrent step, and (2) utilizing an MLP with global max-pooling instead of computing occupancy grids for each pedestrian.

  9. Knowl 9 — Performance Impact of Variety Loss versus Test-Time Multi-Sampling

    empirical result

    Comparing models trained without variety loss (SGAN-1V-NN, trained with k=1k=1 and evaluated with NN test samples) against models trained directly with variety loss (SGAN-NNV-NN, trained with k=Nk=N and evaluated with NN test samples) demonstrates that test-time sampling alone is insufficient for multimodal trajectory coverage.

    On the ETH, ZARA1, and ZARA2 datasets, increasing test samples NN up to 100 for SGAN-1V-NN yields minimal displacement error reduction because GANs trained with standard loss suffer from mode collapse and output near-identical trajectories regardless of latent noise zz. In contrast, SGAN-NNV-NN exhibits consistent accuracy gains as NN increases, achieving an average error reduction of approximately 33% at N=100N=100 relative to SGAN-1V-100 (e.g., ADE improves by 34% on ETH, 29% on ZARA1, and 33% on ZARA2).

  10. Knowl 10 — Disentangled Latent Space Encoding of Trajectory Speed and Heading

    empirical result

    Exploration of the generator's latent noise vector space z∼N(0,I)z \sim \mathcal{N}(0, \mathbf{I}) reveals that distinct linear directions along the learned latent manifold correspond to interpretable physical trajectory attributes:

    1. Heading / Direction: Traversing specific axes in the latent space shifts predicted paths smoothly between leftward and rightward deviations relative to the observed heading.
    2. Velocity / Speed: Traversing orthogonal axes modulates the predicted speed along the trajectory, producing faster (longer displacement per time step) or slower (shorter displacement per time step) forward motions.

    Varying the latent vector zz along these directions while keeping the observed past trajectory fixed allows the network to generate globally consistent avoidance maneuvers—such as accelerating to overtake, slowing down to yield right-of-way, or veering to avoid colliding with other pedestrians.

Coverage note — None was omitted; all primary methodological contributions, equations, algorithmic workflows, quantitative benchmark tables, speed comparisons, and empirical findings on diversity and latent space analysis are covered.

References

  1. 1.A. Alahi, K. Goel, V. Ramanathan, A. Robicquet, L. Fei-Fei, and S. Savarese. Social lstm: Human trajectory prediction in crowded spaces. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 961–971, 2016.
  2. 2.G. Antonini, M. Bierlaire, and M. Weber. Discrete choice models of pedestrian walking behavior. Transportation Research Part B: Methodological, 40(8):667–687, 2006.
  3. 3.L. Ballan, F. Castaldo, A. Alahi, F. Palmieri, and S. Savarese. Knowledge transfer for scene-specific motion prediction. In European Conference on Computer Vision, pages 697–713. Springer, 2016.
  4. 4.F. Bartoli, G. Lisanti, L. Ballan, and A. Del Bimbo. Context-aware trajectory prediction. arXiv preprint arXiv:1705.02503, 2017.
  5. 5.W. Choi and S. Savarese. A unified framework for multitarget tracking and collective activity recognition. In Computer Vision–ECCV 2012, pages 215–230. Springer, 2012.
  6. 6.W. Choi and S. Savarese. Understanding collective activitiesof people from videos. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 36(6):1242–1257, 2014.
  7. 7.J. Chorowski, D. Bahdanau, K. Cho, and Y. Bengio. End-to-end continuous speech recognition using attention-based recurrent nn: First results. arXiv preprint arXiv:1412.1602, 2014.
  8. 8.J. Chung, K. Kastner, L. Dinh, K. Goel, A. C. Courville, and Y. Bengio. A recurrent latent variable model for sequential data. In Advances in neural information processing systems, pages 2980–2988, 2015.
  9. 9.P. Coscia, F. Castaldo, F. A. Palmieri, L. Ballan, A. Alahi, and S. Savarese. Point-based path prediction from polar histograms. In Information Fusion (FUSION), 2016 19th International Conference on, pages 1961–1967. IEEE, 2016.
  10. 10.Y. Du, W. Wang, and L. Wang. Hierarchical recurrent neural network for skeleton based action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1110–1118, 2015.
  11. 11.H. Fan, H. Su, and L. Guibas. A point set generation network for 3d object reconstruction from a single image. arXiv preprint arXiv:1612.00603, 2016.
  12. 12.T. Fernando, S. Denman, S. Sridharan, and C. Fookes. Soft+ hardwired attention: An lstm framework for human trajectory prediction and abnormal event detection. arXiv preprint arXiv:1702.05552, 2017.
  13. 13.J. Gauthier. Conditional generative adversarial nets for convolutional face generation. Class Project for Stanford CS231N: Convolutional Neural Networks for Visual Recognition, Winter semester, 2014(5):2, 2014.
  14. 14.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In Advances in neural information processing systems, pages 2672–2680, 2014.
  15. 15.A. Graves and N. Jaitly. Towards end-to-end speech recognition with recurrent neural networks. In Proceedings of the 31st International Conference on Machine Learning (ICML-14), pages 1764–1772, 2014.
  16. 16.K. Gregor, I. Danihelka, A. Graves, D. J. Rezende, and D. Wierstra. Draw: A recurrent neural network for image generation. arXiv preprint arXiv:1502.04623, 2015.
  17. 17.D. Helbing and P. Molnar. Social force model for pedestrian dynamics. Physical review E, 51(5):4282, 1995.
  18. 18.W. Hu, D. Xie, Z. Fu, W. Zeng, and S. Maybank. Semantic-based surveillance video retrieval. Image Processing, IEEE Transactions on, 16(4):1168–1181, 2007.
  19. 19.P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros. Image-to-image translation with conditional adversarial networks. arXiv preprint arXiv:1611.07004, 2016.
  20. 20.A. Karpathy, A. Joulin, and F. F. F. Li. Deep fragment embeddings for bidirectional image sentence mapping. In Advances in neural information processing systems, pages 1889–1897, 2014.
  21. 21.K. Kim, D. Lee, and I. Essa. Gaussian process regression flow for analysis of motion trajectories. In Computer Vision (ICCV), 2011 IEEE International Conference on, pages 1164–1171. IEEE, 2011.
  22. 22.D. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  23. 23.D. P. Kingma and M. Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  24. 24.K. M. Kitani, B. D. Ziebart, J. A. Bagnell, and M. Hebert. Activity forecasting. In Computer Vision–ECCV 2012, pages 201–214. Springer, 2012.
  25. 25.L. Leal-Taixe, M. Fenzi, A. Kuznetsova, B. Rosenhahn, and S. Savarese. Learning an image-based motion context for multiple people tracking. In CVPR, pages 3542–3549. IEEE, 2014.
  26. 26.L. Leal-Taixe, G. Pons-Moll, and B. Rosenhahn. Everybody needs somebody: Modeling social and grouping behavior on a linear programming multiple people tracker. In ICCV Workshops, 2011.
  27. 27.C. Ledig, L. Theis, F. Huszar, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang, et al. Photo-realistic single image super-resolution using a generative adversarial network. arXiv preprint arXiv:1609.04802, 2016.
  28. 28.N. Lee, W. Choi, P. Vernaza, C. B. Choy, P. H. Torr, and M. Chandraker. Desire: Distant future prediction in dynamic scenes with interacting agents. arXiv preprint arXiv:1704.04394, 2017.
  29. 29.J. Liu, A. Shahroudy, D. Xu, and G. Wang. Spatio-temporal lstm with trust gates for 3d human action recognition. In European Conference on Computer Vision, pages 816–833. Springer, 2016.
  30. 30.M. Luber, J. A. Stork, G. D. Tipaldi, and K. O. Arras. People tracking with human motion predictions from social forces. In Robotics and Automation (ICRA), 2010 IEEE International Conference on, pages 464–469. IEEE, 2010.
  31. 31.R. Mehran, A. Oyama, and M. Shah. Abnormal crowd behavior detection using social force model. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on, pages 935–942. IEEE, 2009.
  32. 32.M. Mirza and S. Osindero. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784, 2014.
  33. 33.B. T. Morris and M. M. Trivedi. Trajectory learning for activity understanding: Unsupervised, multilevel, and long-term adaptive approach. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 33(11):2287–2301, 2011.
  34. 34.A. Odena, C. Olah, and J. Shlens. Conditional image synthesis with auxiliary classifier gans. arXiv preprint arXiv:1610.09585, 2016.
  35. 35.H. S. Park and J. Shi. Social saliency prediction.
  36. 36.S. Pellegrini, A. Ess, and L. Van Gool. Improving data association by joint modeling of pedestrian trajectories and groupings. In Computer Vision–ECCV 2010, pages 452–465. Springer, 2010.
  37. 37.C. R. Qi, H. Su, K. Mo, and L. J. Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. arXiv preprint arXiv:1612.00593, 2016.
  38. 38.A. Robicquet, A. Sadeghian, A. Alahi, and S. Savarese. Learning social etiquette: Human trajectory understanding in crowded scenes. In European conference on computer vision, pages 549–565. Springer, 2016.
  39. 39.T. Shu, S. Todorovic, and S.-C. Zhu. Cern: confidenceenergy recurrent network for group activity recognition. Proc. of CVPR, Honolulu, Hawaii, 2017.
  40. 40.N. Srivastava, E. Mansimov, and R. Salakhudinov. Unsupervised learning of video representations using lstms. In International Conference on Machine Learning, pages 843–852, 2015.
  41. 41.M. K. C. Tay and C. Laugier. Modelling smooth paths using gaussian processes. In Field and Service Robotics, pages 381–390. Springer, 2008.
  42. 42.A. Treuille, S. Cooper, and Z. Popovic. Continuum crowds. In ACM Transactions on Graphics (TOG), volume 25, pages 1160–1168. ACM, 2006.
  43. 43.O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. Show and tell: A neural image caption generator. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3156–3164, 2015.
  44. 44.J. M. Wang, D. J. Fleet, and A. Hertzmann. Gaussian process dynamical models for human motion. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 30(2):283–298, 2008.
  45. 45.K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. In International Conference on Machine Learning, pages 2048–2057, 2015.
  46. 46.K. Yamaguchi, A. C. Berg, L. E. Ortiz, and T. L. Berg. Who are you with and where are you going? In Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on, pages 1345–1352. IEEE, 2011.
  47. 47.S. Yi, H. Li, and X. Wang. Understanding pedestrian behaviors from stationary crowd groups. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3488–3496, 2015.
  48. 48.H. Zhang, T. Xu, H. Li, S. Zhang, X. Huang, X. Wang, and D. Metaxas. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. arXiv preprint arXiv:1612.03242, 2016.
  49. 49.B. Zhou, X. Wang, and X. Tang. Random field topic model for semantic region analysis in crowded scenes from tracklets. In Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on, pages 3441–3448. IEEE, 2011.

Citation

MLA
Gupta, A., et al. “Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks”. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 2255–64, https://doi.org/10.1109/CVPR.2018.00240.
APA
Gupta, A., Johnson, J., Fei-Fei, L., Savarese, S., & Alahi, A. (2018). Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2255–2264. https://doi.org/10.1109/CVPR.2018.00240
Chicago
Gupta, A., J. Johnson, L. Fei-Fei, S. Savarese, and A. Alahi. 2018. “Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks”. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2255–64. https://doi.org/10.1109/CVPR.2018.00240.
Harvard
Gupta, A. et al. (2018) “Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, pp. 2255–2264. Available at: https://doi.org/10.1109/CVPR.2018.00240.
Vancouver
1. Gupta A, Johnson J, Fei-Fei L, Savarese S, Alahi A (2018) Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, pp 2255–2264

BibTeX

@inproceedings{Gupta_2018, title={Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks}, url={http://dx.doi.org/10.1109/CVPR.2018.00240}, DOI={10.1109/cvpr.2018.00240}, booktitle={2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition}, publisher={IEEE}, author={Gupta, Agrim and Johnson, Justin and Fei-Fei, Li and Savarese, Silvio and Alahi, Alexandre}, year={2018}, month=June, pages={2255–2264} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE