Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks
Agrim GuptaJustin JohnsonLi Fei-FeiSilvio SavareseAlexandre Alahi
Proposes a generative adversarial framework that combines a novel pooling mechanism with a variety loss to predict diverse, socially compliant pedestrian trajectories with high computational efficiency.
Autonomous platforms such as self-driving cars and social robots must predict how pedestrians will move through crowded spaces. Human motion depends on nearby people, follows unspoken social rules, and admits many equally valid futures given the same past observations. Existing methods either ignore distant interactions, produce only average paths, or run too slowly for real-time use.
The article evaluates a generative adversarial network called Social GAN that addresses these issues. An encoder-decoder recurrent generator produces candidate trajectories while a recurrent discriminator scores their social plausibility. A new global pooling module shares information across all people in a scene, and a variety loss encourages the generator to output diverse yet realistic futures. The model was trained and tested on the ETH and UCY pedestrian datasets using a leave-one-out protocol, with trajectories observed for 3.2 seconds and predicted for 3.2 or 4.8 seconds.
Social GAN produces lower average and final displacement errors than prior recurrent baselines, with the largest gains on longer horizons. It generates multiple distinct, collision-free paths that respect social conventions such as yielding or group cohesion. The pooling module improves qualitative realism while the variety loss raises accuracy by roughly one-third when many samples are drawn. Computation is sixteen times faster than the leading prior method because pooling occurs only once rather than at every time step.
These results indicate that generative models can supply the multiple plausible futures required by downstream planners without sacrificing speed or safety margins. Real-time deployment on vehicles or robots becomes more feasible, and the learned social norms reduce the risk of awkward or dangerous maneuvers.
Further validation on larger and more diverse scenes, including integration with vehicle dynamics and perception noise, would strengthen before widespread adoption. The main limitations are reliance on perfect position data at the pooling step and training without synthetic augmentation, which may limit generalization to rare events.
- Paper: Social LSTM: Human Trajectory Prediction in Crowded Spaces, Alexandre Alahi et al. (2016). Social GAN directly extends Social LSTM’s recurrent pedestrian predictor and social-pooling mechanism with adversarial training, global interaction pooling, and multimodal trajectory generation.
- Paper: Generative Adversarial Networks, Ian J. Goodfellow et al. (2014). This foundational GAN paper supplies the generator–discriminator framework that Social GAN adapts to socially plausible trajectory prediction.
- Paper: Improved Techniques for Training GANs, Tim Salimans et al. (2016). Its practical GAN-stabilization techniques provide important context for understanding how Social GAN can train a recurrent generator and discriminator effectively.
- Paper: Improved Training of Wasserstein GANs, Ishaan Gulrajani et al. (2017). WGAN-GP’s treatment of adversarial-training instability clarifies the broader optimization problems that motivate and contextualize Social GAN’s GAN design.
- Paper: GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium, Martin Heusel et al. (2017). The two-time-scale analysis explains how separately paced generator and discriminator updates can support stable adversarial learning, a useful lens for Social GAN’s training procedure.
- Paper: Generative Agents: Interactive Simulacra of Human Behavior, Joon Sung Park et al. (2023). Generative Agents extends Social GAN’s learned social behavior from short-horizon trajectory prediction to persistent, interactive simulations of human groups.
- Paper: Learning Latent Dynamics for Planning from Pixels, Danijar Hafner et al. (2018). PlaNet carries Social GAN’s multimodal future-prediction idea into latent world modeling, using imagined dynamics to support downstream planning and control.
- Paper: Dream to Control: Learning Behaviors by Latent Imagination, Danijar Hafner et al. (2019). Dreamer continues predictive modeling toward action by learning policies through latent imagination rather than using generated futures only as trajectory forecasts.
- Paper: Robust Reinforcement Learning via Genetic Curriculum, Yeeho Song et al. (2022). Genetic Curriculum applies learned scenario generation to robust control, extending Social GAN’s concern with rare, difficult futures into systematic safety-oriented training.
- Paper: WorldSimBench: Towards Video Generation Models as World Simulators, Yiran Qin et al. (2025). WorldSimBench generalizes predictive video models toward actionable world simulation, providing evaluation of whether generated futures support embodied decision-making.
