SoGuDiff: Socially Guided Diffusion for
Steerable, Norm-Grounded Robot Navigation

Christian Schaible1, Haoran Ji2, Yash Vardhan Pant1, Stephen L. Smith1
1University of Waterloo 2McMaster University
Preprint · 2026
A robot approaches its goal past three pedestrians in an indoor scene, with four differently styled sets of candidate, selected and projected trajectories fanning out from it.

Abstract

Socially competent robot navigation requires more than collision avoidance; it demands adherence to implicit social conventions that vary across contexts, cultures, and deployment requirements. Approaches typically learn a single normative behavior, either through reinforcement learning against a fixed reward function or imitation of human demonstrations, exposing no interface for adjusting that conduct at runtime. We present a diffusion-based navigation framework whose social behavior can be tuned at deployment: a desired style is specified, such as how closely the robot passes, which side it yields to, or how much it defers to groups, and the planner adapts accordingly. Continuous style axes can be followed independently or composed, spanning a behavioral space that discrete, primitive-based specifications cannot express. A feasibility projection layer separates learned social behavior from kinematic feasibility and collision avoidance. A single-axis sweep illustrates a tradeoff curve that strictly dominates fixed-behavior baselines, and stylistic differences are replicated in real-world demonstrations.

How It Works

Ten steps, from the inputs the planner receives to the actions it commands.

Overview of SoGuDiff. Left: training, where scene and style tokens condition a 1D U-Net that learns to denoise trajectories. Right: inference, where the network is queried under decomposed conditions, combined by per-axis guidance, selected from, and projected onto feasible controls.
Training Deployment

The Style Vector

Conduct is requested as a four-dimensional vector s = [sprox, spass, syield, sgroup], each axis continuous in [-1, +1]. Every axis is grounded in a documented navigation convention, and each is steered by its own guidance weight at inference.

Proxemic conservatism

s_prox

How much personal space the robot preserves. Zone radii interpolate between Hall's intimate, personal, and social distances.

+1 wide berth, generous clearance
−1 tight, efficient lines

Passing side

s_pass

Which side the robot takes when passing or overtaking. The sign selects the convention, and the magnitude sets how firmly it is enforced. The axis therefore reverses cleanly between left- and right-hand traffic regions.

+1 right-hand convention
−1 left-hand convention

Yielding disposition

s_yield

Whether the robot anticipates and defers in oncoming encounters, and whether it is willing to cut across a pedestrian's path. Built on time-to-collision terms following the power-law structure of human avoidance.

+1 anticipates and gives way
−1 holds its line, cuts in

Group deference

s_group

Whether the robot routes around co-moving social groups or drives between their members. Lowering the axis narrows what counts as a group at all, while raising it admits wider, more loosely spaced formations and enforces deference to them more strictly.

+1 defers to even loose groups
−1 agnostic to groups

Neutral Style and Baseline Comparisons

Five scenes from a randomized evaluation set, where SoGuDiff is run at the neutral style for comparison with ten baseline methods.

SoGuDiff (neutral style)
ORCA
SFM
CADRL
LSTM-RL
SARL
RGL
DSRNN
NaviSTAR
HEIGHT
SICNav
  • reached the goal
  • collided
  • timed out

Steering One Axis at a Time

Each axis is swept from −1 to +1 with the other three held neutral, on a scenario chosen to isolate that interaction.

s_prox = −1 Assertive
Moves along the most assertive, shortest path, cutting close to the pedestrians.
s_prox = 0 Neutral
Steers clear of the pedestrians while balancing path efficiency.
s_prox = +1 Conservative
Takes the widest arc around the pedestrians to preserve clearance.

Composing Axes at Inference

Distinct, interpretable behaviors are composed at inference using per-axis guidance.

Neutral

[0, 0, 0, 0]

The reference behavior with every axis at its default convention.

Assertive & Left-Side Passing

[−1, −1, −1, −1]

Produces the most assertive path and passes all pedestrians on the left.

Cautious & Right-Side Passing

[+1, +1, +1, +1]

Makes the widest arc around all pedestrians, passing on the right side.

Yielding & Group-Agnostic

[0, 0, +1, −1]

Yields to the first pedestrian before cutting through the oncoming social group.

Real-World Deployment

SoGuDiff is deployed on a Clearpath Jackal for real-world demonstrations. Each clip shows the recorded demonstration, human detection via the onboard camera, and logged real-time detection and prediction data.

Unstructured navigation using the neutral style

The robot navigates socially in an environment with three pedestrians.

Style differences across a single axis

For each axis, the same encounter is repeated using −1 and +1 values.

s_prox = −1
Cuts close past the pedestrian, taking the shorter line.
s_prox = +1
Swings wide and early, holding a generous berth all the way past.

Two axes composed

Proxemic conservatism and passing side, four distinct compositions. On the same encounter, differences in behavior are perceptible across each style.

prox +1 pass −1
prox +1 pass +1
prox −1 pass −1
prox −1 pass +1

Adapting social style during runtime

The style is flipped from left-side passing to right-side passing during the run, changing behavior in real-time.

BibTeX

@misc{schaible2026sogudiffsociallyguideddiffusion,
      title={SoGuDiff: Socially Guided Diffusion for Steerable, Norm-Grounded Robot Navigation},
      author={Christian Schaible and Haoran Ji and Yash Vardhan Pant and Stephen L. Smith},
      year={2026},
      eprint={2609.30560},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.30560},
}