Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning

NeurIPS 2026

1 School of Information and Control Engineering, China University of Mining and Technology

2 School of Computer Science and Engineering, South China University of Technology

Selected OGBench evaluation environments: maze navigation, ball control, cube manipulation, scene interaction, and puzzle solving.

Abstract

Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoal decomposition; however, high-level decision-making still relies on noise-sensitive value estimates, leading to unstable behavior in complex environments. We address this limitation by proposing Diffusion Subgoal Planning (DSP), a diffusion-based framework for high-level subgoal generation. DSP casts high-level planning as guided generative inference over goal-conditioned subgoals and learns both conditional and unconditional flows, enabling classifier-free guidance to introduce a goal-directed bias at inference time. By removing explicit value-based guidance from high-level planning, DSP generates reachable and goal-directed subgoals through a generative model while retaining hierarchical execution. Experiments on offline GCRL benchmarks demonstrate that DSP outperforms prior methods on a range of navigation and manipulation tasks, with particularly strong performance in maze environments that require multi-step subgoal planning.

Method

DSP learns to generate subgoals from offline trajectories using flow matching. At inference time, classifier-free guidance steers generation toward the desired goal, while a low-level policy executes each subgoal. High-level planning uses no explicit value-function guidance; the low-level learning retains a value function.

Subgoal Planning

AntMaze Giant

HIQLw/o

HIQL without subgoal representation in AntMaze Giant, with the paper’s red box marking a cluster of subgoals along the plan.

DSP Ours

DSP: a numbered subgoal sequence extends through the AntMaze Giant maze toward the marked goal, overlaid on a value heatmap for visualization.
The AntMaze Giant task 1 comparison used in the paper’s main text. Numbered markers show subgoals; the red box marks a local cluster in the HIQLw/o plan. The background shows the goal-conditioned value function for visualization; DSP does not query it for high-level guidance.

Experiments

OGBench

We evaluate DSP on 18 locomotion and manipulation tasks from OGBench, comparing it with flat and hierarchical offline goal-conditioned RL methods.

Scroll horizontally to compare all methods.

Main results: goal-reaching success rate (%)
TaskFlat policiesHierarchical policies
GCBCCFGRLGCIVLOTAPi-HIQLHIQLHIQLw/oDSP
pointmaze-medium4.8 ± 6.965.2 ± 5.370.2 ± 5.985.4 ± 5.062.2 ± 7.370.6 ± 7.074.6 ± 6.190.6 ± 3.8
pointmaze-large25.6 ± 6.374.6 ± 2.942.6 ± 5.387.6 ± 9.280.2 ± 13.339.4 ± 2.350.0 ± 10.493.6 ± 2.9
pointmaze-giant2.2 ± 4.93.4 ± 2.51.0 ± 2.269.4 ± 11.63.0 ± 2.95.0 ± 5.00.0 ± 0.043.4 ± 7.2
pointmaze-teleport23.8 ± 5.749.2 ± 5.143.6 ± 2.540.2 ± 5.433.6 ± 9.08.8 ± 6.48.2 ± 6.251.0 ± 5.8
antmaze-medium32.6 ± 8.534.8 ± 5.574.6 ± 8.091.6 ± 2.490.8 ± 2.793.2 ± 1.392.8 ± 2.498.4 ± 0.5
antmaze-large23.0 ± 2.018.4 ± 5.715.0 ± 6.689.2 ± 2.381.8 ± 3.187.8 ± 1.588.6 ± 2.192.0 ± 4.1
antmaze-giant0.0 ± 0.00.0 ± 0.00.0 ± 0.068.4 ± 3.450.4 ± 3.758.2 ± 4.750.4 ± 8.069.2 ± 2.9
antmaze-teleport26.4 ± 2.235.2 ± 7.535.6 ± 5.748.4 ± 2.649.4 ± 2.743.4 ± 6.640.8 ± 4.159.8 ± 1.3
humanoidmaze-medium6.8 ± 1.37.0 ± 2.332.6 ± 4.082.4 ± 3.676.2 ± 3.874.8 ± 3.052.6 ± 4.389.6 ± 2.9
humanoidmaze-large1.0 ± 1.41.2 ± 1.62.6 ± 1.876.4 ± 4.039.0 ± 3.223.6 ± 2.97.2 ± 3.469.4 ± 3.9
humanoidmaze-giant0.0 ± 0.00.4 ± 0.50.0 ± 0.081.4 ± 4.934.2 ± 3.35.0 ± 1.60.0 ± 0.066.0 ± 3.4
antsoccer-arena5.4 ± 3.012.2 ± 4.147.4 ± 7.539.8 ± 6.517.8 ± 2.960.6 ± 4.556.6 ± 5.576.6 ± 4.0
antsoccer-medium1.4 ± 0.53.6 ± 2.111.6 ± 2.316.6 ± 3.52.4 ± 1.58.4 ± 1.57.2 ± 1.515.2 ± 2.4
cube-single5.2 ± 2.27.6 ± 2.353.4 ± 5.410.0 ± 4.21.2 ± 1.111.8 ± 2.321.8 ± 4.041.0 ± 6.1
cube-double1.0 ± 1.01.6 ± 1.131.6 ± 3.02.2 ± 0.80.0 ± 0.03.6 ± 2.29.4 ± 4.430.8 ± 8.4
scene4.8 ± 2.617.2 ± 3.344.6 ± 2.923.4 ± 7.116.6 ± 3.137.2 ± 4.242.0 ± 5.262.6 ± 5.9
puzzle-3x33.0 ± 1.22.0 ± 1.24.2 ± 2.54.6 ± 3.25.6 ± 1.58.8 ± 1.316.0 ± 2.913.4 ± 2.3
puzzle-4x40.0 ± 0.00.0 ± 0.07.8 ± 2.63.4 ± 1.813.6 ± 8.47.8 ± 2.65.4 ± 2.327.6 ± 2.9
Total167.0333.6518.4920.4658.0648.0623.61090.2

Results average success over five test-time goals and five seeds; ± denotes standard deviation. Bold means reach at least 95% of the highest mean in each task, following the paper’s highlighting rule. This threshold is not a statistical significance test. Total sums the 18 task means. Task names omit -navigate-v0 for navigation and -play-v0 for manipulation. HIQLw/o denotes HIQL without subgoal representation.

DSP has the highest mean success rate on 11 of the 18 tasks in this table. Performance varies by task: OTA leads on several giant-maze tasks, and GCIVL leads on the two cube tasks.

Learning curves

Full-size figure
Success rate over training steps on PointMaze Giant, AntMaze Giant, HumanoidMaze Giant, and AntSoccer Arena, comparing DSP, HIQL, and HIQL without subgoal representation.
Success rate over training steps on four selected OGBench navigation tasks, comparing DSP with HIQL and HIQLw/o.

Citation

@article{zhang2026dsp,
  title={Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning},
  author={Zhang, Hengrui and Cheng, Yuhu and Chen, C. L. Philip and Wang, Xuesong},
  journal={arXiv preprint arXiv:2609.34575},
  year={2026}
}