Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to the entire generation, are poorly suited to correcting these inconsistencies over long horizons. We introduce LoGo, which blends global and spatially localized rewards for camera-controlled video models. The local reward gives fine-grained credit assignment, which substantially improves 3D consistency, while the global reward preserves camera following and video quality. Across three base models, LoGo shows a clear advantage on DL3DV and TrajectoryBench, a new benchmark for long-horizon, complex camera trajectories that current evaluations lack. LoGo effectively reduces local object shifts and artifacts, as well as global scene changes, illustrating the importance of credit assignment in post-training video models.
As video generation models advance toward longer horizon, various 3D inconsistency failure modes emerge, such as 1) scene layout changes upon revisit, 2) artifacts or corruption, 3) objects appearing or disappearing, and 4) objects changing geometry or appearance. While prior methods cannot effectively fix them, LoGo resolves these inconsistencies.
Input image
Base
VideoGPA
World-R1
LoGo
WASD denotes camera translation, and arrows denote rotation.
The most naive method to post-train for 3D consistency is to use a global reward. However, this is not effective enough (see ablation below). Modern video models fail in largely local ways, including inconsistent object appearance, hallucinated or disappearing objects, and local artifacts. A good reward should penalize these failures directly rather than merely ranking full-length videos. Motivated by this, we design a reward localization method for accurate credit assignment during post-training. LoGo introduces three components: 1) reward localization, 2) Local-Global blending, 3) reward interleaving.
Reward Localization. LoGo is based on 3D scene reprojection. Given a video, we first construct a scene point cloud by passing keyframes to VGGT-Ω and unproject each pixel using the predicted depth into a shared coordinate system. As shown in the figure below, we voxelize this shared 3D space into a voxel grid, and calculate a spatially localized error per voxel: each voxel's error is defined as the average reprojection error of every pixel unprojected into the voxel.
Below is a visualization of the localized reward projected to keyframes. It catches local inconsistencies such as the changing plant appearance upon revisit in f2 (top row), the hallucinated shelf (second row), and the yellow artifact floating on the desk (last row).
Local-Global Blending. While the local reward is highly effective at improving 3D consistency, using it in isolation causes video quality to degrade over the course of training. We attribute this to reward hacking: the local reward fixes inconsistencies at the expense of overall quality, reducing sharpness and vibrancy, since blurrier, smoother edges lower depth errors and less vibrant (closer-to-grey) colors lower RGB errors. LoGo, via blending local with global rewards, improves geometry without degrading video quality, as shown in the reward and eval curves below.
LoGo extends the Pareto frontier in the consistency-video quality tradeoff.
Reward Interleaving. To ensure our post-trained models remain competitive on metrics complementary to geometric consistency, such as video quality and camera control, we interleave post-training with rewards targeting these properties. We adopt a cyclic schedule of consistency, aesthetic, and camera following rewards. In the pareto curve above, the dotted lines represent interleaved rewards, which comprises the "higher video quality" (up left) part of the curves.
We compare global-only reward with LoGo. Global-only reward is insufficient in fixing many localized inconsistency failures, such as artifacts or corruption, objects appearing or disappearing, or objects changing geometry or appearance.
Input image
Global
LoGo
WASD denotes camera translation, and arrows denote rotation.
Current benchmarks lack evaluation of long-horizon, camera-controlled generation. We introduce TrajectoryBench, a new benchmark for long-horizon video generation with complex camera trajectories. Compared to existing benchmarks, it contains challenging scenarios along three axes: 1) expansion: trajectories that move far from the initial view, testing whether a model can generate unseen content; 2) revisit: trajectories that revisit the same area from multiple viewing angles, testing 3D consistency; and 3) transition: trajectories with space transitions, such as passing through a door, testing whether a model can synthesize entirely new scenes. The visualization below compares a camera trajectory from WorldScore with two TrajectoryBench trajectories, one focusing on expansion and one focusing on revisit.
We evaluate LoGo on 3 different base models on 2 benchmarks, with three post-training methods. TrajectoryBench trajectories are split into “easy”, “medium” and “hard” to comprehensively capture a model's performance. The figure below shows the depth reprojection PSNR (higher better) of various methods on easy, medium, and hard splits of TrajectoryBench across three base models.
Below is the DL3DV evaluation result for three models. We show PSNR from depth reprojection, PSNR from Gaussian reconstruction, and epipolar error. Full metric tables can be found in the paper.
We evaluate LoGo for video generation with increasing lengths, up to 400 frames. While epipolar error generally increases as generation becomes longer, LoGo's advantage strengthens for longer videos, showing the increasing importance of credit assignment for challenging, long-horizon generation.
LoGo shows the importance of credit assignment in effective post-training for long-horizon video generation. However, very long horizon generation remains challenging and might require better memory design or planning. Please reach out if you are interested in this topic too!
@misc{ma2026logo,
title={LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation},
author={Ziqi Ma and Shreya Sharma and Mohamed El Banani and Katja Schwarz and Chongjie Ye and Chao-Yuan Wu and Li Fei-Fei and Ben Mildenhall and Georgia Gkioxari and Justin Johnson and Gowthami Somepalli},
year={2026},
eprint={2610.03636},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.03636},
}