LoGo: Local-Global Rewards for Consistent
Long‑Horizon Video Generation

Ziqi Ma1,2*, Shreya Sharma2, Mohamed El Banani2, Katja Schwarz2, Chongjie Ye2, Chao-Yuan Wu2, Li Fei-Fei2, Ben Mildenhall2, Georgia Gkioxari1, Justin Johnson2, Gowthami Somepalli2
1California Institute of Technology, 2World Labs
*Work done during internship at World Labs

We present LoGo: a post-training method to improve 3D-consistency for long-horizon camera-controlled video generation.

WASD denotes camera translation, and arrows denote rotation.

Abstract

Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to the entire generation, are poorly suited to correcting these inconsistencies over long horizons. We introduce LoGo, which blends global and spatially localized rewards for camera-controlled video models. The local reward gives fine-grained credit assignment, which substantially improves 3D consistency, while the global reward preserves camera following and video quality. Across three base models, LoGo shows a clear advantage on DL3DV and TrajectoryBench, a new benchmark for long-horizon, complex camera trajectories that current evaluations lack. LoGo effectively reduces local object shifts and artifacts, as well as global scene changes, illustrating the importance of credit assignment in post-training video models.

LoGo vs. Prior Methods

As video generation models advance toward longer horizon, various 3D inconsistency failure modes emerge, such as 1) scene layout changes upon revisit, 2) artifacts or corruption, 3) objects appearing or disappearing, and 4) objects changing geometry or appearance. While prior methods cannot effectively fix them, LoGo resolves these inconsistencies.

Input image

Input image

Base

VideoGPA

World-R1

LoGo

WASD denotes camera translation, and arrows denote rotation.

Method

The most naive method to post-train for 3D consistency is to use a global reward. However, this is not effective enough (see ablation below). Modern video models fail in largely local ways, including inconsistent object appearance, hallucinated or disappearing objects, and local artifacts. A good reward should penalize these failures directly rather than merely ranking full-length videos. Motivated by this, we design a reward localization method for accurate credit assignment during post-training. LoGo introduces three components: 1) reward localization, 2) Local-Global blending, 3) reward interleaving.

Reward Localization. LoGo is based on 3D scene reprojection. Given a video, we first construct a scene point cloud by passing keyframes to VGGT-Ω and unproject each pixel using the predicted depth into a shared coordinate system. As shown in the figure below, we voxelize this shared 3D space into a voxel grid, and calculate a spatially localized error per voxel: each voxel's error is defined as the average reprojection error of every pixel unprojected into the voxel.

Voxel-level reward localization: pixels from several frames unproject into a shared voxel grid

Below is a visualization of the localized reward projected to keyframes. It catches local inconsistencies such as the changing plant appearance upon revisit in f2 (top row), the hallucinated shelf (second row), and the yellow artifact floating on the desk (last row).

Localized reward projected to keyframes (red = bad, blue = good)

Local-Global Blending. While the local reward is highly effective at improving 3D consistency, using it in isolation causes video quality to degrade over the course of training. We attribute this to reward hacking: the local reward fixes inconsistencies at the expense of overall quality, reducing sharpness and vibrancy, since blurrier, smoother edges lower depth errors and less vibrant (closer-to-grey) colors lower RGB errors. LoGo, via blending local with global rewards, improves geometry without degrading video quality, as shown in the reward and eval curves below.

Geometry reward and validation video quality over training for Local-only and LoGo

LoGo extends the Pareto frontier in the consistency-video quality tradeoff.

Pareto frontier of 3D consistency vs. video quality for Global-only, Local-only and LoGo

Reward Interleaving. To ensure our post-trained models remain competitive on metrics complementary to geometric consistency, such as video quality and camera control, we interleave post-training with rewards targeting these properties. We adopt a cyclic schedule of consistency, aesthetic, and camera following rewards. In the pareto curve above, the dotted lines represent interleaved rewards, which comprises the "higher video quality" (up left) part of the curves.

Ablation: Global vs. LoGo

We compare global-only reward with LoGo. Global-only reward is insufficient in fixing many localized inconsistency failures, such as artifacts or corruption, objects appearing or disappearing, or objects changing geometry or appearance.

Input image

Input image

Global

LoGo

WASD denotes camera translation, and arrows denote rotation.

TrajectoryBench

Current benchmarks lack evaluation of long-horizon, camera-controlled generation. We introduce TrajectoryBench, a new benchmark for long-horizon video generation with complex camera trajectories. Compared to existing benchmarks, it contains challenging scenarios along three axes: 1) expansion: trajectories that move far from the initial view, testing whether a model can generate unseen content; 2) revisit: trajectories that revisit the same area from multiple viewing angles, testing 3D consistency; and 3) transition: trajectories with space transitions, such as passing through a door, testing whether a model can synthesize entirely new scenes. The visualization below compares a camera trajectory from WorldScore with two TrajectoryBench trajectories, one focusing on expansion and one focusing on revisit.

WorldScore vs. TrajectoryBench camera trajectories (expansion and revisit)

Quantitative Results

We evaluate LoGo on 3 different base models on 2 benchmarks, with three post-training methods. TrajectoryBench trajectories are split into “easy”, “medium” and “hard” to comprehensively capture a model's performance. The figure below shows the depth reprojection PSNR (higher better) of various methods on easy, medium, and hard splits of TrajectoryBench across three base models.

PSNR by difficulty

Below is the DL3DV evaluation result for three models. We show PSNR from depth reprojection, PSNR from Gaussian reconstruction, and epipolar error. Full metric tables can be found in the paper.

DL3DV PSNR-depth, PSNR-Gaussian and epipolar error for Base, VideoGPA, World-R1 and LoGo on three base models

Longer Videos

We evaluate LoGo for video generation with increasing lengths, up to 400 frames. While epipolar error generally increases as generation becomes longer, LoGo's advantage strengthens for longer videos, showing the increasing importance of credit assignment for challenging, long-horizon generation.

Epipolar error vs. video length

What's Next?

LoGo shows the importance of credit assignment in effective post-training for long-horizon video generation. However, very long horizon generation remains challenging and might require better memory design or planning. Please reach out if you are interested in this topic too!

BibTeX

@misc{ma2026logo,
  title={LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation},
  author={Ziqi Ma and Shreya Sharma and Mohamed El Banani and Katja Schwarz and Chongjie Ye and Chao-Yuan Wu and Li Fei-Fei and Ben Mildenhall and Georgia Gkioxari and Justin Johnson and Gowthami Somepalli},
  year={2026},
  eprint={2610.03636},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2610.03636},
}