Qualitative Demo on DAVIS
We qualitatively demonstrate our method's forecasting ability in unstructured scenes. We train the model on the pseudolabeled Kinetics dataset and evaluate it on DAVIS. For this experiment only, we use a model at the DiT-B scale. In the examples below, the model tracks the first 8 frames and forecasts the next 16; see the bottom of the page for failure-case analysis.