FOFPred v1.0
FOFPred is a diffusion-based model that predicts future optical flow from a single image guided by natural language instructions. Given an input image and a text prompt describing a desired action (e.g., "Moving the water bottle from right to left"), FOFPred generates four sequential optical flow frames showing how objects would move.
Diffusion-based model for future optical flow prediction
Guided by natural language instructions
Generates sequential optical flow frames from a single image