This fork is tested only with Linux (ubuntu 22.04 and 24.04 in my case) and WSL 24.04 on Windows, pure Windows "should" work for gui versions but pipeline_master_gui and all auto batches uses bash scripts that needs to be reimplemented using powershell.
All GUI works too in WSL so moving to Windows is quite pointless
Aim to this Fork is to create a full sbs 1080p 3d content as result with a one-click solution but keeping ways to customize it.
Supports only 8 bit format (inpainting is 8 bit anyway so it's actually impossible to get 10 bit hdr end to end) with hardcoded yuv444p during all steps to preserve colors transitions.
All runner scripts are tuned for intel 265k, 48GB ram and RTX 4090 and specifically for 1920*800 content
Most of the job is done by VIBE CODING using VS code + Codex 5.3/5.4
All the extra scripts are made multithreading/parallel when possible, helping reducing time (but is still a very very slow process)
python pipeline_master_gui.pyjust run scenedetect TAB manually and Verify(quick) then try a test run on the last tab, it will pick the first 5 unfinished clips and do all the steps, if everything went well press the run/resume button and go on vacation, when you are back it should have finished :)
NOTE about PYTORCH_ALLOC_CONF max_split_size_mb:xx, while this command can save from VRAM OOM be warned that a low value such as 64 can increase inference times by up to 200%, i have implemented those exports to be enabled only when there are fails.
garbage_collection_threshold:0.8 and expandable_segments:True are very light so i kept them on as default for depthcrafting and inpainting steps.
python runners/pipeline_master_headless.py --helpThe headless runner uses a single work_dir as the source of truth.
If <work_dir>/config_pipeline_master_gui.json exists it loads the saved Pipeline Master settings from there automatically.
If <work_dir>/pipeline_state.json exists it also resumes the DONE/VERIFY state from there automatically.
If the local config json does not exist yet, it starts from the normal Pipeline Master defaults but still keeps paths/state inside that same work_dir.
The supported CLI flags and step-selection options are documented directly by the script help above.
All the scripts can be launched as stand-alone, or you can run the gui versions like the originating fork where i've started but some features will be missing (sharpness csv, autoct csv, "prepare seg mono to sbs" and sharpen step has no GUI, mask for merge step is embeddded into merging_gui)
Below the full manual pipeline (using all features)
As a general rule:
- open the gui version and check values you want to use
- open the corresponding sh runner and change values (they are all available at the top) and eventually paths
- run
- next step
Scenedetect CSV is NOT available as stand alone, is a single line command btw. Use just csv creation, then go on as below
python Utilities/split_scenes_from_csv.py
# manually move mono files to seg-mono folder
./runners/run_depthcrafter_nogui_batch.sh
./runners/run_splatting_runner_parallel.sh
python Utilities/analyze_inpaint_sharpness.py
./runners/run_inpainting_runner.sh
./runners/run_inpaint_sharpen_runner.sh
./runners/run_mask_formerge_nogui.sh
python Utilities/analyze_auto_ct_csv.py
./runners/run_merging_nogui_batch_parallel.sh #or runners/run_merging_nogui_batch.sh for single thread
python Utilities/prepare_seg_mono_to_sbs.py # if you have mono files on seg-mono
./Utilities/Rejoin_HEVC_NVENC.sh
./Utilities/remux_replace_video_mkvtoolnix.shIf you find this useful, you can leave a tip through my Ko-fi page by using the Sponsor button at the top of this repository. Every contribution is appreciated and helps me dedicate more time to updates, fixes, and new projects.
You can learn more about DepthCrafter GUI Seg here.
-
GIT: Ensure Git is installed and added to your system’s PATH.
Download here: https://git-scm.com/downloads/win
You can check the installation by running the command:
git --version
If it shows a version, Git is installed and on PATH. -
CUDA ToolKit: Ensure CUDA 12.8 is installed and added to your PATH.
Download here: https://developer.nvidia.com/cuda-12-8-0-download-archive?target_os=Windows&target_arch=x86_64 -
FFMPEG: Ensure FFMpeg is installed and added to your PATH.
See Here for a tutorial on how to install.
- Run script from folder where you want StereoCrafter installed
- Download and extract model "weights" to StereoCrafter folder (use qBittorrent to download)
For Manual Install Instructions Click Here
StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos
Sijie Zhao*
Wenbo Hu*
Xiaodong Cun*
Yong Zhang†
Xiaoyu Li†
Zhe Kong
Xiangjun Gao
Muyao Niu
Ying Shan
* equal contribution † corresponding author
We propose a novel framework to convert any 2D videos to immersive stereoscopic 3D ones that can be viewed on different display devices, like 3D Glasses, Apple Vision Pro and 3D Display. It can be applied to various video sources, such as movies, vlogs, 3D cartoons, and AIGC videos.
2024/12/27We released our inference code and model weights.2024/09/11We submitted our technical report on arXiv and released our project page.
Here we show some examples of input videos and their corresponding stereo outputs in Anaglyph 3D format.
We run our code on Python 3.8 and Cuda 11.8. You can use Anaconda or Docker to build this basic environment.
# use --recursive to clone the dependent submodules
git clone --recursive https://github.com/TencentARC/StereoCrafter
cd StereoCrafterpip install -r requirements.txtcd ./dependency/Forward-Warp
chmod a+x install.sh
./install.sh
1. Download the SVD img2vid model for the image encoder and VAE.
# in StereoCrafter project root directory
mkdir weights
cd ./weights
git lfs install
git clone https://huggingface.co/stabilityai/stable-video-diffusion-img2vid-xt-1-12. Download the DepthCrafter model for the video depth estimation.
git clone https://huggingface.co/tencent/DepthCrafter3. Download the StereoCrafter model for the stereo video generation.
git clone https://huggingface.co/TencentARC/StereoCrafterScript:
# in StereoCrafter project root directory
sh run_inference.shThere are two main steps in this script for generating stereo video.
Execute the following command:
python depth_splatting_inference.py --pre_trained_path [PATH] --unet_path [PATH]
--input_video_path [PATH] --output_video_path [PATH]Arguments:
--pre_trained_path: Path to the SVD img2vid model weights (e.g.,./weights/stable-video-diffusion-img2vid-xt-1-1).--unet_path: Path to the DepthCrafter model weights (e.g.,./weights/DepthCrafter).--input_video_path: Path to the input video (e.g.,./legacy/source_video/camel.mp4).--output_video_path: Path to the output video (e.g.,./outputs/camel_splatting_results.mp4).--max_disp: Parameter controlling the maximum disparity between the generated right video and the input left video. Default value is20pixels.
The first step generates a video grid with input video, visualized depth map, occlusion mask, and splatting right video, as shown below:
Execute the following command:
python inpainting_inference.py --pre_trained_path [PATH] --unet_path [PATH]
--input_video_path [PATH] --save_dir [PATH]Arguments:
-
--pre_trained_path: Path to the SVD img2vid model weights (e.g.,./weights/stable-video-diffusion-img2vid-xt-1-1). -
--unet_path: Path to the StereoCrafter model weights (e.g.,./weights/StereoCrafter). -
--input_video_path: Path to the splatting video result generated by the first stage (e.g.,./outputs/camel_splatting_results.mp4). -
--save_dir: Directory for the output stereo video (e.g.,./outputs). -
--tile_num: The number of tiles in width and height dimensions for tiled processing, which allows for handling high resolution input without requiring more GPU memory. The default value is1(1$\times$ 1 tile). For input videos with a resolution of 2K or higher, you could use more tiles to avoid running out of memory.
The stereo video inpainting generates the stereo video result in side-by-side format and anaglyph 3D format, as shown below:
We would like to express our gratitude to the following open-source projects:
- Stable Video Diffusion: A latent diffusion model trained to generate video clips from an image or text conditioning.
- DepthCrafter: A novel method to generate temporally consistent depth sequences from videos.
@article{zhao2024stereocrafter,
title={Stereocrafter: Diffusion-based generation of long and high-fidelity stereoscopic 3d from monocular videos},
author={Zhao, Sijie and Hu, Wenbo and Cun, Xiaodong and Zhang, Yong and Li, Xiaoyu and Kong, Zhe and Gao, Xiangjun and Niu, Muyao and Shan, Ying},
journal={arXiv preprint arXiv:2409.07447},
year={2024}
}



