Hi, thank you for sharing this great work.
I have a question regarding the RTF results in the ablation table. (Table 4)
Most ablation variants seem to share the same inference architecture, yet their reported RTF values differ considerably (e.g., Full: 0.50, duplicate weight: 0.69, w/o training: 0.76).
Could you clarify what causes these differences, and whether optical-flow preprocessing, decoding length, or post-processing time are included in the RTF measurement?
Hi, thank you for sharing this great work.
I have a question regarding the RTF results in the ablation table. (Table 4)
Most ablation variants seem to share the same inference architecture, yet their reported RTF values differ considerably (e.g., Full: 0.50, duplicate weight: 0.69, w/o training: 0.76).
Could you clarify what causes these differences, and whether optical-flow preprocessing, decoding length, or post-processing time are included in the RTF measurement?