I would like to try some ideas i’ve had in my previous job and recently while researching video world models. Similar to how LLMs use speculative decoding to speed-up inference with near perfect alingment with the bigger models, Speculative Decoding for auto-regressive diffussive generation of videos is starting to appear in the research of mid 2026.
My idea is not only to implement it using WAN 2.1 or some other SOTA small model making 1) a contribution to the model zoo potentially, but also pushing a research idea that I had while reading this paper on speculative decoding (https://arxiv.org/html/2604.17397v1#:~:text=T%2DStitch%2C%20SRDiffusion%2C%20and%20HybridStitch%20use%20fixed%20step,changes%20to%20either%20the%20drafter%20or%20target) of using V-JEPA 2 / SigILP2 as model router instead of the ImageReward model they use 2) to improve upon the work of Hu. et Al.
I know its quite ambitious given that as far as I’ve seen I dont see many transformers in the model zoo nor video diffusers (Like Wan 2 or OpenSora) nor V-JEPA 2, but If I get more information about the ONNX operations compatibility that the voyage compiler offers and I fight a bit with it, we could get it and push the frontier of which models can run on Metis. I’ve already did the computations of the weights and 16GB should be way more than enough, so do not worry on that front.
So I would like to get more feedback on feasibility because although its more research oriented rather than product oriented like other ideas on the forum, I would love to work on it if given the chance :)
Question
Video
Sign up
Already have an account? Login
Log in, or create an Axelera AI account
Log In or Register HereLogin to the community
No account yet? Create an account
Log in, or create an Axelera AI account
Log In or Register HereEnter your E-mail address. We'll send you an e-mail with instructions to reset your password.
