Just ran a quick side-by-side, with telling results.
LTX 2.5:
Silhouette Matte ML + Matte Assist + Matte Refine:
Flame AutoMatte
Source clip: https://artlist.io/stock-footage/search?terms=697723
This test was done at 3200x1900 resolution, for 49 frames.
Flame ran with it’s usual parameters on Linux.
Silhouette ran on Linux, used about 10GB NVRAM
LTX ran on Linux, used about 90GB of NVRAM, using bf16 version of LTX 2.5 (42GB model file)
Flame results were almost instantenous after the ML model loading
Silhouette took about 1min
LTX ran in about 2.5min
The LTX model was a more powerful version than what Flame or Autodesk distribute. So it had the advantage.
LTX model has some frame count limitations, so running it on long and large clips will take a bit of extra work. Also on long/large clips you run into RAM limitations of Comfy itself.
I had existing LTX workflows for Comfy, so setup was less than 10min.
All workflows read/wrote EXR sequences, only final export for posting here was .mp4
Is LTX 2.5 using same process as VideoMaMa: Mask-Guided Video Matting via Generative Prior, or something else? I was able to find this: “The IC-LoRA conditions on the latents of the RGB reference video and generates the corresponding alpha matte as a video, with an empty text prompt. Because the reference stays attached for the entire denoise (stage-1-only inference at native resolution), the matte stays aligned with the source down to fine edges.”
I don’t know the exact mechanics of LTX. But it is one core model and then a number IC loras for different workflows, different from the earlier loras which just bias style.
And the prompt was empty indeed on my test. Works like AutoMatte in finding the dominant object in frame.
It’s worth noting the business model here. LTX is not like Qwen or Wan which have been open models under GPL or similar license, until their owners think they had enough PR and lock future version up behind paywalls as happened with Wan 3.0.
LTX is free for commercial use under a certain turnover threshold. Above that you have to pay for a commercial license. The company’s goal is to make models, they’re not a byproduct of something else. I hope there will be more like that.
Not that paying reasonable API rates is a problem. As long as we have a local option.
Essentially the Nuke Indie model.
ADSK needs to stop spending resources making in-house sub-standard models, and focus on seamless integration of 3rd party SOTA models. Batch should have been what ComfyUI now is.
You say that autodesk should have turned batch into comfyui, but think how grubby the pearls would have become from all of the fervent clutching.
Reflect on just how vibrant and productive the community development has become.
Years of domain specific knowledge and insight has brought about significant advances, backed up by twenty years of Reddit generality and an electricity bill you didn’t have to pay for.
I’m happy that elbows are finally pythonic.
It helps organize the feeds into the comfyui batch nodes.
I see. Thanks for the info. I’ll have to look more into it.
Some possible pointers on why this model performed so well, and also implicit limitations.
The first one is that is appears to attention over the whole clip, rather than frame by frame. That helps it deal with frames where conditions from frame-by-frame models get worse because of contrast and color changes. It also means it can handle occlusions better. Temporal stability is inherent to its design.
The downside is that this requires significantly more compute and doesn’t scale with clip length. The model is currently limited to 145 frames, and because of whole clip attention you may have limited luck breaking longer clips in to segments without getting noticeable issues at the seams. May have to pick strategic seams.
The other factor is - classic loras used to just nudges the models weights to favor a specific trained preference. An IC lora adds the additional reference input (IC = in context), and with it additional training data.
This has two consequences in this particular model: this model by-passes the hallucination phase of the model processing. The matte pixels you see come from the reference input, they’re not pixels the model has hallucinated, as all the other models do we are familiar with.
Other matting models decide for each pixel, what matte should look like based on their training. Hence the flickering and inconsistency. This model just decides if a pixel is foreground or background based on its training data. Foreground = output is max(r,g,b) for all 3 channels. Background is black.
The second consequence is that the core model is a 22B diffusion model, which has seen several magnitudes more training data than all the matting models ADSK and BorisFX have trained. This is at the scale of other diffusion models, not matting aux models. That’s an unfair advantage towards the training capabilities of the other models.
And in my case I ran the bf16 version of the model, not the further quantized smaller variants. So it had additional precision to work with.
So a lot more memory, a whole lot more compute, and some stricter boundaries on what can be processed. But if you stay within those boundaries, holy cow, this really works.
So this model does take a whole different approach to ML matting than what we’ve seen until now. Similar to how SmartRoto has taken a whole different approach by pursuing ML-masking instead of ML-matting.
A question for ADSK is that currently they’re shipping models that are tuned to a wide range of GPUs. Which is nice if you have a A5000 ADA card and get nice automatte results (maybe not from fine hair, but a wide range of objects - we’ve seen the kudos to AutoMatte).
But it also means that if someone goes out and sources a bigger GPU, that could handle a lot more compute, they’re not getting better results for the most part. It’s one model for all. There is a point to be made that there should the S, M, and XL models.
If you need to crack a hard nut and you have the compute, I don’t mind downloading a 42GB model file and having it go for it. And I doubt any of you would.