AI SDR->HDR 16fp EXR - via Beeble (and a let down)

OK, so lots of you probably got the email about Beeble’s new features in Switch-X that can restore and SDR clip to an HDR clip and export it with 16fp EXR in ACES2065-1.

Finally a solution to the 8 bit AI purgatory?

Unfortunately not entirely… Still usable, but not the solution we’re looking for. It’s a generative image edit model, not a discriminative AI model.

I ran a test case:

Arri Mini clip in LogC with a bit of a blown out sky and just a bit of clouds left.

I graded it as reference.

Then I exported it without grade as Rec709 (applying Arri IDT/DRT/ODT). Took that clip into Comfy and removed one person which was overlapping with the sky. A realistic scenario.

Then I took the result, which was a Rec709 render from Comfy into Beeble and ran it through the new model, exported EXR sequences in ACES2065-1 and brought it back into Flame, next to the original:

Left Beeble clip, right, original graded. Notice something? The clouds are different.

What Beeble does is NOT enhance your original SDR pixels to rejuvenate them to HDR. It hallucinates new HDR pixels that could fit in that space via a prompt. Well, we could have done that via a standard prompt and some HDR material. It’s just faster via Beeble.

But what we really want is the original pixels with more bit depth and gamut restored, including that removed area that got clobbered. The in-paint would have used 8 bit pixels which could have been rejuvenated based on the same information from the neighbors and looked good.

And you can see that in their workflow:

You need to adjust two sliders to show the model what you consider highlights and shadows. And then you need to provide prompts of what should go there. It does have a button to auto-generate the prompts based on image analysis.

But yes, these are brand new pixels, not re-furbed pixels.

So close, and yet so far.

Hey Jan,

we did some tests with beeble and… it’s not perfect. Here are my chaotic 4 euro cents.
TL:DR;
Ignore Beeble. It does not deliver production results, not their fault in my oppinion, more poor dataset. Train your own model 8bit > 16bit

Full Banana:
It ended with me training our own model. Unfortunately, I cannot share it due to footage I used for training.
What I can tell you that training the model to handle this task is actually much easier and not super scary :slight_smile:

I first did a version using NAFNet (I know it’s not perfect choice due to architecture and did force some workarounds) becuase I needed something fast. I’ve trained the model to go from rec709 rgb (sRGB is too similar so it does not matter in the end run :D) to acescct and, because this is what NAFNet does really well, clean up compression artifacts.

Now I am doing a more intelligent version (not AI intelligent, I learned enough not to be an idiot :smiley:) and decided to decouple the process into separate stages. First cleanup artifacts from compressed source, the source being usually mp4 from Higgsfield/Weavy, than pass it through a model that is build around NAFNet UNet with separate global context extractor and FiLM conditioning with sigmoid output.
From what I’ve learned going rec > acescg is… optimistic.
ACEScg has linear unbounded highlights and shadows. When I tried training a model I had to arbitrarily decide where is my white point (7? 11? maybe 8.3?) and boyyyyyy. Did it produce funny results. After a week of trying to get it to work I ended back with my original idea.
8bit source > AI rec to acescct > acescct to any other aces using OCIO.
This way I am doing a proper conversion into aces world but I avoid negative values in shadows (cct has a linear segment in deep shadows so clean black at 0), I have a nice [0,1] log curve with middle gray at ~0.41, highlights with scene 1 nit at ~0.55 and high speculars in 0.6-0.8 range. And to top it all off a well defined finite range, no linear crazy
I know it seems counter intuitive but most colour scientists that shouted at me over the years did tell me (with a varying levels of desperation in voice) to never go rec > acescg directly. Which of course, I always ignored.

As a side quest I tried to also train a non-prompt (to have it in flame) diffusion stage to hallucinate details (shadows and highlights) but I gave up. Even with the dataset I posses it was not big enough to have a significant result.

So. My solution? Ignore beeble which tries well but most likely has limited or synthetic dataset, train your own model. While I cannot share the resulting models I can share the methodology and basic architecture of training scripts I used. I train on linux boxes that are ada 5000 with 32gb vram after artists go to bed :smiley: so my days are really long these past weeks.

Forgot to write. Hallucinating details can be, and should be, handled in Comfy. The tensors are 16 or 32bit and you can save exr files. So you can plonk your trained model into Comfy and use it before save to exr.

If you have to use some online model like seedance. Yeah we are buggered. But for a lot of “details in bright highlights” wan/ltx is good enough. This is a fascinating subject to be honest and if you do not stop me I will talk about it for hours.

This is fantastic. I’m not opposed to doing some training of models if that is the answer.

Weavy and Beeble seemed to be more aligned with the VFX community than the garden variety models out there. But they still fall short too often what we need.

I can send you an invite to a Discord server we setup for folks doing this type of work (coding and building AI tools for serious post pipelines). That’s the perfect place to geek out on this stuff.

His Majesty @andymilkis did invite me to the main Discord server, i’m there as Panda_VFX. I’m not too discordy, because there is only 24hours in a day. A serious mistake IMHO.
But yes, please add me :slight_smile: If people are actually interested I might record a short (knowing me ~4h) YouTube semi - tutorial on how to start working with this problem.

Hi, I got in contact with Gunwoo Kim from Beeble on LinkedIn. I asked if there is an actual HDR encoded clip available and he gave me this YouTube Link:

I must say, yes the images are brighter in HDR, but I don’t understand what the source was for each shot. The SDR clips look weirdly clipped or even a bit like log footage.

But the HDR version just looks terrible on an Apple XDR screen. A good HDR grade should give me the feeling that there was more light energy in the scene. So I would expect that the whole image should look brighter plus some additional bright highlights. Right now it feels like: push highlights up and gamma the entire image down at the same time.

This is a comp right? That they’re taking from SDR to HDR? This looks really bad to me. She looks very unbalanced with the BG and the perspective seems wonky the HDR looks objectively weird and worse to me.

I’ll say, I don’t honestly understand the fascination with Beeble and relighting. The push doesn’t need to be a tool to relight everything, the push should just be better dialogue between directors and DPs and VFX artists so that we shoot smartly, light correctly, and make the whole process smooth without the need to slather photography with digital goop all the time. Like just taking rotten ingredients and seasoning the shit out of them to hide the rank taste is a sad way to have to work in the kitchen (and often times just ends up nauseating anyway). But alas, I know the industry we work in and fix it in post seems to be no longer a safety net for production but a lazy inevitability.

I think there are two scenarios…

  1. Improve the material for a true HDR grade. That bar would be very high.
  2. Adding enough dynamic range back into the footage, so that it can hold up to camera file neighbors in an SDR grade.

I was mostly hoping for the latter, but alas.

In an ideal world, I would agree. But in that same world we wouldn’t need to do much beauty retouch if the hair and make-up department would take a bit more time, or any other number of things that could be done with less effort on set.

One good example I had recently, where I had to relight an AI generated clip. The subject needed to be backlit to fit the night scene. But the AI can only produce good looking front-lit daytime versions. So it was either relight, or it ain’t gonna happen at all.

I’m happy to have it in the toolbox. But just like beauty work, it’s a fix-it tool, not a design tool.

I’ve used beeble, don’t get me wrong here. For what it is it is what it is. I can swing around my AI hammers like plenty of folks. Not going to say it’s not the same frustration I have had for years now of trying to make food that people expect to taste like gourmet food with rotten ingredients (the rotten ingredients used to be poorly lit green screen shots and seemingly arbitrarily chosen stock footage, now we have a whole new batch that needs fixing). It’s tiresome to start from a deficit so often and end up just trying to make something work as opposed to make something look good.

I’ll add that lighting is not really anything comparable to pimples or flyaways. Lighting is literally the photography itself. No make up I have images. No lighting I have black.

My take is that these ‘HDR’ models are just for niche use case where you have an AI generated clip, and you just want to add some ‘better’ highlights. So not for comp round trips.

For round tripping from comp on shots where the highlights fall apart in AI, I typically use a Look node to bring the highlights close to 1 and adjusting the gamma to taste before baking in the view transform and exporting as rec709. The logic being you just need to put the data you want to process in the bucket. It’s going to get sloshed around anyway, you might as well consciously choose what data you want sloshed about.

Then coming back in it’s a reverse view transform, then the Look node again with invert checked. Then usually divide the clip by a blurred version of itself, and multiply by an (equally) blurred version of the original to do the last 5% of colour matching.

Look → View Transform → AI junk → Reverse View Transform → Inverted Look → Divide/Multiply

Until there are models trained on HDR material, everything is just a hack.