@Curly_rincewind's banner p

Curly_rincewind


				

				

				
0 followers   follows 1 user  
joined 2025 August 18 14:30:44 UTC

				

User ID: 3895

Curly_rincewind


				
				
				

				
0 followers   follows 1 user   joined 2025 August 18 14:30:44 UTC

					

No bio...


					

User ID: 3895

The bar for me is not that it's recognizably consistent. It's actual consistence. For something like this to cross the commercial viability threshold stuff needs to stay on model.

The character needs to stay consistent in different lighting conditions, angles and FOVs.

Finally it needs to be able to handle unique appearences, not average pretty faces and clothes.

The issue isn't that it's impossible to make a video of a character from an image be consistent with that image. Although in my opinion we're still not there. The difficulty arrises from the fact that such a video will inevitably have to conjure up new details in the process. Keeping the newly created information consistent with the next generated clip gets exponentially harder with each new clip and required context. Similar to how LLMs fail if the context is long enough.

I doubt you can make something like bill gates wearing tiger face paint and a floppy sleeping cap from a flat front shot, to an over the shoulder partial view, to a side view without messing up the direction of the flop of the cap or the position/amount of tiger stripes in the make-up.

Not going to bet money on it because I'm sure with enough tries it's doable, I'm just illustrating a point that the amount of stripes and flops or whatever is essentially the same as subtle facial features like the angle of the jaw or the tilt of the eyes.

The technology is fundamentally just not designed for this sort of thing. There's tons of workarounds and it will still be very impactful, you can work within the constraint to achieve amazing stuff, but the constraints are still there.

I think by and large they are terrible at it and don't. There are a few different techniques that claim to achieve this, but as someone who follows this closely it's all still fairly bad. By far one of the biggest remaining hurdles of mass commercial use.

Matching Eye colour hair colour, clothes etc are doable with stuff like retraining the model, a reference or prompting with a well known actor/figure

God forbid you try to recreate a character that passes the filter of someone who's not faceblind