WEBVTT

00:00.095 --> 00:02.945
This is ComfyUI for
advertising: talent casting.

00:04.105 --> 00:08.005
This workflow takes a talent headshot
and generates a casting previs, a

00:08.005 --> 00:11.675
still frame and a short video clip
to conceptualize how the talent would

00:11.675 --> 00:13.545
look in the actual production scenario.

00:13.795 --> 00:15.375
Start by loading the talent headshot.

00:15.885 --> 00:17.735
Then the workflow splits into two steps.

00:18.225 --> 00:22.255
Run step one first, then use the
output as the input for step two.

00:22.885 --> 00:26.525
Step one generates the previs
still frame using Gemini Image 2.

00:26.995 --> 00:31.535
Describe the scene in the text node,
the camera framing, location, action,

00:31.785 --> 00:33.375
and any other relevant details.

00:33.795 --> 00:37.245
Optionally, load a product image
and an environment reference to

00:37.245 --> 00:38.695
give the model more to work with.

00:38.885 --> 00:40.365
Bypass them if not needed.

00:40.575 --> 00:43.685
The output is a production quality
image placing the talent in the

00:43.685 --> 00:47.055
described scene with the product
and environment as references.

00:49.275 --> 00:51.315
Step two generates the previs video.

00:51.755 --> 00:55.745
Load the still frame from step one as
the first frame, and Kling generates

00:55.745 --> 00:57.475
a short clip animating the scene.

00:58.035 --> 01:02.055
Describe the action in the text node,
what the talent does and how the camera

01:02.055 --> 01:06.025
moves, and the output will visualize
the input still frame in motion.
