Video summary
Clean Character Swap with Flux.2 Klein + LoRA Detailing + 4K Upscale (Full Workflow)
Main summary
Key takeaways
Main purpose
- A full Flux.2 “clean character swap” workflow (follow-up to prior Flux.2 Klein workflows) showing how to: 1) swap a subject from one image into the pose/background of a target image, 2) optionally improve face quality with masking, 3) apply LoRA-based detailing, and 4) upscale to 4K.
Model options and sampling-step guidance
- Uses the Klein 9B distilled model, expected to work in about 4–8 sampling steps.
- The speaker notes this may be a mistake and suggests sampling steps may need to be higher.
- Recommended adjustment mentioned:
- Start small (e.g., 4) but likely increase up to around 20–25, depending on results.
- Mentions alternative model formats:
- FP8 precision distilled models (planned/linked similarly).
- GGUF models loaded via a U-Net loader (GGUF placed inside a U-Net folder).
Core workflow: Subject + target scene/pose conditioning
The workflow explicitly sets required components:
- Text encoder (CLIP) via Load-Clip (select the red-colored node).
- VAE via the VAE node (also indicated by red-colored nodes).
LoRA usage
- One LoRA can be selected to be used across the workflow.
- A float/value is adjusted for image scaling consistency—especially when using larger images.
- Later, a different LoRA may be used for the detailing pass.
Node group 1: “Processing subject”
Inputs
- Subject image (image 1)
- resized using an image resize / pixel-size parameter
- prompt to change/remove background
Prompt concept
- Uses a prompt like replacing background with white to isolate/standardize the subject.
Output
- A generated subject image without the original background, so it can serve as the character reference while keeping the target scene.
Node group 2: “Processing target pose / target scene”
Inputs
- Target image (image 2)
- model + CLIP encoder + VAE connections
- megapixel size / image scale adjustment
Prompt concept
- Intended to remove clothes from the person in the target scene.
Notes
- OpenPose was tried but made the workflow more complex; a simpler technique “works.”
Preview handling
- The speaker suggests bypassing preview nodes and saving only the final output.
Node group 3: Character swap output (“subject + target pose”)
Inputs
- Both generated images from groups 1 and 2
Prompt concept
- A trigger-style prompt (speaker references a Civit AI example) instructing the model to place the character from image 1 into the pose/scene of image 2.
Logic described
- Using both images + the prompt conditions so the character matches the target.
Output
- Result -1: a clean character swap result.
Optional module: Improving face (mask-based refinement)
- An “improving result -1 face” node group exists.
- Condition
- Optional; disabled unless needed.
- How it works
- Uses image 1 result as the base
- Reloads the original subject image to create a face mask
- Requires a mask—otherwise the node errors
- Goal prompt
- Change image 1 face to match the face from image 2
- Outcome
- In the speaker’s run, they didn’t see changes, but they provide it as a fix if face quality is wrong.
- Outputs
- Produces an “improved face” image, which can replace the earlier face reference in later steps.
LoRA detailing pass (adding realism/details)
- After the base swap, a details pass generates “result -2” (or “result two”) using an added LoRA.
- Speaker notes:
- A prompt acts as a trigger word for the LoRA.
- The LoRA connection is not always directly wired to the same graph section; diffusion may begin from the model, with another LoRA used above.
- LoRA source
- Downloaded from Hugging Face as a .safetensors file (labeled realistic).
- Testing note
- Works reasonably for a 3D character, but they recommend decreasing LoRA weight slightly.
- Resolution note
- After detailing, you may need to increase resolution, then upscale to 4K.
4K upscale (Seed VR2 / DIT + VAE)
- Upscaling step to produce a 4K detailed image.
- To make “seed VR2” work:
- Requires a DIT model and VAE (downloaded automatically by the nodes).
- VAE
- A preferred VAE model is selected from available options.
Hardware considerations
- Mentions smaller GGUF models for lower-memory GPUs.
- Refers viewers to “Seed VR2 DIT models” and notes there are 3B and 7B sizes.
- Guidance:
- FP16 is good with a strong GPU; otherwise use smaller models.
Runtime result
- Example run took about 175 seconds and produced a 4K image.
Final claimed outcome / comparison
The workflow’s final output produces a character swap that is:
- clean
- with details added
- and 4K upscale results that look good in side-by-side comparison against the target/expected images.
Main speakers / sources
- Main speaker: The video’s creator/instructor (no name provided in the subtitles).
External sources referenced
- Civit AI (example prompt reference)
- Hugging Face (LoRA download)
- GitHub (Seed VR2 model details referenced)
- GGUF models (loaded via U-Net loader; model files placed in the U-Net folder)