comfyui-character-gen

Build identity-preserving character generation workflows and pipelines in ComfyUI. Selects the optimal identity method (InfiniteYou, FLUX Kontext, PuLID, InstantID, IP-Adapter) based on use case requirements. Handles face preservation, likeness transfer, cross-domain conversion (3D to photo), multi-

By mckruz · 541 installs

npx skills add mckruz/comfyui-expert --skill comfyui-character-gen

Source repository · Upstream listing

ComfyUI Character Generation Expert Build production ready ComfyUI workflows for consistent character generation across image, video, and voice modalities. Quick Decision: Which Approach? Starting from reference images (like 3D renders)? → InfiniteYou (state of the art 2025) or InstantID + IP Adapter (proven, lower VRAM) Need highest identity fidelity? → FLUX.2 (NEW 2026: up to 10 ref images) or PuLID Flux II (no model pollution) Want iterative editing without retraining? → FLUX Kontext (context aware, maintains consistency across edits) Creating video content? → LTX 2 (NEW 2026: 4K production ready), Wan 2.2 MoE (film level), or FramePack (60 sec on 6GB!) Need voice for character? → TTS Audio Suite (unified platform, 23 languages) or F5 TTS Cross Lingual (NEW 2026) Core Workflow Patterns Pattern 1: Zero Shot Character Generation (No Training) Best for: Quick iteration, 3D to photorealism conversion, limited reference images Critical settings: CFG: 4 5 (prevents burning with InstantID) Resolution: 1016×1016 (avoids watermark artifacts) IP Adapter weight: 0.6 0.8 InstantID noise injection: 35% to negative See references/workflows.md for complete node configurations. Pattern 2: LoRA + Identity Methods (Maximum Consistency) Best for: Production work, character series, video generation base Training requirements: 15 30 images, varied poses/expressions/lighting Unique trigger word (e.g., "sage character") See references/lora training.md for full parameters Pattern 3: Video Generation Pipeline Best for: Talking heads, character animation, promotional content Model selection: Wan 2.1 14B: Best quality, 24GB+ VRAM, slower Wan 2.1 1.3B: 8GB VRAM, good quality, faster AnimateDiff Lightning: Fastest, best for iteration Pattern 4: Talking Head with Voice Best for: Character dialogue, presentations, social content Two approaches available: See references/talking head workflows.md for complete workflows and references/voice synthesis.md for voice creation options. Model Recommendations (2026 Updated) Image Generation Use Case Model Notes Best photorealism FLUX.1 dev Slow but superior quality Multi reference consistency FLUX.2 NEW 2026 : Up to 10 ref images, strong identity preservation Fast iteration RealVisXL V5.0 Good balance speed/quality Character editing FLUX Kontext Context aware, maintains consistency across edits Iterative refinement FLUX Kontext Pro/Max 8x faster than GPT Image (API) Identity Preservation (2026 State of Art) Method Best For VRAM Notes FLUX.2 Multi reference consistency 24GB+ NEW 2026 : Up to 10 ref images, branded content InfiniteYou Highest identity match 24GB ICCV 2025 Highlight, SIM/AES variants FLUX Kontext Iterative editing 12 32GB Built in consistency, no retraining PuLID Flux II Dual characters, no pollution 24 40GB Contrastive alignment solves model pollution AuraFace Commercial identity encoding 12GB NEW 2026 : Open source ArcFace alternative InstantID Style transfer, 3D→realistic 12GB Maintenance mode but still excellent IP Adapter FaceID Speed, lower VRAM 6GB+ Good baseline approach Video Generation Model Quality Speed VRAM Notes LTX 2 ★★★★★ Medium 16GB+ NEW 2026 : First open source 4K audio+video, production ready Wan 2.2 MoE ★★★★★ Slow 24GB+ Film level aesthetics, first+last frame control FramePack ★★★★★ Medium 6GB 60 sec videos, VRAM invariant breakthrough Wan 2.1 1.3B ★★★★ Medium 8GB+ Consumer friendly AnimateDiff V3 ★★★ Fast 8GB Motion/camera LoRAs, infinite length Voice/TTS Tool License Quality Features TTS Audio Suite Multi ★★★★★ Unified platform, 23 languages, emotion control F5 TTS MIT ★★★★ Zero shot from <15 sec samples, Cross Lingual 2026 Chatterbox MIT ★★★★★ Paralinguistic tags ( [laugh] , [sigh] ), 4 voices IndexTTS 2 MIT ★★★★ 8 emotion vector control ElevenLabs Commercial ★★★★★ Production quality (API) Essential Custom Nodes Install via ComfyUI Manager: RTX 50 Series Optimization (NEW 2026) With 32GB VRAM on RTX 5090, run most workflows without optimization. ComfyUI v0.8.1 adds major RTX 50 Series enhancements: NEW v0.8.1 Features: NVFP4/NVFP8 precision formats : 3x faster performance, 60% VRAM reduction on RTX 50 Series Weight streaming : Uses system RAM when VRAM exhausted, enables larger models on mid range GPUs Enable tiled VAE for 8K+ upscaling Batch 4× 1024×1024 generations in parallel Run Wan 2.2 14B + LTX 2 natively Use FP8 quantization for FLUX (50% VRAM reduction) Workflow Generation Process When building a workflow for a user: 1. Clarify the goal : Image only? Video? With voice? What's the source material? 2. Select the pipeline pattern from above based on requirements 3. Generate the workflow following node configurations in references/workflows.md 4. Include model downloads with exact filenames and paths from references/models.md 5. Provide parameter recommendations specific to their hardware/use case Reference Files references/research log.md Latest techniques : InfiniteYou, FLUX Kontext, PuLID Flux II, Wan 2.2 MoE, FramePack, FLUX.2, LTX 2.3, Wan 2.6, Qwen3 TTS, and more references/models.md Complete model list with HuggingFace/Civitai links, file paths, and compatibility notes references/workflows.md Detailed node by node workflow templates for each pattern references/lora training.md LoRA training guide with Kohya/AI Toolkit parameters references/voice synthesis.md Voice cloning, TTS, and lip sync pipeline details references/talking head workflows.md Complete talking head workflows : Image→Talking Head (SadTalker, LivePortrait) and Video→Add Voice (Wav2Lip) with production scripts references/evolution.md Update sources, changelog, and user specific learnings Skill Evolution This skill is designed to evolve. When helping the user: Before starting a workflow: Check if new models have dropped that might be better (search HuggingFace/Civitai if uncertain) Consider if user's past successes/failures inform the approach After completing a workflow: Note what worked well or poorly for future reference If user discovers better settings, update the relevant reference file Proactive updates: When the user mentions a new model or technique, research and integrate it Periodically suggest checking for updates to key dependencies See references/evolution.md for monitoring sources and update protocols. Example: 3D Render to Photorealistic Character For converting stylized 3D renders (like game/VN characters) to photorealistic images: Recommended approach: InstantID + IP Adapter FaceID on FLUX This converts the stylized look to photorealism while preserving the core identity features.