AI Character Consistency: 4 Emotions in One Image Prompt
Discover a four-panel AI image prompt designed to preserve character identity across different emotions, poses, and camera angles. Learn how identity anchoring and precise spatial prompting can help create consistent visual content for social media campaigns.
1 October 20263 min read

Can your brand's AI persona genuinely laugh, scowl, and celebrate without turning into a completely different person in every frame? For social media managers and creative directors, character consistency across diverse emotional states has long been the holy grail—and one of the most frustrating bottlenecks—in generative visual marketing.
Prompt of the Week
"Using the uploaded reference image, generate a four-panel composition (upper-left, upper-right, lower-left, lower-right).
Keep the person’s face, hairstyle, and clothing exactly the same as in the uploaded image for all panels. Maintain a realistic photo style, consistent identity, and natural lighting.
Upper-left (Joy / 喜び): A close-up of her face from a slightly high angle. She is smiling brightly with sparkling eyes, expressing pure joy and happiness. Her expression feels open, warm, and genuinely delighted.
Upper-right (Anger / 怒り): A medium shot of her upper body. She is visibly angry, with furrowed brows and a sharp, intense expression. Her posture is tense, and she slightly clenches her fists or crosses her arms, clearly showing irritation or frustration.
Lower-left (Sadness / 哀しみ): An over-the-shoulder or slightly downward-angled shot. Her body language is subdued and closed off. She looks down with a sorrowful, quiet expression, conveying loneliness or emotional pain, with a soft, muted mood.
Lower-right (Fun / 楽しい): A casual selfie-style shot. She looks playful and energetic, laughing or making a cheerful, fun expression. The mood is lighthearted and carefree, showing her enjoying the moment."
Category: Social Media Content
Difficulty: ⭐⭐⭐⭐⭐ (Advanced)
Recommended Model: Nano Banana Pro
How This Prompt Works
Generating multiple variations in a single generation requires precise spatial prompting. By instructing the model to construct a defined 2x2 grid (upper-left, upper-right, lower-left, lower-right), you bypass the challenge of running multiple individual generations and hoping the identity aligns afterward.
The prompt relies on three core mechanisms: strict identity anchoring (locking hair, face, and clothing), varied camera framing (high-angle close-up, medium shot, over-the-shoulder, selfie), and multisensory emotional cues. Instead of merely saying "angry" or "happy," the prompt dictates posture, eye details, and body language (e.g., clenched fists, subdued stance), which helps advanced image models like Nano Banana Pro render authentic, high-fidelity micro-expressions.
Marketing Use Case: The Virtual Brand Ambassador
Consider a D2C beauty brand rolling out an interactive Instagram and TikTok campaign titled "Four Moods of Monday." In traditional production, capturing an influencer across four wildly different emotional beats with coordinated lighting and framing requires hours in a studio, multiple wardrobe adjustments, and costly retakes.
With this workflow, marketing teams can generate a four-panel character sheet to anchor interactive story polls, carousel posts, or meme-style relatable reactions (e.g., "When your order ships vs. when it says delivered but isn't there"). The result? Ultra-relatable, campaign-ready social assets delivered in minutes, preserving the face of the brand with zero continuity drift.
Variations and Customization
Adapt this multi-panel approach across various industries by tweaking the framing and emotional triggers:
- SaaS / B2B Tech (User Journey Grid): Swap the emotional states to map user onboarding: Upper-left: Overwhelmed by spreadsheets (Confusion); Upper-right: Discovering your tool (Curiosity); Lower-left: Setting up automations (Deep Focus); Lower-right: Seeing the monthly analytics ROI (Triumphant Relief).
- Fashion & Retail (Style Moodboard): Keep the model's identity consistent while varying lighting and attitude across settings: Upper-left: Clean morning commuter; Upper-right: High-fashion boardroom confidence; Lower-left: Moody evening lounge; Lower-right: Energetic weekend casual.
- Fitness & Wellness (Workout Lifecycle): Track an athlete's physical and mental arc: Upper-left: Pre-workout focus and determination; Upper-right: Mid-workout grit and exertion; Lower-left: Exhausted cool-down; Lower-right: Post-workout endorphin rush with a branded shaker cup.
Tips & Best Practices
- Vary camera angles deliberately: Don't leave framing to chance. Pairing specific angles (close-ups vs. medium shots) with specific emotions prevents the grid from looking like a repetitive passport photo sheet.
- Anchor stable variables first: Explicitly state what must not change (lighting tone, wardrobe, hairstyle) before defining individual quadrant variations.
- Crop into carousels: Once generated, split the 4-panel image into individual 1:1 or 4:5 crops for sequential Instagram Carousel slides or dynamic video cutaways.
Automate Your Visual Pipeline with VIVID
Mastering multi-expression character generation opens up entirely new avenues for consistent, narrative-driven visual marketing. If you want to streamline this workflow without spending hours tweaking prompts, explore how VIVID helps agencies and marketing teams automate on-brand image generation at scale.
Try it with
your product
Sign up and verify your email: you get {{credits}} credits to try it on your product, no card needed.
Start for free →