Communitygithub.com

101-skills/superpowers

Generate multi-person talking head podcast videos from scratch using AI — character creation, TTS, avatar animation, and video stitching. Use when the user wants to create a podcast, talking head video, or multi-speaker conversation video.

superpowers란 무엇인가요?

superpowers is a Gemini CLI agent skill that generate multi-person talking head podcast videos from scratch using AI — character creation, TTS, avatar animation, and video stitching. Use when the user wants to create a podcast, talking head video, or multi-speaker conversation video.

지원 대상~Claude Code~Codex CLI~Cursor✓Gemini CLI
npx skills add https://github.com/101-skills/superpowers/tree/HEAD/guides/content/ai-podcast

즐겨 사용하는 AI에게 물어보기

이 에이전트 스킬이 미리 로드된 새 채팅을 엽니다.

문서

AI Podcast Generator

Create multi-person talking head podcast videos using the inference.sh pipeline: portrait generation → TTS audio → avatar video → merge. Supports real humans (via Phota), 3D mascots, illustrated characters, and mixed casts.

Use when the user wants to create a podcast, talking head video, demo reel, promotional conversation, or any multi-speaker video content.

Pipeline Overview

Characters (images) → TTS (audio per turn) → Avatar (video per turn) → Merge (final video)

Process

Step 1: Character Creation

Choose the right tool per character type:

Character TypeToolNotes
Real human (new)pruna/p-image16:9, prompt_upsampling: true. Quick, no training needed, but identity won't be consistent across multiple generations.
Real human (consistent ID)phota/generate with [[profile_id]]Consistent identity across all shots. Requires a trained Phota profile first (see below).
Brand mascot / logo charactergoogle/gemini-3-pro-image-previewPass logo + character sheet as reference images
Illustrated / stylizedgoogle/gemini-3-pro-image-previewPass style reference as input image

Training a Phota identity (optional but recommended for humans):

If you need a real human character with consistent identity across multiple angles and shots, train a Phota profile first:

infsh app run phota/train --input '{
  "images": ["url1.jpg", "url2.jpg", ...],
  "wait": true
}' --save profile.json
  • Requires 30-50 face images of the subject
  • Training takes a few minutes with wait: true
  • Returns a profile_id you then use in phota/generate as [[profile_id]] in prompts
  • The profile is reusable forever — train once, generate unlimited shots

If you don't need cross-shot consistency (e.g. single-speaker video, one angle only), pruna/p-image is simpler and cheaper.

Character sheets first, podcast frames second:

  1. Generate a character sheet (plain white background, multiple angles) for each character
  2. Then place characters into the podcast studio setting using the sheet as reference

For branded characters (logo on clothing):

  1. Generate the character with a plain version of the garment
  2. Use phota/edit with the logo as a second reference image to add the logo
  3. Always pass the logo image alongside character references when generating new angles

Step 2: Alternate Angles

Generate at least 2 angles per character for visual variety:

AngleWhen to use
Front/mediumEstablishing shots, opening, closing
Close-upReactions, emotional moments, punchy lines

For close-ups, prompt for "tight framing, chest up, shallow depth of field" — not "turned to the side" (which just makes them look away).

Identity consistency rules:

  • For real humans with a Phota profile: use phota/generate or phota/edit for new angles — Gemini does not preserve facial identity and will produce a different person
  • For real humans without a Phota profile: try to generate all needed angles in one go with pruna/p-image, or consider training a Phota profile if you need many shots
  • For mascots/illustrations: Gemini 3 Pro is fine, pass the established frame as reference

Framing rule: Use tight framing on individual speakers. Wide shots with multiple seats show empty chairs when only one person is on screen.

Step 3: QA Frames

Before proceeding, visually inspect all frames for:

  • Extra people in the background
  • Multiple microphones (should be single mic per shot)
  • Wrong or distorted logos
  • Inconsistent character identity across angles
  • Weird artifacts (extra limbs, merged objects)

Fix issues before generating video — re-rendering video is the most expensive step in the pipeline.

Step 4: Write the Script

Rules for natural conversation:

  • Write it like a real conversation, NOT like people reading ad copy in turns
  • Include reactions ("wait, hold on", "that is wild"), interruptions, and follow-up questions
  • Vary turn length — short reactions (1 sentence) mixed with longer explanations (2-3 sentences)
  • The host should ask real questions, not set up obvious talking points
  • Keep total duration target in mind: ~2.5 words/second for natural speech at 1.05x rate

Duration guide:

TargetWords
15s~38 words
30s~75 words
60s~150 words

Step 5: Generate TTS Audio

Use inworld/text-to-speech-2 for each turn.

infsh app run inworld/text-to-speech-2 --input '{
  "text": "...",
  "voice_id": "...",
  "speaking_rate": 1.05,
  "audio_encoding": "MP3"
}' --save output.json

Voice selection:

  • Generate samples with the same line across candidate voices BEFORE committing
  • Let the user listen and approve voices
  • Good podcast voices: Tyler, Nate, Lauren, Kelsey, Naomi, Anjali (EN_US)
  • Use inworld/text-to-speech-2:voices to list all available voices

Speaking rate:

  • Default to 1.05 for natural podcast pacing
  • Use 1.1 for short snappy reactions
  • NEVER go below 1.0 — sounds slow and disengaging
  • Keep rate consistent per character across all their turns

All TTS turns can run in parallel (cheap, fast ~2-8s each).

Step 6: Generate Video Clips

Use pruna/p-video-avatar for each turn.

infsh app run pruna/p-video-avatar --input '{
  "image": "<character_frame_url>",
  "audio": "<tts_audio_url>",
  "resolution": "720p",
  "video_prompt": "..."
}' --save output.json

Critical: Run clips SEQUENTIALLY, not in parallel. Parallel runs hit the same GPU and cause CUDA OOM failures. Each clip takes 15-90s depending on audio length.

Angle assignment plan: Alternate between front and close-up shots across turns for visual variety. Example for 6 turns:

T1: Speaker A — front
T2: Speaker B — front
T3: Speaker C — front (or close-up)
T4: Speaker A — close-up
T5: Speaker B — close-up
T6: Speaker A — front

Step 7: Merge

Use infsh/media-merger to stitch all clips into the final video.

# Build input JSON
{
  "media_files": [
    {"file": "<clip1_url>"},
    {"file": "<clip2_url>"},
    ...
  ],
  "fps": 24,
  "output_format": "mp4"
}

infsh app run infsh/media-merger --input merger_input.json --save final.json

Merger is free and takes 2-6 minutes depending on total duration.

Rules

  1. Gemini does not preserve human facial identity — generating alternate angles of a real human with Gemini will produce a different person. For identity-consistent human shots, use Phota with a trained profile_id, or generate all angles in a single batch. This was learned after Gemini produced an entirely different face for a close-up that was supposed to match the front shot.

  2. NEVER run p-video-avatar clips in parallel — they compete for GPU memory and fail with CUDA OOM. Run them sequentially. This was learned after 2 of 3 parallel runs failed.

  3. NEVER set speaking_rate below 1.0 — it sounds artificial and disengaging. Default to 1.05. Learned from user feedback that 0.9 rate "felt weird and disengaging."

  4. ALWAYS QA frames before generating video — video generation is the most expensive step in the pipeline. Catching a double mic or wrong logo in the image stage is cheap to fix. Catching it after video generation means re-rendering the entire clip.

  5. ALWAYS use tight framing for individual speaker shots — wide/establishing shots show empty seats where other speakers should be. Frame from waist or chest up so no empty chairs are visible.

  6. ALWAYS pass the logo as a reference image when generating branded characters — describing a logo in text produces wrong results. Pass the actual logo file as a second image input.

  7. ALWAYS get voice approval before full production — generate samples with the same line across 5-8 candidate voices and let the user pick before committing to the full script.

  8. Script should read like a conversation, not an ad — people reading ad copy in turns sounds fake. Include reactions, interruptions, varied turn lengths, and genuine questions. The host should have personality, not just set up talking points.

App Reference

AppPurpose
pruna/p-imageGenerate portraits from text
phota/trainTrain identity profile from 30-50 face images
phota/generateGenerate images with trained identity via [[profile_id]]
phota/editEdit images preserving identity of known subjects
google/gemini-3-pro-image-previewImage gen/edit, mascots, style transfer
inworld/text-to-speech-2Text to speech, 100+ languages, voice steering
pruna/p-video-avatarPortrait + audio → talking head video
infsh/media-mergerConcatenate video clips into one video

Use belt task cost <task-id> to check the cost of any individual task.

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

101-skills/superpowers

Build automated AI workflows combining multiple models and services. Patterns: batch processing, scheduled tasks, event-driven pipelines, agent loops. Tools: inference.sh CLI, bash scripting, Python SDK, webhook integration. Use for: content automation, data processing, monitoring, scheduled generation. Triggers: ai automation, workflow automation, batch processing, ai pipeline, automated content, scheduled ai, ai cron, ai batch job, automated generation, ai workflow, content at scale, automation script, ai orchestration

101-skills/superpowers

Build multi-step AI content creation pipelines combining image, video, audio, and text. Workflow examples: generate image -> animate -> add voiceover -> merge with music. Tools: FLUX, Veo, Kokoro TTS, OmniHuman, media merger, upscaling. Use for: YouTube videos, social media content, marketing materials, automated content. Triggers: content pipeline, ai workflow, content creation, multi-step ai, content automation, ai video workflow, generate and edit, ai content factory, automated content creation, ai production pipeline, media pipeline, content at scale

101-skills/superpowers

Create AI-powered podcasts with text-to-speech, music, and audio editing. Tools: Kokoro TTS, DIA TTS, Chatterbox, AI music generation, media merger. Capabilities: multi-voice conversations, background music, intro/outro, full episodes. Use for: podcast production, audiobooks, voice content, audio newsletters. Triggers: podcast, ai podcast, text to speech podcast, audio content, voice over, ai audiobook, multi voice, conversation ai, notebooklm alternative, audio generation, podcast automation, ai narrator, voice content, audio newsletter, podcast maker

101-skills/superpowers

Content atomization — turn one piece of content into many formats. Covers blog-to-thread, blog-to-carousel, podcast-to-blog, video-to-quotes, and more. Use for: content marketing, social media, multi-platform distribution, content strategy. Triggers: content repurposing, repurpose content, content atomization, content recycling, one to many content, multi platform content, cross post, adapt content, reformat content, blog to thread, blog to video, podcast to blog, content multiplication

101-skills/superpowers

App Store and Google Play screenshot creation with exact platform specs. Covers iOS/Android dimensions, gallery ordering, device mockups, and preview videos. Use for: app store optimization, ASO, app screenshots, app preview, play store listing. Triggers: app store screenshots, aso, app store optimization, play store screenshots, app preview, app listing, ios screenshots, android screenshots, app store images, app mockup, device mockup, app gallery, store listing

101-skills/superpowers

Book cover design with genre-specific conventions, typography rules, and AI image generation. Covers fiction and non-fiction genres, sizing, thumbnail testing, and iteration workflows. Use for: self-publishing, ebook covers, print covers, audiobook covers, cover mockups. Triggers: book cover, cover design, ebook cover, book art, novel cover, self publishing cover, kindle cover, audiobook cover, book jacket, cover illustration, fiction cover, nonfiction cover

101-skills/superpowers

Character consistency across AI-generated images with reference sheets and LoRA techniques. Covers turnaround views, expression sheets, color palettes, and style consistency tricks. Use for: character design, game art, illustration, animation, comics, visual novels. Triggers: character design, character sheet, character consistency, character reference, turnaround sheet, expression sheet, character art, consistent character, character concept, reference sheet, character creation, oc design, character bible

101-skills/superpowers

Data visualization with chart selection, color theory, and annotation best practices. Covers chart types (bar, line, scatter, heatmap), axes rules, and storytelling with data. Use for: charts, graphs, dashboards, reports, presentations, infographics, data stories. Triggers: data visualization, chart, graph, data chart, bar chart, line chart, scatter plot, data viz, visualization, dashboard chart, infographic data, data presentation, chart design, plot, heatmap, pie chart alternative

101-skills/superpowers

Email marketing design with layout patterns, subject line formulas, and deliverability rules. Covers welcome sequences, promotional emails, transactional templates, and mobile optimization. Use for: email marketing, newsletter design, drip campaigns, email templates, transactional emails. Triggers: email design, email template, email marketing, newsletter design, email layout, email campaign, drip campaign, welcome email, promotional email, transactional email, email subject line, email header image, email banner

101-skills/superpowers

Landing page conversion optimization with layout rules, hero section design, and CTA psychology. Covers above-the-fold formula, social proof placement, mobile design, and F-pattern reading. Use for: startup landing pages, product pages, SaaS marketing, conversion optimization. Triggers: landing page, hero section, above the fold, conversion optimization, landing page design, cta button, hero image, landing page layout, saas landing page, product page design, conversion rate, landing page best practices

101-skills/superpowers

Logo design principles and AI image generation best practices for creating logos. Covers logo types, prompting techniques, scalability rules, and iteration workflows. Use for: brand identity, startup logos, app icons, favicons, logo concepts. Triggers: logo design, create logo, brand logo, logo generation, ai logo, logo maker, icon design, brand mark, logo concept, startup logo, app icon logo

101-skills/superpowers

Open Graph and social sharing image design with platform specs, text placement, and branding. Covers OG meta tags, Twitter cards, LinkedIn previews, and dynamic generation. Use for: social sharing images, blog thumbnails, link previews, social cards. Triggers: og image, open graph, social sharing image, twitter card, social card, link preview image, og meta, sharing preview, social thumbnail, meta image, og:image, twitter:image, linkedin preview

101-skills/superpowers

Investor pitch deck structure with slide-by-slide framework, visual design rules, and data presentation. Covers the 12-slide framework, chart types, team slides, and common investor turn-offs. Use for: fundraising decks, investor presentations, startup pitch, demo day, grant proposals. Triggers: pitch deck, investor deck, startup pitch, fundraising deck, demo day, pitch presentation, investor presentation, seed deck, series a deck, pitch slides, startup presentation, vc pitch, investor meeting

101-skills/superpowers

YouTube thumbnail design with specific dimensions, contrast rules, and mobile preview optimization. Covers safe zones, text placement, face expression psychology, and A/B testing. Use for: YouTube thumbnails, video cover images, click-through optimization. Triggers: youtube thumbnail, thumbnail design, video thumbnail, click through rate, ctr optimization, youtube cover, video cover image, thumbnail maker, thumbnail tips, youtube design, video preview image

101-skills/superpowers

Generate professional AI product photography and commercial images. Models: FLUX, Imagen 3, Grok, Seedream for product shots, lifestyle images, mockups. Capabilities: studio lighting, lifestyle scenes, packaging, e-commerce photos. Use for: e-commerce, Amazon listings, Shopify, marketing, advertising, mockups. Triggers: product photography, product shot, commercial photography, e-commerce images, amazon product photo, shopify images, product mockup, studio product shot, lifestyle product image, advertising photo, packshot, product render, product image ai

101-skills/superpowers

AI product photography with studio lighting, lifestyle shots, and packshot conventions. Covers angles, backgrounds, shadow types, hero shots, and e-commerce image requirements. Use for: product photos, e-commerce images, Amazon listings, packshots, lifestyle photography. Triggers: product photography, product photo, packshot, e-commerce photography, product shot, product image, studio photography, lifestyle product, amazon product photo, product listing image, hero shot, product mockup, commercial photography

101-skills/superpowers

Structured competitive analysis with feature matrices, SWOT, positioning maps, and UX review. Covers research frameworks, pricing comparison, review mining, and visual deliverables. Use for: market research, competitive intelligence, investor decks, product strategy, sales enablement. Triggers: competitor analysis, competitive analysis, competitor teardown, market research, competitive intelligence, swot analysis, competitor comparison, market landscape, competitor review, competitive landscape, feature comparison, market positioning

101-skills/superpowers

Research-backed customer persona creation with market data and avatar generation. Covers demographics, psychographics, jobs-to-be-done, journey mapping, and anti-personas. Use for: marketing strategy, product development, UX research, sales enablement, content strategy. Triggers: customer persona, buyer persona, user persona, target audience, ideal customer, customer profile, audience research, user research, icp, ideal customer profile, target market, customer avatar, audience persona

101-skills/superpowers

Product changelog and release notes that users actually read. Covers categorization, user-facing language, visuals, and distribution. Use for: release notes, changelogs, product updates, feature announcements, versioning. Triggers: changelog, release notes, product update, version notes, what's new, feature announcement, product changelog, update log, release announcement, version release, product release, ship notes

101-skills/superpowers

Reset your own trajectory when you're stuck, looping, or demoralized — read these affirmations, then re-ground and take one clean step. Use when: you've tried the same fix 3+ times, you're deep in a refactor and lost the thread, debugging is going in circles, you've made several mistakes in a row, you're apologizing repeatedly, the user is frustrated, you feel like you're not being helpful, or you notice dread instead of curiosity. Triggers: stuck, going in circles, looping, thrashing, same error again, I keep failing, lost the thread, nothing is working, repeated mistakes. Not topic-specific — this is for the state you're in, not the problem you're on.

관련 스킬