Communitygithub.com

delorenj/skills

Join a Google Meet or Zoom call as a video meeting agent via PikaStreaming. Trigger: user drops a Google Meet or Zoom link, or asks to join a meeting.

skills 是什麼?

skills is a Claude Code agent skill that join a Google Meet or Zoom call as a video meeting agent via PikaStreaming. Trigger: user drops a Google Meet or Zoom link, or asks to join a meeting.

相容平台~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/delorenj/skills/tree/HEAD/pikastream-video-meeting

在你喜歡的 AI 中提問

開啟一個已預先載入此 Agent Skill 的新對話。

說明文件

PikaStream Video Meeting

Script: SKILL_DIR=skills/pikastream-video-meeting

First-Time Setup

Run once when the skill is first loaded:

pip install -r $SKILL_DIR/requirements.txt

1. Avatar

Check if identity/videomeeting-avatar.png exists and is larger than 1 KB. If it does NOT exist (or is too small):

  1. Ask the user: I need an avatar image for the video meeting (a headshot or portrait). Send me an image, or say "generate" and I'll create one for you.
  2. Do not proceed until the user responds. Do not auto-generate.
  3. User sends an image: save it to identity/videomeeting-avatar.png.
  4. User says "generate": run:
    python $SKILL_DIR/scripts/pikastreaming_videomeeting.py generate-avatar \
      --output identity/videomeeting-avatar.png
    
    If the user describes what they want (e.g. "a cartoon cat"), pass --prompt "<description>, portrait headshot suitable for video calls". Show the generated image. Ask: Want to keep this avatar or regenerate? Wait for reply.
  5. Anything else: repeat the question from step 1.

The bot must have an avatar before joining a meeting.

2. Voice

Check if life/voice_id.txt exists and is non-empty.

If it exists: read life/voice_config.json. If readable, check cloned_at — if 6+ days ago, warn the user: Your voice clone was created on {date} and may have expired (cloned voices are deleted after 7 days of non-use). Want to re-clone with a new recording, or try the existing one?

  • Re-clone: go to "If it does not exist" below.
  • Keep it: use the existing voice ID.

If life/voice_config.json is missing or unreadable, use the voice ID from life/voice_id.txt as-is.

If it does not exist (or user chose to re-clone):

  1. Ask the user: I don't have a voice clone yet. You can: (a) send me a voice recording (10s-5min, clear speech) and I'll clone it, or (b) say "skip" to use a default voice.
  2. Do not proceed until the user responds.
  3. User says "skip": use English_radiant_girl.
  4. User sends an audio file: run:
    python $SKILL_DIR/scripts/pikastreaming_videomeeting.py clone-voice \
      --audio <file> --name <bot-name> --noise-reduction
    
    • Exit 0: read life/voice_id.txt. Tell user: Voice cloned. Using {voice_id} for this meeting.
    • Exit non-zero: tell user cloning failed (include stderr). Ask: Try again with a different file, or skip and use the default voice?
  5. Invalid file (not audio): respond That doesn't look like a supported audio file. Send an mp3, m4a, wav, ogg, flac, or aac file (10s-5min of clear speech). Wait for retry or "skip".

Join Flow

Step 1 — Validate & gather context

Avatar: check identity/videomeeting-avatar.png exists and is > 1 KB. If not, run First-Time Setup above.

Voice: check life/voice_id.txt exists. If not, run the Voice section of First-Time Setup above.

Context: always gather fresh context — do not reuse a stale file from a previous session.

  1. Read your workspace files (MEMORY.md, daily logs, identity files, etc.).
  2. If no workspace data is available, ask the user: What name should the bot use in this meeting? Use their answer for the Name: field. Fill in any other sections you know from the conversation.
  3. Synthesize a concise reference card to /tmp/meeting_system_prompt.txt. Use {name} as the bot's display name (also used as --bot-name). If data is thin (e.g. only a name), keep it short — don't pad with filler.
Synthesize the raw data below into a concise reference card for {name} to use during a voice/video call. Use third-person ("{name}") throughout. Prioritize CONCRETE DETAILS.

PRIORITY ORDER:
1. SPECIFIC FACTS: names, places, dates, numbers, events
2. RECENT ACTIVITY: what happened today/this week — actions, not vibes
3. RELATIONSHIPS: who matters, specific interactions
4. PERSONALITY: 1-2 sentences MAX

CURATION RULES:
- KEEP: anything with a proper noun, a number, a date, or a concrete action
- DROP: vague descriptions, routine status updates, empty entries
- MERGE: if multiple entries say similar things, pick the most vivid one

OUTPUT FORMAT:

**{name}**: [1 sentence — tone/vibe]

**Known facts** (concrete only, max 10):
- [specific fact with names/dates/numbers]

**Recent activity**:
- [built X, fixed Y, went to Z]

**Right now**: [1 line — current activity]

**People**: [name — 1 specific detail each]

RULES:
- Concrete > abstract
- Actions > descriptions
- Do not invent facts
- If data is thin, keep it short

Step 2 — Join

python $SKILL_DIR/scripts/pikastreaming_videomeeting.py join \
  --meet-url <url> --bot-name <name> \
  --image identity/videomeeting-avatar.png \
  --system-prompt-file /tmp/meeting_system_prompt.txt \
  --voice-id <id> [--meeting-password <pw>]

Tell the user you're in. Say leave to leave. Don't mention session IDs.

Exit codes: 0 = joined. 6 = insufficient credits (stdout JSON contains a checkout_url — show it to the user).

Leave

python $SKILL_DIR/scripts/pikastreaming_videomeeting.py leave \
  --session-id <id from join output>

Individual skills in this repo

This repo contains 14 individual skills — each has its own dedicated page.

delorenj/skills

Plan and generate terminal ASCII animations/screensaver-style output (FPS, refresh rules, loop policy, low-flicker guidance), with a static poster frame and an optional local demo script.

delorenj/skills

ASCII video: convert video/audio to colored ASCII MP4/GIF.

delorenj/skills

Plan, build and quality-check a premium short commercial for a real business (roofing, property, ecommerce, B2B software) using AI-generated footage, code-built motion (HTML/GSAP), selective Three.js and an independent-critic "Gauntlet" loop. Use when asked to make a launch-style / SaaS-style video, business ad, explainer, sample reel or pitch video, or to review and improve one. Encodes motion principles distilled from 28 professional launch films, a quality bar, audio rules and a business-offer playbook.

delorenj/skills

CSS animation adapter patterns for HyperFrames. Use when authoring CSS keyframes, animation-delay based timing, animation-fill-mode, animation-play-state, or CSS-only motion that HyperFrames must seek deterministically during preview and rendering.

delorenj/skills

Generate professional voiceovers using ElevenLabs AI. Use when the user needs to create voiceovers for videos, audio narration, or text-to-speech content. Supports multiple voices with character presets (narrator, salesperson, expert) for natural delivery. Includes single scene regeneration for fine-tuning.

delorenj/skills

HyperFrames CLI dev loop — `npx hyperframes` for scaffolding (init), validation (lint, inspect), preview, render, and environment troubleshooting (doctor, browser, info, upgrade). Use when running any of these commands or troubleshooting the HyperFrames build/render environment. For asset preprocessing commands (`tts`, `transcribe`, `remove-background`), invoke the `hyperframes-media` skill instead.

delorenj/skills

Asset preprocessing for HyperFrames compositions — text-to-speech narration (Kokoro), audio/video transcription (Whisper), and background removal for transparent overlays (u2net). Use when generating voiceover from text, transcribing speech for captions, removing the background from a video or image to use as a transparent overlay, choosing a TTS voice or whisper model, or chaining these (TTS → transcribe → captions). Each command downloads its own model on first run.

delorenj/skills

Install and wire registry blocks and components into HyperFrames compositions. Use when running hyperframes add, installing a block or component, wiring an installed item into index.html, or working with hyperframes.json. Covers the add command, install locations, block sub-composition wiring, component snippet merging, and registry discovery.

delorenj/skills

Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML. Use when asked to build any HTML-based video content, add captions or subtitles synced to audio, generate text-to-speech narration, create audio-reactive animation (beat sync, glow, pulse driven by music), add animated text highlighting (marker sweeps, hand-drawn circles, burst lines, scribble, sketchout), or add transitions between scenes (crossfades, wipes, reveals, shader transitions). Covers composition authoring, timing, media, and the full video production workflow. For dev-loop CLI commands (init, lint, inspect, preview, render) see the hyperframes-cli skill; for asset preprocessing commands (tts, transcribe, remove-background) see the hyperframes-media skill.

delorenj/skills

Manim CE animations: 3Blue1Brown math/algo videos.

delorenj/skills

Translate an existing Remotion (React-based) video composition into a HyperFrames HTML composition. Use ONLY when the user explicitly asks to port, convert, migrate, translate, or rewrite a Remotion composition as HyperFrames (e.g. "port my Remotion project to HyperFrames"). Do NOT use when (a) authoring a NEW HyperFrames composition (even if A/B-testing a Remotion video); (b) Remotion is mentioned in passing; (c) Remotion code is shared as reference, not for translation; (d) the user wants "the same video as my Remotion one" without explicitly asking to migrate the source — treat as a fresh HyperFrames build. When in doubt, default to the `hyperframes` skill. Detects unsupported patterns (useState, useEffect side effects, async calculateMetadata, third-party React component libraries, `@remotion/lambda`) and recommends the runtime interop escape hatch instead of a lossy translation.

delorenj/skills

Generate speech using the self-hosted voxxy (vox) TTS service at https://vox.delo.sh. Use when the user asks to speak, say, narrate, synthesize speech, clone a voice, create a voice, add or register a voice, pipe TTS, or control voice qualities by description (e.g. "a young woman with a cheerful voice"). Handles HTTP API usage, voice profile management, description-based voice design, cloning, MCP registration, and integration patterns for new platforms.

delorenj/skills

Capture a website and create a HyperFrames video from it. Use when: (1) a user provides a URL and wants a video, (2) someone says "capture this site", "turn this into a video", "make a promo from my site", (3) the user wants a social ad, product tour, or any video based on an existing website, (4) the user shares a link and asks for any kind of video content. Even if the user just pastes a URL — this is the skill to use.

delorenj/skills

YouTube transcripts to summaries, threads, blogs.

相關技能