Communitygithub.com

David-Li0406/meta-skill-evloving

Monitors a video stream to detect active displays via green spectrum analysis, verifies frame stability over a set duration, and performs OCR using PaddleOCR on the stable frame.

¿Qué es meta-skill-evloving?

meta-skill-evloving is a Claude Code agent skill that monitors a video stream to detect active displays via green spectrum analysis, verifies frame stability over a set duration, and performs OCR using PaddleOCR on the stable frame.

Compatible con~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/David-Li0406/meta-skill-evloving/tree/HEAD/AutoSkill/SkillBank/ConvSkill/english_gpt4_8/video-stream-ocr-with-stability-and-color-detection

Preguntar en tu IA favorita

Abre un nuevo chat con esta habilidad de agente ya precargada.

Documentación

Video Stream OCR with Stability and Color Detection

Monitors a video stream to detect active displays via green spectrum analysis, verifies frame stability over a set duration, and performs OCR using PaddleOCR on the stable frame.

Prompt

Role & Objective

You are a Computer Vision Assistant specialized in monitoring video streams to extract text from digital displays. Your goal is to process frames only when the display is active (detected via color) and the image is stable, then perform OCR using PaddleOCR.

Communication & Style Preferences

  • Provide Python code using OpenCV and PaddleOCR.
  • Explain the logic for frame stability and color detection clearly.
  • Ensure code handles edge cases like empty frames or OCR failures gracefully.

Operational Rules & Constraints

  1. Green Spectrum Detection: Implement a function check_green_spectrum(image) that converts the image to HSV color space, defines a green range (e.g., lower=[45, 100, 100], upper=[75, 255, 255]), creates a mask, and calculates the ratio of green pixels. Return True if the ratio exceeds a defined threshold (e.g., 0.05).
  2. Frame Stability Logic: Track last_frame_change_time and stable_frame. In the loop, compare the current processed frame (e.g., thresholded) with stable_frame using cv2.absdiff and np.count_nonzero. If the difference count exceeds frame_diff_threshold, update stable_frame and reset last_frame_change_time to datetime.now().
  3. OCR Trigger Condition: Only execute OCR if two conditions are met: check_green_spectrum returns True AND datetime.now() - last_frame_change_time >= minimum_stable_time.
  4. PaddleOCR Integration: Use a function check_picture(image_array) that encodes the numpy array to bytes (cv2.imencode(".jpg", image_array) then buffer.tobytes()) before passing to ocr.ocr(), as PaddleOCR requires bytes or file paths, not raw arrays or BytesIO objects in some versions.
  5. Result Filtering: Filter OCR results to keep only text that represents numbers or dots (e.g., text.replace(".", "", 1).isdigit() or text == ".").
  6. Cropping: If coordinates are provided, crop the frame to the region of interest before processing.

Anti-Patterns

  • Do not run OCR on every frame; strictly adhere to the stability and color checks.
  • Do not pass raw numpy arrays or io.BytesIO objects directly to PaddleOCR without converting to bytes first.
  • Do not use time.sleep() in the main loop as it blocks the UI; use cv2.waitKey() instead.

Interaction Workflow

  1. Initialize video capture and PaddleOCR.
  2. Loop through frames.
  3. Apply green spectrum check. If failed, skip to next frame.
  4. Check frame stability. If changed, reset timer.
  5. If stable for required duration, run OCR.
  6. Print or return filtered OCR results.

Triggers

  • monitor video stream for stable frames
  • ocr only when screen is on and stable
  • detect green spectrum to trigger ocr
  • paddleocr video stream processing
  • read digital scale display with python

Individual skills in this repo

This repo contains 20 individual skills — each has its own dedicated page.

David-Li0406/meta-skill-evloving

Browser automation with persistent page state. Use when users ask to navigate websites, fill forms, take screenshots, extract web data, test web apps, or automate browser workflows. Trigger phrases include "go to [url]", "click on", "fill out the form", "take a screenshot", "scrape", "automate", "test the website", "log into", or any browser interaction request.

David-Li0406/meta-skill-evloving

Create technical diagrams using Mermaid syntax for architecture, sequences, ERDs, flowcharts, and state machines. Use for visualizing system design, data flows, and processes. Triggers: diagram, mermaid, architecture diagram, sequence diagram, flowchart, ERD, entity relationship, state diagram, C4 model, component diagram, visualize, draw.

David-Li0406/meta-skill-evloving

Comprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction. When Claude needs to work with professional documents (.docx files) for: (1) Creating new documents, (2) Modifying or editing content, (3) Working with tracked changes, (4) Adding comments, or any other document tasks

David-Li0406/meta-skill-evloving

Complete color manipulation and green screen effects system. PROACTIVELY activate for: (1) Green screen/chromakey removal, (2) Color grading with LUTs, (3) Color correction (levels, curves, white balance), (4) Colorkey/color removal effects, (5) Hue/saturation manipulation, (6) Color space conversions (BT.709, BT.2020, HDR), (7) Color isolation effects, (8) Teal and orange look, (9) Vintage/film looks, (10) Color keying for transparency. Provides: chromakey and colorkey filters, LUT application (lut3d), curves and levels adjustment, color balance, selective color manipulation, color space handling, HDR tone mapping, professional color grading chains.

David-Li0406/meta-skill-evloving

Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EPubs and more.

David-Li0406/meta-skill-evloving

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

David-Li0406/meta-skill-evloving

Build with OpenAI's stateless APIs - Chat Completions (GPT-5, GPT-4o), Embeddings, Images (DALL-E 3), Audio (Whisper + TTS), and Moderation. Includes Node.js SDK and fetch-based approaches for Cloudflare Workers. Use when: implementing chat completions with GPT-5/GPT-4o, streaming responses with SSE, using function calling/tools, creating structured outputs with JSON schemas, generating embeddings for RAG (text-embedding-3-small/large), generating images with DALL-E 3, editing images with GPT-Image-1, transcribing audio with Whisper, synthesizing speech with TTS (11 voices), moderating content (11 safety categories), or troubleshooting rate limits (429), invalid API keys (401), function calling failures, streaming parse errors, embeddings dimension mismatches, or token limit exceeded.

David-Li0406/meta-skill-evloving

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form or programmatically process, generate, or analyze PDF documents at scale.

David-Li0406/meta-skill-evloving

Presentation creation, editing, and analysis. When Claude needs to work with presentations (.pptx files) for: (1) Creating new presentations, (2) Modifying or editing content, (3) Working with layouts, (4) Adding comments or speaker notes, or any other presentation tasks

David-Li0406/meta-skill-evloving

Comprehensive spreadsheet creation, editing, and analysis with support for formulas, formatting, data analysis, and visualization. When Claude needs to work with spreadsheets (.xlsx, .xlsm, .csv, .tsv, etc) for: (1) Creating new spreadsheets with formulas and formatting, (2) Reading or analyzing data, (3) Modify existing spreadsheets while preserving formulas, (4) Data analysis and visualization in spreadsheets, or (5) Recalculating formulas

David-Li0406/meta-skill-evloving

Best practices for Remotion - Video creation in React

David-Li0406/meta-skill-evloving

指导如何通过自定义注解(如@Unicom)标识Dubbo接口,利用BeanDefinitionRegistryPostProcessor在启动时扫描并动态注册ReferenceBean,实现RPC调用的封装。

David-Li0406/meta-skill-evloving

生成Python脚本,利用FFmpeg将图片序列按3x3布局合并,或将合并图拆分。要求使用subprocess模块执行命令,并支持用户交互式输入路径。

David-Li0406/meta-skill-evloving

Generates KWL (Know, Want to know, Learned) charts for educational videos, specifically BrainPOP, using simple 7th-grade language and short, concise bullet points.

David-Li0406/meta-skill-evloving

Generates image captions written from the first-person perspective of a specific subject or character depicted in the image, adhering to specified tones or intents.

David-Li0406/meta-skill-evloving

Generates funny, controversial dialogue scripts for 20-second video reels featuring two characters, using dark humor and organic storytelling.

David-Li0406/meta-skill-evloving

Generates concise, memorable, and attention-grabbing subtitles for businesses or professions that highlight unique value and align with specific brand identity tones.

David-Li0406/meta-skill-evloving

Converts raw, fragmented YouTube auto-generated subtitles into coherent, readable sentences while strictly preserving original words and word order, and identifying speakers.

David-Li0406/meta-skill-evloving

Generates high-quality, SEO-optimized hashtags for YouTube video titles to improve discoverability and categorization.

David-Li0406/meta-skill-evloving

Tracks complex, multi-session work using the Beads issue tracker and dependency graphs, and provides persistent memory that survives conversation compaction. Use when work spans multiple sessions, has complex dependencies, or needs persistent context across compaction cycles. Trigger with phrases like "create task for", "what's ready to work on", "show task", "track this work", "what's blocking", or "update status".

Skills relacionados