bigdata-analysis-skill 是做什么的?
AI coding skill for Hive/Impala/Spark ETL — 10 rules to prevent silent data bugs on HDFS/YARN
适用于Hive/Impala/Spark ETL的AI编程技能——10条规则,防止HDFS/YARN上的静默数据错误。
Oak-B/bigdata-analysis-skill 是一款专注于大数据ETL开发的AI编程技能,特别针对Hive、Impala和Spark框架。该技能通过10条精心设计的规则,帮助开发者在HDFS和YARN环境下预防静默数据错误(silent data bugs),这些错误往往难以检测且可能导致严重的数据质量问题。技能覆盖了从查询编写、数据分区管理、资源优化到任务调度的全流程最佳实践,并提供了可执行的代码建议和优化策略。无论是数据工程师、数据分析师还是AI agent开发者,都能从中受益,尤其在构建稳定、高效的大数据管道时。该技能支持主流AI编程平台,如Claude Code、Cursor和Codex,用户可以通过简单的提示词激活这些规则来审查或生成更加健壮的ETL代码。通过遵循这些规则,团队可以显著减少因数据不一致或丢失导致的线上故障,提升数据管道的可靠性和可维护性。
npx skills add Oak-B/bigdata-analysis-skillInstalled? Explore more 研究与数据分析 skills: obra/superpowers, affaan-m/quarkus-verification, affaan-m/uspto-database · View all 6 →
AI coding skill for Hive/Impala/Spark ETL — 10 rules to prevent silent data bugs on HDFS/YARN
Agent skill repository discovered by 10x-chat research.
Verification loop for Quarkus projects: build, static analysis, tests with coverage, security scans, native compilation, and diff review before release or PR.
USPTO patent and trademark data workflow for official record lookup, PatentSearch queries, TSDR checks, assignment data, and reproducible IP research logs.
Structured scholarly-work evaluation for papers, proposals, literature reviews, methods sections, evidence quality, citation support, and research-writing feedback.
Systematic literature-review workflow for academic, biomedical, technical, and scientific topics, including search planning, source screening, synthesis, citation checks, and evidence logging.
Evidence-first current-state research workflow for ECC. Use when the user wants fresh facts, comparisons, enrichment, or a recommendation built from current public evidence and any supplied local context.