Communitygithub.com

G1Joshi/Agent-Skills

Expert Dask distributed computing assistance covering Dask DataFrames, Arrays, Futures, and cluster scaling. Use when analyzing datasets too large for pandas on a single machine or multi-node cluster.

Agent-Skills とは?

Agent-Skills is a Claude Code agent skill that expert Dask distributed computing assistance covering Dask DataFrames, Arrays, Futures, and cluster scaling. Use when analyzing datasets too large for pandas on a single machine or multi-node cluster.

対応~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/G1Joshi/Agent-Skills/tree/HEAD/skills/ai-ml/dask

お気に入りのAIに質問する

このエージェントスキルを事前に読み込んだ状態で新しいチャットを開きます。

ドキュメント

Dask

Dask scales Python. It looks like Pandas/NumPy but runs on clusters. 2025 updates focus on High Performance Shuffle and GPU integration.

When to Use

  • Larger-than-Memory Tabular Computation: Processing 10GB-1TB CSV/Parquet datasets on a single machine or multi-node cluster.
  • Parallelizing Custom Python Workflows: Using dask.delayed to parallelize arbitrary Python loops and task graphs.
  • Distributed Machine Learning: Integrating with Scikit-learn, XGBoost, and LightGBM across distributed worker nodes.
  • NumPy & Pandas Scaling: Drop-in familiar APIs for out-of-core array and DataFrame calculations.

Quick Start

import dask.dataframe as dd

# Read multi-gigabyte partitioned CSV files lazily
df = dd.read_csv('data/transactions_*.csv')

# Compute aggregated statistics in parallel across all CPU cores
result = df.groupby('category').amount.mean().compute()
print(result)

Core Concepts

#Out-of-Core Processing with Dask DataFrame

Loading partitioned Parquet files and computing aggregations lazily:

import dask.dataframe as dd
from dask.distributed import Client

# Initialize local distributed cluster
client = Client(n_workers=4, threads_per_worker=2, memory_limit='4GB')
print(f"Dask Dashboard running at: {client.dashboard_link}")

# Lazy load multi-file dataset
df = dd.read_parquet(
    's3://analytics-bucket/transactions/*.parquet',
    columns=['customer_id', 'amount', 'status', 'country'],
    engine='pyarrow'
)

# Build computational graph lazily
filtered = df[df['status'] == 'completed']
metrics = filtered.groupby('country')['amount'].agg(['mean', 'sum', 'count'])

# Execute graph and return concrete Pandas DataFrame
result_df = metrics.compute()
print(result_df)

#Custom Task Graphs with dask.delayed

Parallelizing independent function executions:

from dask import delayed, compute
import time

@delayed
def fetch_api_data(endpoint: str) -> dict:
    time.sleep(0.5) # Simulating I/O
    return {'endpoint': endpoint, 'status': 200, 'items': [1, 2, 3]}

@delayed
def transform_payload(data: dict) -> int:
    return sum(data['items'])

@delayed
def aggregate_totals(totals: list[int]) -> int:
    return sum(totals)

# Construct dependency DAG
endpoints = ['users', 'orders', 'inventory', 'payments']
fetched = [fetch_api_data(ep) for ep in endpoints]
transformed = [transform_payload(f) for f in fetched]
final_total = aggregate_totals(transformed)

# Execute all tasks in parallel across the cluster
total_result = final_total.compute()
print("Grand Total:", total_result)

#Dask Array for Distributed Linear Algebra

Manipulating massive multi-dimensional arrays:

import dask.array as da

# Create a 50,000 x 50,000 array split into 5,000 x 5,000 chunks
x = da.random.normal(10, 0.1, size=(50000, 50000), chunks=(5000, 5000))

# Perform calculations lazily
y = x + x.T
z = y[::2, ::2].mean(axis=0)

# Compute result
result = z.compute()
print("Computed array mean shape:", result.shape)

Common Patterns

Distributed Futures for Irregular Parallel Task Graphs

Problem: Custom parallel computations that don't fit neatly into DataFrame/Array row/column structures.

Solution: Use the Dask Distributed Client with asynchronous futures:

from dask.distributed import Client

client = Client(n_workers=4, threads_per_worker=2)

def train_partition(part_id: int):
    # Train local sub-model
    return f"Model {part_id} trained"

futures = [client.submit(train_partition, i) for i in range(10)]
results = client.gather(futures)
print(results)

Best Practices (2026)

  • Do check the Dask Web Dashboard (typically on port 8787) to monitor memory pressure, task streams, and bottlenecks.
  • Do choose chunk sizes between 100MB and 300MB in memory for optimal parallel efficiency.
  • Do persist intermediate DataFrames (df = df.persist()) when querying the same transformed data repeatedly.
  • Do filter columns and rows early using column projections and predicate pushdown.
  • Don't call .compute() inside loops; build the complete task graph and call compute() once.
  • Don't use Dask if data fits comfortably in RAM; native Pandas and Polars are significantly faster for in-memory tasks.
  • Don't create millions of tiny delayed tasks; excessive task overhead degrades scheduler performance.

Troubleshooting

ErrorCauseSolution
KilledWorker: Worker exceeded 95% memory budgetTask partition too large to fit in worker RAM during compute.Increase number of partitions: df.repartition(npartitions=100).
UserWarning: Sending large object to workersLarge NumPy array or DataFrame passed as direct function argument.Scatter data once with client.scatter(data) before submitting tasks.
Slow compute: Too many tiny tasksThousands of micro-partitions causing task scheduling overhead.Merge small partitions using repartition(partition_size="100MB").

References

Individual skills in this repo

This repo contains 11 individual skills — each has its own dedicated page.

G1Joshi/Agent-Skills

Expert dbt (data build tool) assistance covering SQL modeling, Jinja macros, tests, documentation, and semantic layer. Use when building analytics engineering pipelines on BigQuery, Snowflake, or PostgreSQL.

G1Joshi/Agent-Skills

Expert JAX assistance covering Autograd, XLA compilation (`jit`), vectorization (`vmap`), and parallelization (`pmap`). Use when building high-performance numerical computing and cutting-edge deep learning research.

G1Joshi/Agent-Skills

Expert Ray distributed computing assistance covering Ray Core (actors, tasks), Ray Train, Ray Tune, and Ray Serve. Use when scaling Python compute and ML training across multi-node clusters.

G1Joshi/Agent-Skills

Expert Git version control assistance covering branching, interactive rebase, cherry-pick, submodules, worktrees, and conflict resolution. Use when managing source code history and collaboration workflows.

G1Joshi/Agent-Skills

Expert K9s CLI assistance covering terminal Kubernetes cluster navigation, real-time log streaming, port forwarding, and pod debugging. Use when managing and troubleshooting Kubernetes clusters with speed.

G1Joshi/Agent-Skills

Expert SWC assistance covering high-performance Rust-based JavaScript/TypeScript compilation, .swcrc configuration, Jest testing via @swc/jest, and minification. Use when replacing Babel with SWC for faster builds, accelerating test execution, or compiling modern ECMAScript features.

G1Joshi/Agent-Skills

Expert Tig assistance covering text-mode interface for Git, interactive staging, commit graph visualization, diff exploration, and blame navigation. Use when navigating Git commit history in the terminal, staging hunks interactively, browsing file changes, or reviewing revisions.

G1Joshi/Agent-Skills

Expert Vim assistance covering modal editing, .vimrc configuration, registers, macros, search/replace, and plugin management via vim-plug. Use when editing text efficiently in terminal environments, writing Vimscript, recording macros, or configuring core Vim settings.

G1Joshi/Agent-Skills

Expert Zed editor assistance covering high-performance Rust-based text editing, multi-buffer editing, language server protocols (LSP), and AI assistant integrations. Use when configuring Zed settings.json, setting up language extensions, collaborating in real-time channels, or optimizing editor startup speed.

G1Joshi/Agent-Skills

Expert Zsh assistance covering shell customization, Oh My Zsh plugins, Zinit plugin manager, prompt engineering (Starship/Powerlevel10k), and shell scripting. Use when configuring .zshrc, writing Zsh automation scripts, optimizing shell startup time, or configuring tab-completion.

G1Joshi/Agent-Skills

Expert [skill-name] assistance covering [feature 1], [feature 2], and [feature 3]. Use when [working with X], [debugging Y], or [implementing Z].

関連スキル