Communitygithub.com

G1Joshi/Agent-Skills

Expert Ray distributed computing assistance covering Ray Core (actors, tasks), Ray Train, Ray Tune, and Ray Serve. Use when scaling Python compute and ML training across multi-node clusters.

Agent-Skills とは?

Agent-Skills is a Claude Code agent skill that expert Ray distributed computing assistance covering Ray Core (actors, tasks), Ray Train, Ray Tune, and Ray Serve. Use when scaling Python compute and ML training across multi-node clusters.

対応~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/G1Joshi/Agent-Skills/tree/HEAD/skills/ai-ml/ray

お気に入りのAIに質問する

このエージェントスキルを事前に読み込んだ状態で新しいチャットを開きます。

ドキュメント

Ray

Ray is the compute layer for AI. It powers ChatGPT training and massive scale workloads. v3.0 (2025) improves efficiency and adds an MCP Server for agents.

When to Use

  • Distributed Python Applications: Scaling compute-heavy Python tasks seamlessly from a single laptop to large cloud clusters.
  • Distributed ML Training with Ray Train: Orchestrating multi-node PyTorch, XGBoost, and LightGBM training jobs.
  • Production Model Serving with Ray Serve: Serving LLMs and multi-model microservice pipelines with dynamic autoscaling.
  • Hyperparameter Optimization with Ray Tune: Running Bayesian, Optuna, and population-based training sweeps across hundreds of workers.

Quick Start

import ray

ray.init()

# Remote distributed task
@ray.remote
def square(x):
    return x * x

# Launch tasks in parallel across cluster
futures = [square.remote(i) for i in range(10)]
results = ray.get(futures)
print("Computed squares:", results)

Core Concepts

#Distributed Tasks & Stateful Actors

Scaling functions and classes across cluster workers:

import ray
import time

# Initialize local or connect to remote Ray cluster
ray.init(ignore_reinit_error=True)

# 1. Stateless Remote Function (Distributed Task)
@ray.remote
def process_data_chunk(chunk_id: int) -> dict:
    time.sleep(0.5) # Simulating CPU/IO bound computation
    return {"chunk_id": chunk_id, "processed_items": 1000 * chunk_id}

# Launch tasks concurrently across cluster
futures = [process_data_chunk.remote(i) for i in range(8)]
results = ray.get(futures) # Gather results
print("Processed Chunks:", results)

# 2. Stateful Remote Class (Distributed Actor)
@ray.remote
class GlobalCounter:
    def __init__(self):
        self.count = 0

    def increment(self) -> int:
        self.count += 1
        return self.count

counter_actor = GlobalCounter.remote()
inc_futures = [counter_actor.increment.remote() for _ in range(5)]
print("Final actor count:", ray.get(inc_futures[-1]))

#Production Model Serving with Ray Serve

Deploying scalable REST endpoints with dynamic request batching:

from ray import serve
import torch
from transformers import pipeline

@serve.deployment(num_replicas=2, ray_actor_options={"num_cpus": 1, "num_gpus": 0})
class SentimentClassifier:
    def __init__(self):
        self.classifier = pipeline("sentiment-analysis")

    async def __call__(self, request):
        json_data = await request.json()
        text = json_data.get("text", "")
        return self.classifier(text)

app = SentimentClassifier.bind()
# serve.run(app)

#Distributed Hyperparameter Tuning with Ray Tune

Running parallel hyperparameter trials:

from ray import tune

def training_function(config):
    # Simulated model training
    for step in range(10):
        intermediate_score = config["alpha"] * step + config["beta"]
        tune.report(loss=1.0 / (intermediate_score + 0.1))

search_space = {
    "alpha": tune.uniform(0.1, 1.0),
    "beta": tune.choice([1, 2, 5])
}

tuner = tune.Tuner(
    training_function,
    param_space=search_space,
    tune_config=tune.TuneConfig(num_samples=10, metric="loss", mode="min")
)
results = tuner.fit()
print("Best hyperparameters:", results.get_best_result().config)

Common Patterns

Stateful Distributed Actors for Model Serving

Problem: Re-loading large model weights into memory on every function invocation.

Solution: Use Ray Actors to hold in-memory state across requests:

@ray.remote(num_gpus=1)
class ModelServer:
    def __init__(self):
        # Load weights once into GPU VRAM
        self.model = load_large_model()

    def predict(self, sample):
        return self.model(sample)

server_actor = ModelServer.remote()
prediction = ray.get(server_actor.predict.remote(sample_data))

Best Practices (2026)

  • Do check the Ray Dashboard (port 8265) to monitor node CPU/GPU utilization, actor placement, and object store memory.
  • Do pass large read-only datasets via the Ray plasma object store (ray.put(data)) to prevent repetitive serializations.
  • Do specify resource requirements explicitly in decorators (@ray.remote(num_cpus=2, num_gpus=1)).
  • Do use Ray Serve for production multi-model microservices with built-in auto-batching.
  • Don't call ray.get() inside remote tasks; pass ObjectRefs directly to other remote tasks to preserve pipelining.
  • Don't create millions of micro-tasks; batch fine-grained tasks together to minimize scheduler overhead.
  • Don't pass large stateful objects (like database connections or open files) inside remote arguments.

Troubleshooting

ErrorCauseSolution
RaySystemError: Object store fullTasks producing more data than allocated shared memory plasma store.Use ray.get() to consume results, or store large arrays in disk storage.
Actor died unexpectedlyWorker process crashed due to unhandled exception or OOM.Check logs in Ray dashboard (http://localhost:8265) and increase worker RAM.
Task failed to schedule: Unsatisfiable resource requestRequested more CPUs/GPUs (num_gpus=2) than available on cluster.Verify available resources via ray.cluster_resources().

References

Individual skills in this repo

This repo contains 11 individual skills — each has its own dedicated page.

G1Joshi/Agent-Skills

Expert Dask distributed computing assistance covering Dask DataFrames, Arrays, Futures, and cluster scaling. Use when analyzing datasets too large for pandas on a single machine or multi-node cluster.

G1Joshi/Agent-Skills

Expert dbt (data build tool) assistance covering SQL modeling, Jinja macros, tests, documentation, and semantic layer. Use when building analytics engineering pipelines on BigQuery, Snowflake, or PostgreSQL.

G1Joshi/Agent-Skills

Expert JAX assistance covering Autograd, XLA compilation (`jit`), vectorization (`vmap`), and parallelization (`pmap`). Use when building high-performance numerical computing and cutting-edge deep learning research.

G1Joshi/Agent-Skills

Expert Git version control assistance covering branching, interactive rebase, cherry-pick, submodules, worktrees, and conflict resolution. Use when managing source code history and collaboration workflows.

G1Joshi/Agent-Skills

Expert K9s CLI assistance covering terminal Kubernetes cluster navigation, real-time log streaming, port forwarding, and pod debugging. Use when managing and troubleshooting Kubernetes clusters with speed.

G1Joshi/Agent-Skills

Expert SWC assistance covering high-performance Rust-based JavaScript/TypeScript compilation, .swcrc configuration, Jest testing via @swc/jest, and minification. Use when replacing Babel with SWC for faster builds, accelerating test execution, or compiling modern ECMAScript features.

G1Joshi/Agent-Skills

Expert Tig assistance covering text-mode interface for Git, interactive staging, commit graph visualization, diff exploration, and blame navigation. Use when navigating Git commit history in the terminal, staging hunks interactively, browsing file changes, or reviewing revisions.

G1Joshi/Agent-Skills

Expert Vim assistance covering modal editing, .vimrc configuration, registers, macros, search/replace, and plugin management via vim-plug. Use when editing text efficiently in terminal environments, writing Vimscript, recording macros, or configuring core Vim settings.

G1Joshi/Agent-Skills

Expert Zed editor assistance covering high-performance Rust-based text editing, multi-buffer editing, language server protocols (LSP), and AI assistant integrations. Use when configuring Zed settings.json, setting up language extensions, collaborating in real-time channels, or optimizing editor startup speed.

G1Joshi/Agent-Skills

Expert Zsh assistance covering shell customization, Oh My Zsh plugins, Zinit plugin manager, prompt engineering (Starship/Powerlevel10k), and shell scripting. Use when configuring .zshrc, writing Zsh automation scripts, optimizing shell startup time, or configuring tab-completion.

G1Joshi/Agent-Skills

Expert [skill-name] assistance covering [feature 1], [feature 2], and [feature 3]. Use when [working with X], [debugging Y], or [implementing Z].

関連スキル