CommunityResearch & Data Analysisgithub.com

Unknown-333/building-dagster-assets

Build Dagster pipelines using software-defined assets — asset dependencies, partitions, resources and IO managers, asset checks, and schedules/sensors. Use when creating Dagster assets or jobs, modeling data as assets, adding partitions or backfills, wiring resources/IO managers, or migrating from task-based orchestration to assets.

What is building-dagster-assets?

building-dagster-assets is a Claude Code agent skill that build Dagster pipelines using software-defined assets — asset dependencies, partitions, resources and IO managers, asset checks, and schedules/sensors. Use when creating Dagster assets or jobs, modeling data as assets, adding partitions or backfills, wiring resources/IO managers, or migrating from task-based orchestration to assets.

Works with~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/Unknown-333/awesome-data-engineering-skills/tree/main/skills/building-dagster-assets

Installed? Explore more Research & Data Analysis skills: obra/superpowers, affaan-m/quarkus-verification, affaan-m/uspto-database · View all 6 →

Ask in your favorite AI

Open a new chat with this agent skill pre-loaded.

Documentation

Building Dagster Assets

When to use

  • Creating or refactoring Dagster software-defined assets and jobs.
  • Modeling tables/files/ML models as assets with lineage.
  • Adding partitions, backfills, asset checks, schedules, or sensors.
  • Do NOT use for Airflow (use the Airflow skills).

Workflow

- [ ] Model each output as an @asset; declare deps via function args
- [ ] Add partitions for time/category-sliced data
- [ ] Move IO (reads/writes) into IO managers or resources
- [ ] Add asset checks for data quality
- [ ] Schedule/sensor to materialize
  1. Think in assets, not tasks. An asset is a persistent object (a table, file, model). Declare dependencies by referencing upstream assets as function parameters — Dagster builds the lineage graph automatically.
  2. Partition assets that are naturally sliced (by day, region) so you can materialize/backfill one slice at a time.
  3. Resources and IO managers hold connections and read/write logic, keeping asset bodies focused on transformation and making them testable.
  4. Asset checks attach data quality assertions to an asset.

Patterns

Partitioned assets with a dependency:

import dagster as dg

daily = dg.DailyPartitionsDefinition(start_date="2026-01-01")

@dg.asset(partitions_def=daily)
def raw_orders(context: dg.AssetExecutionContext) -> None:
    day = context.partition_key
    write_parquet(f"raw/orders/{day}.parquet", fetch_orders(day))

@dg.asset(partitions_def=daily)
def orders_clean(context, raw_orders) -> None:  # depends on raw_orders
    day = context.partition_key
    transform_and_load(day)

@dg.asset_check(asset=orders_clean)
def no_null_ids(context) -> dg.AssetCheckResult:
    n = count_null_order_ids(context.partition_key)
    return dg.AssetCheckResult(passed=n == 0, metadata={"null_ids": n})

Resource / IO manager — inject a warehouse client instead of constructing it inside the asset:

@dg.asset
def orders_summary(context, warehouse: WarehouseResource, orders_clean):
    warehouse.execute("insert into summary select ...")

Schedule a partitioned job so each run materializes the latest partition; use a sensor to materialize when an upstream file/asset appears.

Common pitfalls

  • Task thinking — using bare @op/jobs for everything loses lineage, observability, and partition-aware backfills that assets give for free.
  • IO inside asset bodies — hard-coding connections makes assets untestable; use resources/IO managers.
  • Non-idempotent partitioned assets — materializing a partition must overwrite that partition, not append (see writing-idempotent-transformations).
  • Skipping asset checks — without them, bad data materializes silently; Dagster surfaces check failures in the UI and can block downstream.
  • Mismatched partition definitions between dependent assets — keep the partitions_def consistent so mappings resolve.

Individual skills in this repo

This repo contains 9 individual skills — each has its own dedicated page.

Unknown-333/authoring-airflow-dags

Write production-grade Apache Airflow DAGs using the TaskFlow API — idempotent tasks, correct scheduling and catchup, retries/SLAs, connections/variables, and avoiding top-level code. Use when creating or reviewing Airflow DAGs, scheduling pipelines, wiring task dependencies, configuring retries/backfills, or fixing non-idempotent tasks.

Unknown-333/building-dbt-models

Build well-structured dbt models — staging/intermediate/marts layers, ref() and source(), materializations, and incremental models with the right strategy. Use when creating or refactoring dbt models, choosing table vs view vs incremental, structuring a dbt project, or writing incremental logic.

Unknown-333/building-feature-pipelines

Build ML feature pipelines and feature stores — point-in-time-correct joins to avoid label leakage, offline/online parity, feature freshness and backfills, and materialization with tools like Feast. Use when engineering features for ML, preventing train/serve skew or data leakage, building a feature store, or backfilling historical features for training.

Unknown-333/building-iceberg-tables

Design and operate Apache Iceberg tables — partitioning and hidden partitioning, partition/schema evolution, snapshots and time travel, compaction and small-file cleanup, and MERGE/upsert for lakehouse tables on Spark, Flink, Trino, or Snowflake. Use when creating or maintaining Iceberg tables, choosing partitioning, evolving schema/partitions, or fixing small-file and metadata bloat.

Unknown-333/building-ingestion-pipelines

Build batch and incremental data ingestion (extract-load) pipelines — full vs incremental extraction, change data capture (CDC), watermarks and high-water marks, API pagination and rate limits, and choosing managed EL tools (Fivetran, Airbyte) vs custom code. Use when ingesting data from databases, APIs, files, or SaaS into a warehouse/lake, or designing incremental extraction and CDC.

Unknown-333/building-kafka-consumers

Build reliable Apache Kafka consumers and producers — consumer groups and partition assignment, offset commit strategy, at-least-once vs exactly-once, idempotent/transactional producers, rebalancing, and dead-letter handling. Use when writing Kafka consumers/producers, configuring offset commits or consumer groups, tuning throughput, or handling rebalances and poison messages.

Unknown-333/debugging-data-pipelines

Systematically root-cause data pipeline failures and data incidents — job errors, wrong or missing data, duplicates, and freshness misses — by tracing lineage upstream, isolating the failing stage, reconciling against source, and planning a safe fix and backfill. Use when a pipeline fails, numbers look wrong, data is missing or duplicated, a dashboard is stale, or a stakeholder reports a data discrepancy.

Unknown-333/designing-backfills-and-replays

Plan and run safe data backfills and replays — idempotent reprocessing of historical windows, partition-by-partition execution, isolating backfill compute from production, verifying results, and avoiding double-counting or changed history. Use when backfilling a new or fixed model, reprocessing after a bug, replaying events, or loading history for a new pipeline without corrupting existing data.

Unknown-333/designing-data-contracts

Define and enforce data contracts between producers and consumers — explicit schema, semantics, ownership, SLAs, and versioning — to prevent silent upstream changes from breaking downstream pipelines. Use when a producer schema change could break consumers, defining an interface between teams/services and the warehouse, or adding schema enforcement at ingestion.

Related Skills