DVC (Data Version Control)
Installation
pip install dvc[s3]
pip install dvc[gs]
pip install dvc[azure]
pip install dvc[all]
cd my-ml-project
git init
dvc init
Track Data Files
dvc add data/training_images/
dvc add data/dataset.csv
git add data/training_images.dvc data/dataset.csv.dvc .gitignore
git commit -m "Track training data with DVC"
Configure Remote Storage
dvc remote add -d myremote s3://my-bucket/dvc-storage
dvc remote add -d myremote gs://my-bucket/dvc-storage
dvc remote add -d myremote /mnt/shared/dvc-storage
dvc push
dvc pull
Build Reproducible Pipelines
stages:
prepare:
cmd: python src/prepare.py
deps:
- src/prepare.py
- data/raw/
outs:
- data/processed/
train:
cmd: python src/train.py
deps:
- src/train.py
- data/processed/
params:
- train.epochs
- train.learning_rate
- train.batch_size
outs:
- models/model.pkl
metrics:
- metrics/train.json:
cache: false
evaluate:
cmd: python src/evaluate.py
deps:
- src/evaluate.py
- models/model.pkl
- data/processed/
metrics:
- metrics/eval.json:
cache: false
plots:
- metrics/confusion_matrix.csv:
x: predicted
y: actual
train:
epochs: 50
learning_rate: 0.001
batch_size: 32
dvc repro
dvc repro train
Experiment Tracking
dvc exp run --set-param train.learning_rate=0.01
dvc exp run --set-param train.learning_rate=0.001 --queue
dvc exp run --set-param train.learning_rate=0.01 --queue
dvc exp run --set-param train.learning_rate=0.1 --queue
dvc queue start --jobs 3
dvc exp show
dvc exp diff
dvc exp apply exp-abc123
dvc exp push origin exp-abc123
Metrics and Plots
import json
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score, f1_score
import yaml
import pickle
with open("params.yaml") as f:
params = yaml.safe_load(f)["train"]
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = RandomForestClassifier(n_estimators=params["epochs"])
model.fit(X_train, y_train)
preds = model.predict(X_test)
metrics = {
"accuracy": accuracy_score(y_test, preds),
"f1_score": f1_score(y_test, preds, average="weighted"),
}
with open("metrics/train.json", "w") as f:
json.dump(metrics, f, indent=2)
with open("models/model.pkl", "wb") as f:
pickle.dump(model, f)
dvc metrics show
dvc metrics diff
dvc plots show metrics/confusion_matrix.csv
dvc plots diff
Data Access Without Cloning
dvc get https://github.com/org/ml-repo data/processed/dataset.csv
dvc import https://github.com/org/ml-repo models/model.pkl
import dvc.api
with dvc.api.open("data/dataset.csv", repo="https://github.com/org/ml-repo") as f:
import pandas as pd
df = pd.read_csv(f)
url = dvc.api.get_url("models/model.pkl", repo="https://github.com/org/ml-repo")
Key Concepts
.dvc files: Small pointer files committed to Git that reference large data in remote storage
dvc repro: Reproduce pipelines — only re-runs stages with changed dependencies
- Experiments: Branch-free experiment tracking — run, compare, and apply results
- Params: YAML parameter files tracked by DVC for reproducible configurations
- Metrics: JSON/YAML metrics files with built-in comparison tools
- Remote storage: S3, GCS, Azure, SSH, HDFS — data stays where you want it