Communitygithub.com

austin-starks/Public-Portfolio-Challenge

Run and read a NexusTrade walk-forward out-of-sample study — the certification engine behind the Public Portfolio Challenge. Use when certifying a fixed portfolio (backtest_only) or re-optimizing one (sweep), when setting fold_count / anchored / validation / embargo params, when monitoring a run_walk_forward_study to completion, when reading per-fold OOS returns/Sortino/drawdown, or when a variant's fold calendar needs calendar-alignment against a base control. Invoke with the NexusTrade MCP connected.

O que é Public-Portfolio-Challenge?

Public-Portfolio-Challenge is a Claude Code agent skill that run and read a NexusTrade walk-forward out-of-sample study — the certification engine behind the Public Portfolio Challenge. Use when certifying a fixed portfolio (backtest_only) or re-optimizing one (sweep), when setting fold_count / anchored / validation / embargo params, when monitoring a run_walk_forward_study to completion, when reading per-fold OOS returns/Sortino/drawdown, or when a variant's fold calendar needs calendar-alignment against a base control. Invoke with the NexusTrade MCP connected.

Funciona com~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/austin-starks/Public-Portfolio-Challenge/tree/HEAD/skills/walk-forward-oos

Perguntar na sua IA favorita

Abre um novo chat com esta habilidade de agente já pré-carregada.

Documentação

Walk-Forward OOS Certification

The deterministic out-of-sample engine. Walk-forward, validation mode, anchored, 5 folds, 2022→today — the OOS folds are the verdict, not a single backtest.

The study call

mcp__nexustrade__run_walk_forward_study (tools are deferred — load via ToolSearch). Always run preview_only: true FIRST to see the fold calendar + cost, confirm it spans 2022→today (must include both the 2022 bear and the April-2025 selloff — this book's deepest drawdown is an April-2025 event, so the folds have to span it), then flip to false and record the study_id.

{
  "portfolio_id": "<SUBJECT or a clean clone>",
  "global_start_date": "2022-01-01",
  "global_end_date":   "<TODAY>",
  "fold_count": 5,
  "walk_forward_mode": "anchored",
  "mode": "validation",              // per-fold independent OOS
  "oos_width_days": 252,
  "embargo_days": 14,
  "interval": "Day",
  "inner_mode": "backtest_only",     // certify the FIXED book — no per-fold optimizer
  "preview_only": true               // flip to false to actually run
}
  • Certify a fixed book: inner_mode: "backtest_only" (engine_kind ga/default — a sweep is rejected for a fixed-config cert).
  • Re-optimize (only on FAIL or a structural change): engine_kind: "sweep", inner_mode: "optimize", certification: true, plus gene_intents. GA overfits and is rejected for deploy certification — sweep is the certified path. See sweep-reoptimization.
  • certification: true applies the activity-floor + percentChange ≥ 0 + sortino ≥ 0.5 policy and defaults validation_percent to 50.

Pre-flight before certifying a live book

  1. get_portfolio — capture the strategy set, conditionFieldAudit, and the exit ladder.
  2. Strip the LaunchAgent. A LaunchAgent strategy can't run historically — it fires once and logs thousands of cooldownSkip events, contaminating folds. Certify a LaunchAgent-free clone (rebalance + close rules only) and say you did so.
  3. Spread-shape check the live structures (see options-structure-rules). A long-dated vertical spread is a hard violation → the book cannot certify as-is; go straight to re-opt.
  4. Convexity-cap footgun: if multiple "always" take-profit closes exist (e.g. P/L ≥ 20% / 50% / 200%), the lowest binds first and caps every winner at ~+20%, defeating the outright-call rationale. Flag it — it's the most likely reason returns trail a let-winners-run book.

Monitor the study

Poll list_walk_forward_studies / get_walk_forward_study_results <study_id> until terminal (COMPLETE / ERROR / CANCELLED). Parse with jq — payloads are large, don't dump them. Watch for known failure modes and surface them (don't paper over — that's a bug-protocol trigger):

  • Silent-fail / ERROR — no fold stats / empty validationAggregate. Report it; don't invent a verdict.
  • Wrong window — fold calendar not starting 2022-01-01 / skipping the bear.
  • Backtest contamination — LaunchAgent firing in folds (should have been stripped in pre-flight).
  • Zero-trade / NaN folds — upstream position-marking corruption; report which folds, never average over NaNs.
  • A study genuinely hung past its estimated cost/time → stop, summarize, wait for the human (the engine may be under repair in another tab).

The mandatory per-fold checks (these ARE the verdict)

  1. Per-fold OOS table — for each fold: train window, OOS window, OOS return, OOS maxDD, OOS Sortino, median deployment, distinct underlyings, participation. Plus the validationAggregate.
  2. Independent re-backtest spot-check — backtest_portfolio ($25k) on the full cycle (2022-01-01→today), the 2022 bear (2022-01-01→2023-01-01), and the last 12 months; audit_backtest_posture the full-cycle run. Study folds must agree with a fresh standalone backtest. Baseline ≠ SPY for an options book — use a material underlying or equal-weight universe B&H, and note it.
  3. Degradation — train vs OOS per fold. Large train→OOS collapse = overfit, even if the full-cycle backtest looks great.
  4. Drawdown honesty — OOS maxDD next to OOS return per fold; lead with risk. At this leverage a fold can post a huge return and a 50–77% drawdown.
  5. Breadth gate at FIXED $25k (breadth-audit).
  6. Reproducibility / field audit — conditionFieldAudit matches intended knobs; compare_backtests {tolerance_bps: 0} on a re-run. Verify by field, never display name.
  7. Spread-shape compliance (options-structure-rules) — hard reject on violation.

Calendar alignment (variant vs base)

The walk-forward engine clips the fold calendar to the indicator's / data's coverage, so a variant's study can silently get different folds than the base's. Always run a base CONTROL with the same global_end_date as the variant's clipped calendar and compare fold-for-fold. Report per-fold signal coverage next to per-fold OOS. A variant that only wins where the signal exists is a finding, not a cheat — but say it.

Individual skills in this repo

This repo contains 11 individual skills — each has its own dedicated page.

austin-starks/Public-Portfolio-Challenge

Build alternative-data custom indicators on NexusTrade (Reddit/WSB mentions, congressional disclosures, insider filings, news-flow) and wire them into a certified book as a rank/tilt/filter signal. Use when adding alt-data to a strategy, building a CustomIndicator via compute sessions, auditing lookahead safety or per-ticker coverage, or deciding whether a data series is dense enough to drive a rank. Covers the compute-session workflow, the sparse-series rule, and source recipes (see references/). Invoke with the NexusTrade MCP connected.

austin-starks/Public-Portfolio-Challenge

Measure a NexusTrade options book's TRUE participation at a fixed cold-start capital base, defeating the compounded-NAV breadth illusion. Use whenever a book's headline \"holds N/21 names\" needs validating, when a live book collapses to a single name (the OSCR trap), when checking whether a SelectTop / per-name-allocation / total-budget change degraded simultaneous participation, or when a certification needs a breadth gate. Uses audit_backtest_breadth at held-fixed $25k.

austin-starks/Public-Portfolio-Challenge

The \"loudly declare\" protocol for platform, engine, or data bugs during a NexusTrade certification campaign — bugs are a first-class deliverable, not something to route around. Use whenever a NexusTrade tool errors, hangs, or returns numbers that contradict the config; whenever you're tempted to write \"probably an engine quirk\"; or when a discovered bug invalidates prior backtests and results must be quarantined. Includes the bug hand-off doc template (see references/BUG_TEMPLATE.md).

austin-starks/Public-Portfolio-Challenge

The GATED deploy + cleanup flow for a NexusTrade live book — clone the finalist, preview a delta reconcile, stage UNAPPROVED orders, verify fills. Use ONLY after the human explicitly says \"deploy + clean up\" and names a finalist. Covers the clone-before-reconcile ordering, the single-tick reconcile expectation, stale-pending-order hygiene, the signal-freshness gate, and why you can never approve orders yourself. Nothing here runs during certification.

austin-starks/Public-Portfolio-Challenge

The mandatory pre-flight contract checks (Stage S0) that must pass before trusting NexusTrade's certification engine for any strategy work. Use at the start of a bakeoff or certification campaign to prove the walk-forward engine runs end-to-end, windows don't leak, fold winners persist, and the engine doesn't fabricate values for dead/not-yet-listed names. Each check has a STOP-and-report failure mode; these runs do NOT count toward any certification minimum.

austin-starks/Public-Portfolio-Challenge

The single-touch lockbox — a final anti-overfitting holdout run once, after design freeze, on a window held out from every fold, sweep, and search. Use when finalizing a bakeoff winner before deploy, setting up the A/B/C baselines as OOS bars, or running the S1.5 gate-coherence auto-relax. Covers why looking at the lockbox twice burns it, the lockbox pass conditions, and the three baselines. The lockbox is distinct from the walk-forward OOS folds.

austin-starks/Public-Portfolio-Challenge

The hard structural constraints for the Public Portfolio Challenge momentum-LEAP options book — the spread-shape rule, the take-profit convexity-cap footgun, the affordability ladder, and the known losers not to re-test. Use whenever building or auditing an options strategy structure, checking spread-shape compliance before a certification, choosing DTE/strike rungs, or deciding whether a proposed structure change is even allowed. A violation is an automatic certification FAIL.

austin-starks/Public-Portfolio-Challenge

Orchestrate an out-of-sample certification of a NexusTrade trading strategy or live book — the master discipline behind the Public Portfolio Challenge. Use whenever you must decide PASS/FAIL on whether a portfolio holds up out of sample before deploying real money, replaying the Episode 10 runbooks, or running a \"certify my book\" / \"prove it out of sample\" / \"re-certify the fix\" task with the NexusTrade MCP connected. Pulls in walk-forward-oos, breadth-audit, sweep-reoptimization, options-structure-rules, bug-protocol, and deploy-gate.

austin-starks/Public-Portfolio-Challenge

The single entry point that executes a Public Portfolio Challenge episode or addendum runbook end-to-end, delegating each stage to the functional skills. Use when asked to run/execute/replay Episode 10, its bakeoff, or its addendum with the NexusTrade MCP connected. Reads the target runbook, pins its real artifacts (IDs, the incumbent bar), sequences the stages, and stops at the gated deploy. It orchestrates; the functional skills do the work.

austin-starks/Public-Portfolio-Challenge

Run a multi-family strategy bakeoff — the SEARCH→CERTIFY funnel that screens many candidate mechanisms down to a certified deploy winner without letting the cheap search layer issue a verdict. Use when replaying the Episode 10 bakeoff, when exploring several distinct strategy families before certifying, when deciding whether \"no deployable winner\" is even a legal conclusion, or when building the per-family certification ledger. Enforces verdict-integrity: only certification can end a campaign.

austin-starks/Public-Portfolio-Challenge

Re-optimize a NexusTrade strategy with a walk-forward SWEEP and label parameter provenance — the discipline that prevents deploying inherited knobs. Use whenever a structural change (sizing, rung depth, universe membership, DTE family, adding a rank signal) forces a re-sweep, when authoring gene_intents from get_sweep_surface, when choosing sweep over GA for a deploy cert, or when selecting the cross-fold-robust winner instead of the per-fold argmax. Invoke with the NexusTrade MCP connected.

Habilidades Relacionadas