Communitygithub.com

austin-starks/Public-Portfolio-Challenge

The single-touch lockbox — a final anti-overfitting holdout run once, after design freeze, on a window held out from every fold, sweep, and search. Use when finalizing a bakeoff winner before deploy, setting up the A/B/C baselines as OOS bars, or running the S1.5 gate-coherence auto-relax. Covers why looking at the lockbox twice burns it, the lockbox pass conditions, and the three baselines. The lockbox is distinct from the walk-forward OOS folds.

O que é Public-Portfolio-Challenge?

Public-Portfolio-Challenge is a Claude Code agent skill that the single-touch lockbox — a final anti-overfitting holdout run once, after design freeze, on a window held out from every fold, sweep, and search. Use when finalizing a bakeoff winner before deploy, setting up the A/B/C baselines as OOS bars, or running the S1.5 gate-coherence auto-relax. Covers why looking at the lockbox twice burns it, the lockbox pass conditions, and the three baselines. The lockbox is distinct from the walk-forward OOS folds.

Funciona com~Claude Code~Codex CLI~Cursor
npx skills add https://github.com/austin-starks/Public-Portfolio-Challenge/tree/HEAD/skills/lockbox-holdout

Perguntar na sua IA favorita

Abre um novo chat com esta habilidade de agente já pré-carregada.

Documentação

Single-Touch Lockbox & Baselines

The lockbox is a final holdout that no fold, sweep, or search ever touches — the last line of defense against calendar-fit selection pressure. It is distinct from the walk-forward OOS folds (walk-forward-oos), which are used repeatedly during search; the lockbox is touched exactly once.

The lockbox

  • What: the walk-forward span stops lockbox_width_days = 126 (~6 months) before the run date. That final window (run date − 126d) → run date is the lockbox — held out from ALL folds, ALL sweeps, ALL search.
  • How: only the assembled deploy-shape book that already passed every gate on the walk-forward aggregate runs over the lockbox, exactly once, after design freeze (Stage S2).
  • Pass conditions (frozen at S0, do not move them after seeing results):
    • lockbox OOS return ≥ worst single-fold walk-forward OOS return;
    • lockbox maxDD ≤ 55%;
    • posture clean (audit_backtest_posture, all four conditions);
    • breadth ≥ 9 eligible names (audit_backtest_breadth — see breadth-audit).

Single-touch is absolute

If you look at the lockbox and then iterate the design, the lockbox is burned — move the walk-forward span back and hold out a fresh tail before any deploy. A lockbox failure is treated as calendar-fit selection pressure and means no deploy, except by a logged owner override that explicitly names the lockbox.

The three baselines (A/B/C) as OOS bars (Stage S1)

Loaded from committed snapshots/*.json via create_portfolio. Score A and B on the current engine; copy C's frozen published numbers as the bar to beat.

  • A — equity buy&hold (the full universe). A return bar only, never a Sortino bar — its Sortino reflects an equity-rally smoothness a leveraged long-premium book can't match.
  • B — naive LEAP ladder (canonical spec: all names, no rank/filter, 14-day cadence, 8%-of-portfolio per name under 95% budget, structure = first affordable rung of [Δ0.50 call → Δ0.55/0.25 → Δ0.55/0.40 → Δ0.50/0.42 vertical], all 180–365 DTE, single exit at DTE ≤ 45). A Sortino bar only in folds with ≥ 9/20 breadth.
  • C — frozen incumbent bar with published per-fold OOS numbers, copied verbatim (do not re-run to "improve" them).

S1.5 — Gate-coherence auto-relax (deterministic, no owner round-trip)

Before designing, check whether "beat Baseline B" and the ≤55% posture cap are jointly satisfiable — B is a ~95%-deployed book, so a posture-capped finalist can't out-deploy it. If unsatisfiable, AUTO-RELAX Baseline B: derate B's totalBudget until its own median daily deployment ≤ 55%, re-run its S1 scoring, and log original vs posture-matched B. This is not a gate relaxation and not an owner decision — only the benchmark is made comparable; the thresholds stay frozen.

Individual skills in this repo

This repo contains 11 individual skills — each has its own dedicated page.

austin-starks/Public-Portfolio-Challenge

Build alternative-data custom indicators on NexusTrade (Reddit/WSB mentions, congressional disclosures, insider filings, news-flow) and wire them into a certified book as a rank/tilt/filter signal. Use when adding alt-data to a strategy, building a CustomIndicator via compute sessions, auditing lookahead safety or per-ticker coverage, or deciding whether a data series is dense enough to drive a rank. Covers the compute-session workflow, the sparse-series rule, and source recipes (see references/). Invoke with the NexusTrade MCP connected.

austin-starks/Public-Portfolio-Challenge

Measure a NexusTrade options book's TRUE participation at a fixed cold-start capital base, defeating the compounded-NAV breadth illusion. Use whenever a book's headline \"holds N/21 names\" needs validating, when a live book collapses to a single name (the OSCR trap), when checking whether a SelectTop / per-name-allocation / total-budget change degraded simultaneous participation, or when a certification needs a breadth gate. Uses audit_backtest_breadth at held-fixed $25k.

austin-starks/Public-Portfolio-Challenge

The \"loudly declare\" protocol for platform, engine, or data bugs during a NexusTrade certification campaign — bugs are a first-class deliverable, not something to route around. Use whenever a NexusTrade tool errors, hangs, or returns numbers that contradict the config; whenever you're tempted to write \"probably an engine quirk\"; or when a discovered bug invalidates prior backtests and results must be quarantined. Includes the bug hand-off doc template (see references/BUG_TEMPLATE.md).

austin-starks/Public-Portfolio-Challenge

The GATED deploy + cleanup flow for a NexusTrade live book — clone the finalist, preview a delta reconcile, stage UNAPPROVED orders, verify fills. Use ONLY after the human explicitly says \"deploy + clean up\" and names a finalist. Covers the clone-before-reconcile ordering, the single-tick reconcile expectation, stale-pending-order hygiene, the signal-freshness gate, and why you can never approve orders yourself. Nothing here runs during certification.

austin-starks/Public-Portfolio-Challenge

The mandatory pre-flight contract checks (Stage S0) that must pass before trusting NexusTrade's certification engine for any strategy work. Use at the start of a bakeoff or certification campaign to prove the walk-forward engine runs end-to-end, windows don't leak, fold winners persist, and the engine doesn't fabricate values for dead/not-yet-listed names. Each check has a STOP-and-report failure mode; these runs do NOT count toward any certification minimum.

austin-starks/Public-Portfolio-Challenge

The hard structural constraints for the Public Portfolio Challenge momentum-LEAP options book — the spread-shape rule, the take-profit convexity-cap footgun, the affordability ladder, and the known losers not to re-test. Use whenever building or auditing an options strategy structure, checking spread-shape compliance before a certification, choosing DTE/strike rungs, or deciding whether a proposed structure change is even allowed. A violation is an automatic certification FAIL.

austin-starks/Public-Portfolio-Challenge

Orchestrate an out-of-sample certification of a NexusTrade trading strategy or live book — the master discipline behind the Public Portfolio Challenge. Use whenever you must decide PASS/FAIL on whether a portfolio holds up out of sample before deploying real money, replaying the Episode 10 runbooks, or running a \"certify my book\" / \"prove it out of sample\" / \"re-certify the fix\" task with the NexusTrade MCP connected. Pulls in walk-forward-oos, breadth-audit, sweep-reoptimization, options-structure-rules, bug-protocol, and deploy-gate.

austin-starks/Public-Portfolio-Challenge

The single entry point that executes a Public Portfolio Challenge episode or addendum runbook end-to-end, delegating each stage to the functional skills. Use when asked to run/execute/replay Episode 10, its bakeoff, or its addendum with the NexusTrade MCP connected. Reads the target runbook, pins its real artifacts (IDs, the incumbent bar), sequences the stages, and stops at the gated deploy. It orchestrates; the functional skills do the work.

austin-starks/Public-Portfolio-Challenge

Run a multi-family strategy bakeoff — the SEARCH→CERTIFY funnel that screens many candidate mechanisms down to a certified deploy winner without letting the cheap search layer issue a verdict. Use when replaying the Episode 10 bakeoff, when exploring several distinct strategy families before certifying, when deciding whether \"no deployable winner\" is even a legal conclusion, or when building the per-family certification ledger. Enforces verdict-integrity: only certification can end a campaign.

austin-starks/Public-Portfolio-Challenge

Re-optimize a NexusTrade strategy with a walk-forward SWEEP and label parameter provenance — the discipline that prevents deploying inherited knobs. Use whenever a structural change (sizing, rung depth, universe membership, DTE family, adding a rank signal) forces a re-sweep, when authoring gene_intents from get_sweep_surface, when choosing sweep over GA for a deploy cert, or when selecting the cross-fold-robust winner instead of the per-fold argmax. Invoke with the NexusTrade MCP connected.

austin-starks/Public-Portfolio-Challenge

Run and read a NexusTrade walk-forward out-of-sample study — the certification engine behind the Public Portfolio Challenge. Use when certifying a fixed portfolio (backtest_only) or re-optimizing one (sweep), when setting fold_count / anchored / validation / embargo params, when monitoring a run_walk_forward_study to completion, when reading per-fold OOS returns/Sortino/drawdown, or when a variant's fold calendar needs calendar-alignment against a base control. Invoke with the NexusTrade MCP connected.

Habilidades Relacionadas