Stop backtesting
with yesterday's winners.
Download today's Nasdaq-100 list and backtest it — you silently inflate returns by 208 bps/year. You're only testing on survivors — companies that made it. YHOO, RIMM, WFMI were in the index too. Their data is gone. We collected it before it was.
# ❌ Every "quant" tutorial does this — and gets a lie import yfinance as yf tickers = get_current_ndx_100() # today's 100 survivors data = yf.download(tickers, start="2010-01-01") # +216 bps/yr illusion # ✓ Point-in-time universe — what was ACTUALLY in NDX on each date import pandas as pd pit = pd.read_parquet("ndx_pit_daily.parquet") ohlcv = pd.read_parquet("ndx_ohlcv.parquet") for date in trading_days: universe = pit.loc[date][pit.loc[date]].index # PIT universe features = make_features(ohlcv, pit, as_of_date=date) signals = strategy.get_signals(train, features) # real alpha
Your backtest is measuring the wrong thing
Take today's Nasdaq-100 list. Download 15 years of data. Run signals. Report Sharpe.
The implicit assumption: these 100 companies always existed and always traded.
On every date, use only the tickers that were actually in the NDX on that date — including the ones that later failed, got acquired, or were booted from the index.
That gap is survivorship bias. It accumulates silently every year you backtest on only today's winners. At 216 bps/yr, a strategy that looks like it beats the market by 2% actually doesn't — the edge was in the data selection, not the signal.
Everything in the box
One ZIP. Drop the parquets into crucible, point Claude at the skills files, and you're running institutional-grade backtests in minutes.
This dataset was assembled by cross-referencing Nasdaq-100 component change history with historical price feeds collected over time — including for companies that were later acquired or delisted. That price history is no longer publicly available for most of these tickers. You cannot replicate this dataset by running a script today — which is exactly why it exists.
ndx_ohlcv.parquet
31 MB of OHLCV history. MultiIndex (field, ticker) × date. 210 tickers, 4,900 days. Loads in <1 second with pandas or polars.
ndx_pit_daily.parquet
Boolean membership matrix: pit.loc[date, ticker] is True only if that ticker was in the NDX on that date. 4,900 × 270.
ndx_component_changes.csv
225 add/remove events with exact dates. Study inclusion effects, front-running patterns, or index arbitrage strategies.
ndx_pit_summary.csv
Per-ticker entry/exit dates and trading-day counts. Quick reference for understanding which stocks had short vs long NDX tenures.
7 × CLAUDE.md Skill Files
Markdown methodology guides written for Claude Code to read. Append them to your project's CLAUDE.md and Claude instantly understands:
SKILL_pit_dataset.md— PIT filtering, universe constructionSKILL_triple_barrier.md— López de Prado triple barrier labelsSKILL_cpcv.md— Combinatorial Purged Cross-ValidationSKILL_feature_engineering.md— Stationarity, CS z-scoring, frac diffSKILL_position_sizing.md— Kelly, half-Kelly, vol targetingSKILL_regime_detection.md— 200d MA, vol regime, HMMSKILL_strategy_research.md— Hypothesis-first autonomous research loop
The vibe trader workflow
The dataset is designed to plug directly into crucible, the open-source backtesting framework. Claude reads your CLAUDE.md skills, enforces the methodology, and calls you out on anti-patterns.
Buy + download
You get a ZIP with parquets + 7 skill files. Drop them into crucible/data/.
Install skills
Append the skill files to your CLAUDE.md. Claude Code reads them automatically on every session start.
Build strategies
Edit strategy.py. Claude enforces PIT filtering, triple barrier labels, and CPCV validation — and flags your anti-patterns before you commit.
Trust the OOS Sharpe
CPCV generates 15 independent OOS paths. You get a Sharpe distribution — not a single lucky number from one train/test split.
Loading NDX PIT dataset... ohlcv : 4,900 days × 210 tickers (31.3 MB) pit : 4,900 days × 270 tickers (165 KB) Running CPCV C(6,2) = 15 paths... path 00 sharpe=0.81 sortino=1.24 dd=-17.3% path 01 sharpe=0.74 sortino=1.11 dd=-21.5% path 02 sharpe=0.88 sortino=1.38 dd=-14.9% ... ── OOS Summary ────────────────────────────── oos_sharpe : 0.79 oos_sharpe_std : 0.09 ← tight distribution = robust signal cpcv_paths : 15 folds_passed : 13/15 max_drawdown : -21.5% elapsed : 47.3s
What makes this dataset different
Eliminates survivorship bias
Each ticker is only included in the universe during the exact period it was a member of the Nasdaq-100. Backtesting with only today's 100 winners inflates historical returns by 208 basis points per year — measured over 2010–2026 by comparing equal-weight survivors-only vs. point-in-time CAGR.
19 years of daily data
Coverage from January 2007 through the present day spans multiple full market cycles: the 2008 financial crisis, the 2020 COVID crash, the 2022 rate-hike bear market, and the 2023–2025 AI rally.
Point-in-Time membership table
The ndx_pit_daily file gives you a boolean matrix: for any date from 2007 to today, you can instantly look up which 100 tickers were in the index. This is the key building block for unbiased universe construction.
Ready-to-use Parquet format
All data ships as Apache Parquet files with a standard MultiIndex column structure (field, ticker). Loads in under 1 second with pandas or polars. No API keys, no rate limits, no streaming — just local files.
Full OHLCV — not just close prices
Open, High, Low, Close, and Volume are included for all 210 tickers with available data. This enables intraday entry/exit simulation, ATR-based position sizing, and volume-filter signals.
225+ documented rebalancing events
Every add and remove event is logged with its date, giving you the exact day each ticker entered or exited the index. Use this to study inclusion effects, front-running patterns, or index-arbitrage strategies.
Known limitations
We document what's missing. You deserve to know exactly what you're buying before you run a single backtest.
~62 acquired/delisted tickers have no OHLCV
Price history for companies that were acquired or delisted is no longer publicly available from any free source — the data was removed after these companies stopped trading. Affected tickers (e.g. WFMI, CELG, CERN, YHOO, RIMM) are listed in MISSING_TICKERS.txt. Their NDX membership dates are still correct in the PIT table — only the daily price bars are missing.
History starts February 2007
Component-change history was reconstructed from public Nasdaq-100 records, which only reliably cover back to ~2007. Pre-2007 history is not modelled. Backtests should start no earlier than January 2008 to avoid thin coverage.
Some early windows have <90 active members
In 2007–2010, the component change log has occasional gaps, so the reconstructed membership may undercount by a few names. This affects roughly 5–10% of trading days in that window.
Dual-class shares counted as separate slots
Tickers like LBTYA/LBTYK (Liberty) are treated as two independent NDX members, which can push the daily member count above 100 in some periods.
Prices are split-adjusted (not raw)
All OHLCV prices are split and dividend adjusted. Raw unadjusted prices are not included. This is correct for return calculations but means absolute price levels will differ from contemporary quotes.
Not an official exchange feed
This dataset is assembled from historical market data and public index records. It is not sourced from an official exchange or index provider. Occasional data quality issues may exist. Intended for research and strategy development only — not for compliance, auditing, or official index replication.
One price. Everything included.
No subscription. No API keys. No rate limits. Buy once, use forever.
- ✓
ndx_ohlcv.parquet— 4,900 days × 210 tickers OHLCV - ✓
ndx_pit_daily.parquet— point-in-time membership matrix - ✓
ndx_component_changes.csv— 225 rebalancing events - ✓
ndx_pit_summary.csv— per-ticker entry/exit dates - ✓ 7 × CLAUDE.md skill files (strategy research, PIT, triple barrier, CPCV, features, sizing, regime)
- ✓ Works with crucible open-source backtesting framework
- ✓ Apache Parquet — loads in <1s with pandas / polars
- ✓ Instant download after payment
Instant download after payment. Sold as-is with all known limitations documented openly. For research and education only. Not financial advice.
vs the alternatives
Ready to build strategies that actually work?
Join other Claude Code traders who stopped testing against phantom alpha.