Cboe U.S. Equities Historical Market Volume Data 2025
- Rows
- 4,805
- Columns
- 14
Daily historical market volume, notional value, and trade counts for U.S. equities exchanges and TRFs for 2025, part of comprehensive multi-year series.
AI analysis
Dataset Analysis: Cboe U.S. Equities Historical Market Volume Data 2025
1. Dataset Overview & Analytical Value
This dataset, sourced directly from Cboe's official CDN at cdn.cboe.com/resources/us/equities/market-statistics/historical-market-volume/markethistory2025.csv, represents daily U.S. equity market activity broken down by exchange participant and tape classification (Tape A: NYSE-listed, Tape B: NYSE American/regional, Tape C: Nasdaq-listed). With 4,805 rows spanning 250 trading days and 20 distinct market participants, it forms a rich panel dataset capturing share volume, notional dollar value, and trade counts simultaneously. The Cboe provenance signals institutional-grade reliability and real-time publication cadence, making it highly suitable for market microstructure research, liquidity analysis, and cross-venue competition studies. The multi-year series context noted in the description means temporal correlation studies extending back several years are feasible by joining sibling files from the same CDN path pattern.
---
2. Data Quality Observations
Data quality is exceptionally clean across all 14 columns. The total null cell count is zero — a notable achievement for a financial time-series dataset of this size, eliminating the need for imputation strategies. All 4,805 rows were ingested without type-mismatch failures, with decimal and integer columns correctly typed throughout. The only mild concern is that Tape C Shares and Tape C Trade Count both contain a minimum value of 0, which could represent legitimate non-trading days for specific participants on Nasdaq-listed securities, or potentially placeholder/missing entries masked as zeros — these warrant a filter check. Duplicate row analysis is flagged as pending Phase D recomputation against Parquet, so a definitive duplicate count is not yet available; given the 250 distinct dates × 20 participants = 5,000 theoretical maximum rows versus 4,805 actual rows, roughly 195 date-participant combinations appear absent, suggesting either incomplete reporting for some participants on certain days or intentional exclusions that should be documented.
---
3. Key Column Distributions & Statistical Highlights
The volume and notional columns reveal a consistently right-skewed, heavy-tailed structure across all tape segments, reflecting the well-known power-law dynamics of equity market participation. Total Shares is the most striking: mean of ~915M shares vs. a median of only ~182M shares (a 5× gap), with skewness of 3.46 and 563 flagged outliers — almost certainly driven by a handful of dominant market makers or high-volume event days. Tape C (Nasdaq) dominates in absolute magnitude, with a max of ~8.91B shares and max notional of ~$367B, dwarfing Tape A and B, consistent with Nasdaq's larger share of U.S. equity volume. Trade count columns exhibit comparatively lower skewness (Tape C Trade Count: 2.66, the lowest in the dataset), suggesting trade fragmentation is somewhat more evenly distributed than raw share or notional volume. The interquartile ranges are extremely wide — for Total Notional, Q1 is ~$2.6B and Q3 is ~$36.9B, a 14× spread — confirming that participant-level heterogeneity dominates the dataset and any regression analysis must account for this variance structure, likely via log transformation.
---
4. Recommended Join Key Columns
The Date column is the primary join key (250 distinct values, zero nulls, confirmed as a proper Date type), enabling straightforward time-series alignment with any daily-frequency external dataset. The Market Participant column (20 distinct string values, zero nulls) serves as a secondary categorical key for panel data joins, allowing participant-level matching against exchange registration data, broker-dealer filings, or FINRA member records. Together, the composite key (Date, Market Participant) should yield a near-unique row identifier — the ~195 missing combinations noted above mean this composite key is not perfectly complete, so left-join strategies are recommended when merging with external sources to avoid unintended row loss.
---
5. Recommended Pairing Datasets for Correlation Discovery
Several dataset categories would unlock high-value correlations with this data. VIX and volatility indices (also published by Cboe at the same domain) would directly test whether elevated market volume precedes or follows volatility spikes — a classic market microstructure hypothesis. Federal Reserve FOMC announcement calendars and macroeconomic release schedules (e.g., CPI, NFP dates) could explain the outlier trading days visible in the extreme right tail of all distributions. SEC EDGAR short interest and dark pool reporting data would allow analysis of the relationship between lit-venue volume (captured here) and off-exchange activity. ETF flow data from providers like BlackRock or State Street would pair well given Tape B and C's heavy ETF composition. Finally, earnings release calendars (e.g., from Refinitiv or Bloomberg) could help attribute Tape C volume spikes to specific Nasdaq-listed company events, enabling event-study designs around reporting seasons.
Columns
- Date (date)
- Market Participant (string)
- Tape A Shares (decimal)
- Tape B Shares (decimal)
- Tape C Shares (decimal)
- Total Shares (decimal)
- Tape A Notional (decimal)
- Tape B Notional (decimal)
- Tape C Notional (decimal)
- Total Notional (decimal)
- Tape A Trade Count (integer)
- Tape B Trade Count (integer)
- Tape C Trade Count (integer)
- Total Trade Count (integer)