Cboe U.S. Equities Historical Market Volume Data 2015
- Rows
- 3,302
- Columns
- 14
Daily historical market volume, notional value, and trade counts for U.S. equities exchanges and TRFs for 2015.
AI analysis
Dataset Analysis: Cboe U.S. Equities Historical Market Volume Data (2015)
1. Dataset Overview & Analytical Value
This dataset, sourced from Cboe's official statistics repository at cdn.cboe.com/resources/us/equities/market-statistics/historical-market-volume/, captures daily U.S. equity market activity across 15 distinct market participants throughout the 2015 trading year. The 252 distinct dates align precisely with the standard U.S. trading calendar (approximately 252 trading days per year), confirming completeness of the temporal coverage. With metrics spanning share volume, notional dollar value, and trade counts broken out by Tape A (NYSE-listed), Tape B (NYSE American/regional), and Tape C (Nasdaq-listed) securities, the dataset offers a granular, multi-dimensional view of market structure and liquidity dynamics. Its primary value lies in studying market concentration, participant behavior, and volume-driven correlations across exchange venues — making it particularly compelling for microstructure research and competitive exchange analysis.
2. Data Quality Observations
The dataset exhibits exceptional structural cleanliness: zero null cells across all 14 columns and 3,302 total rows with no type mismatches detected. The 252 distinct dates multiplied by 15 distinct market participants yields exactly 3,780 expected rows under a fully balanced panel — the actual count of 3,302 suggests roughly 12.6% of participant-date combinations are absent, which likely reflects participants that were not active on certain dates (e.g., new entrants, temporary halts, or TRF-specific reporting gaps) rather than data corruption. Minimum values of 0 across share and notional columns (while Total Shares minimum is 100 and Total Notional is 982) are consistent with legitimate inactive-day reporting. Duplicate row counts are pending Phase D recomputation, but the fact that Total Shares and Total Trade Count show 3,302 distinct values — exactly matching total row count — strongly implies no duplicate rows exist, since every row appears unique on those aggregated fields.
3. Key Column Distributions
The volume columns reveal a heavily right-skewed market structure consistent with well-known "bursty" trading behavior. Tape A Shares (mean: ~280M, median: ~150M, σ: ~339M, skew: 1.81, outliers: 463) shows the mean is nearly double the median, indicating frequent high-volume trading days driven by market events pulling the distribution rightward — the 463 outliers deserve attention as potential earnings seasons, Fed announcements, or index rebalancing days. Tape C Shares exhibits the steepest skew at 2.06 with 450 outliers, suggesting Nasdaq-listed names (historically technology-heavy) experience more episodic volume spikes. Trade count columns tell a notably different story: Tape A Trade Count has the lowest skew among all columns at just 0.921 with only 30 outliers, implying trade frequency is more stable than share volume — a hallmark of algorithmic fragmentation where order sizes vary but execution cadence remains consistent. Total Notional (mean: ~$21.2B, median: ~$10.7B, Q1: ~$4.7B, Q3: ~$31.1B) shows the widest relative spread, with the IQR spanning nearly $26B, underscoring extreme day-to-day dollar value variability across participants.
4. Recommended Join Key Columns
The Date column is the primary and most reliable join key, with 252 distinct values mapping cleanly to the 2015 U.S. trading calendar — it should be treated as a foreign key when joining to any time-series financial dataset indexed by trading date. The Market Participant column (15 distinct string values) serves as a secondary dimensional key, enabling participant-level panel joins if a reference table of exchange/TRF identifiers is available. For aggregate market-level analysis, joining solely on Date after summing across participants will yield a clean 252-row daily time series. Note that composite key (Date + Market Participant) should be used for row-level joins to avoid fan-out duplication.
5. Suggested Pairing Datasets for Correlation Discovery
Several dataset categories would unlock high-value correlations with this data. CBOE VIX daily closing values (2015) would be the single highest-priority pairing — testing whether elevated VIX days correlate with volume spikes and increased outlier counts in Tape A/C shares. Federal Reserve FOMC announcement dates could explain discrete outlier clusters in trade counts and notional values. S&P 500 or Nasdaq Composite daily returns would allow testing of the well-established volume-volatility relationship at the exchange level, while also revealing whether TRF (off-exchange) volume share rises or falls during high-volatility periods. SEC Rule 605/606 execution quality data per venue would complement the participant-level breakdown, enabling cost-quality-volume correlations by exchange. Finally, earnings announcement calendars (e.g., from Compustat or Bloomberg) could explain Tape-specific volume surges — particularly Tape C spikes — by linking them to Nasdaq-listed mega-cap earnings events like Apple, Google, or Amazon reporting dates in 2015.
Columns
- Date (date)
- Market Participant (string)
- Tape A Shares (integer)
- Tape B Shares (integer)
- Tape C Shares (integer)
- Total Shares (integer)
- Tape A Notional (decimal)
- Tape B Notional (decimal)
- Tape C Notional (decimal)
- Total Notional (decimal)
- Tape A Trade Count (integer)
- Tape B Trade Count (integer)
- Tape C Trade Count (integer)
- Total Trade Count (integer)