Cboe U.S. Equities Historical Market Volume Data 2013
- Rows
- 3,780
- Columns
- 14
Daily historical market volume, notional value, and trade counts for U.S. equities exchanges and TRFs for 2013.
AI analysis
Dataset Analysis: Cboe U.S. Equities Historical Market Volume Data 2013
1. Dataset Overview and Analytical Value
This dataset, sourced from Cboe's official statistics repository at cdn.cboe.com/resources/us/equities/market-statistics/historical-market-volume/, captures daily U.S. equity market microstructure data across 15 market participants (exchanges and Trade Reporting Facilities) for the full 2013 trading year. The publisher — Cboe Global Markets — is one of the most authoritative sources for U.S. equity market structure data, lending high credibility to these figures. The dataset's 252 distinct dates align precisely with the standard count of U.S. trading days in a calendar year, confirming complete coverage. Its three-dimensional structure — decomposed by Tape A (NYSE-listed), Tape B (regional exchange-listed), and Tape C (Nasdaq-listed) securities — makes it exceptionally valuable for studying fragmentation across listing venues, intra-year volume seasonality, and the competitive dynamics among trading venues during a historically significant post-Reg NMS period.
---
2. Data Quality Observations
The dataset demonstrates outstanding completeness: all 14 columns report zero null cells across 3,780 rows, which is rare for financial market data of this granularity and suggests either rigorous upstream data curation by Cboe or a well-structured API extraction. The distinct count for Total Shares (3,780) and Total Trade Count (3,772) matching or nearly matching the row count indicates these aggregate columns are effectively unique identifiers per row — a strong sign of data integrity at the record level. Notably, the Tape-level columns (Shares, Notional, Trade Count) share a distinct count of approximately 3,480–3,529, slightly below total rows, implying a small number of exact ties across participants or low-volume days — not a data quality concern, but worth noting for deduplication logic. Duplicate row counts are pending Phase D recomputation, but the uniqueness profile of Total Shares and Total Trade Count makes full-record duplication unlikely. No type mismatches are flagged, and the mix of Integer, Decimal, String, and Date types appears semantically correct throughout.
---
3. Key Column Distributions and Notable Patterns
The Total Shares column tells a striking story of market concentration and volatility: its mean (412.5M) is nearly 2.5× the median (163.5M), with a right skew of 2.10 and 259 outliers — pointing to a small number of extremely high-volume trading days that dramatically pull the average upward. This pattern is consistent across all tape-level share columns, with Tape C Shares showing the most extreme skew (2.30, 458 outliers), reflecting Nasdaq-listed securities' well-known dominance in high-frequency and retail trading activity. The interquartile range for Total Shares (Q1: 32.9M → Q3: 598.7M) is extraordinarily wide, spanning nearly 18×, underscoring how profoundly volume differs across participants and dates. Trade Count columns present a comparatively tamer distribution — Total Trade Count skew of 1.21 with only 73 outliers — suggesting that while share volume spikes dramatically, the number of trades is more evenly distributed, hinting at variable average trade sizes worth investigating. Tape A Trade Count's remarkably low outlier count (24) versus Tape C's (145) implies NYSE-listed securities maintain more stable trading patterns than their Nasdaq counterparts.
---
4. Recommended Join Key Columns
The natural composite join key for cross-dataset linkage is {Date, Market Participant}, as these two columns together uniquely identify each row (252 dates × 15 participants = 3,780 rows exactly). Date alone is the strongest single join key for time-series alignment with external datasets, given its clean distinct count of 252 and zero nulls — standard ISO date formatting should be confirmed before joining. Market Participant serves as a categorical dimension key enabling joins to exchange reference data, regulatory filings, or market share rankings. For aggregated analyses where participant-level granularity is unnecessary, Date paired with the Total columns (Total Shares, Total Notional, Total Trade Count) provides a compact, high-quality daily market summary suitable for broad macro-financial correlations.
---
5. Recommended Pairing Datasets for Correlation Discovery
Several dataset categories would pair powerfully with this data. VIX/volatility index data (also published by Cboe) would be a natural first join on Date, testing whether elevated volume and trade counts reliably precede or follow volatility spikes — a hypothesis with strong theoretical grounding. S&P 500 or broad equity index return data (e.g., from CRSP or Yahoo Finance) joined on Date could reveal whether high-notional days correlate with large index moves, distinguishing informed from noise trading. Federal Reserve economic releases and FOMC meeting calendars for 2013 would enable event-study analysis, isolating whether taper tantrum dates (mid-2013) appear as volume outliers. Individual exchange market share reports (e.g., from FINRA or SRO filings) joined on Market Participant could contextualize which venues gained or lost share over the year. Finally, tick-level or order book data from the same period would allow this daily aggregate dataset to serve as a validation and filtering layer, anchoring granular microstructure research to reliable daily totals.
Columns
- Date (date)
- Market Participant (string)
- Tape A Shares (integer)
- Tape B Shares (integer)
- Tape C Shares (integer)
- Total Shares (integer)
- Tape A Notional (decimal)
- Tape B Notional (decimal)
- Tape C Notional (decimal)
- Total Notional (decimal)
- Tape A Trade Count (integer)
- Tape B Trade Count (integer)
- Tape C Trade Count (integer)
- Total Trade Count (integer)