Not a member of Pastebin yet?
Sign Up,
it unlocks many cool features!
- # Statistical Analysis of Competitive EDH League Season 1: Rating Distribution and Turn Order Effects
- **Authors:** isleep2late, AEtheriumSlinky
- **Affiliation:** cEDH League Community
- **Date:** November 14, 2025
- **Season Duration:** September 5 - November 7, 2025
- ---
- ## Abstract
- **Background:** The competitive Elder Dragon Highlander (cEDH) community desires a league to utilize rating systems to track player skill, but the fairness of turn order in multiplayer formats remains contested. cEDHSkill was developed to give players an approximated Elo rating.
- **Objective:** To analyze player rating distribution using OpenSkill/Elo conversion and test whether turn order position confers significant competitive advantage in 4-player cEDH games.
- **Methods:** We analyzed 358 confirmed games (1,432 player-matches) from Season 1, calculating Elo ratings for 59 active players (≥5 games). Turn order effects were evaluated using chi-square goodness-of-fit testing (α = 0.05) on 112 games with complete positional data (81 player-match wins, or the sum of all 4 positional wins).
- **Results:** Mean Elo rating was 1008 (SD = 47, range: 924-1115). Aggregate win rate was 22.0% with a 12% draw rate. Turn order analysis revealed position-specific win rates of 27.7% (1st), 24.2% (2nd), 19.5% (3rd), and 16.7% (4th), but chi-square testing found no significant deviation from expected equal distribution (χ² = 3.20, p = 0.362, df = 3).
- **Conclusions:** The rating system demonstrates healthy skill stratification with a 191-point Elo spread. Despite an 11-percentage-point gradient in positional win rates, statistical analysis does not support the existence of significant turn order advantages. The league's random seating policy appears fair, though larger sample sizes may detect subtle effects in future seasons.
- **Keywords:** cEDH, competitive Magic: The Gathering, Elo rating, turn order bias, chi-square analysis, OpenSkill
- ---
- ## 1. Introduction
- ### 1.1 Background
- Magic: the Gathering Commander (EDH) has evolved from a casual multiplayer format to a competitive scene with organized leagues and tournaments. Unlike traditional 1v1 formats, Commander features 4-player "free-for-all" gameplay, introducing unique strategic dynamics including threat assessment, political negotiation, and sequential turn order. Competitive EDH (cEDH) emphasizes optimized deck construction and play patterns, creating an environment where skill differences become measurable through rating systems.
- The effects of turn order in multiplayer games has been debated in game theory literature[1]. First-player advantage is well-documented in 2-player game modes[2]. Going first enables access to turn-based resources first, letting the player in the first seat play more proactively than if the player were later in turn order. In contrast, the last player in turn order generally has to adopt a reactive strategy. Positional advantages a re difficult to quantify considering all game actions are taken after players know the turn order. This gives a signfiicant window of opportunity to prepare for each position.
- ### 1.2 Rating Systems in cEDH
- Our league employs OpenSkill[3], a Bayesian rating system that, unlike TrueSkill[4], is nonproprietary. Each player maintains two parameters: μ (skill estimate) and σ (uncertainty)[5]. For presentation, we convert these to Elo ratings using the formula:
- **Elo = 1000 + (μ - 25) × 12 - (σ - 8.333) × 4**
- This conversion provides intuitive interpretation while preserving the underlying Bayesian updating mechanism. The σ penalty ensures that players with fewer games (higher uncertainty) are not overrated.
- ### 1.3 Research Questions
- This study addresses two primary questions:
- 1. **Rating Distribution:** How is skill distributed among competitive cEDH players after an inaugural season? What is the magnitude of skill stratification?
- 2. **Turn Order Effects:** Does turn order position (1st, 2nd, 3rd, 4th) significantly affect win probability in 4-player cEDH games?
- ### 1.4 Hypotheses
- **H₀ (Null):** Turn order position does not affect win rates; all positions should win something less than 25% of games (accounting for draws).
- **H₁ (Alternative):** Turn order position affects win rates, with systematic deviations from expected equal distribution.
- ---
- ## 2. Methods
- ### 2.1 Study Design and Data Collection
- This prospective observational cohort study analyzed Season 1 data from a Discord-based cEDH league spanning September 5 to November 7, 2025. The league uses a free-play structure where players self-organize into 4-player pods and submit results via Discord bot commands.
- **Inclusion Criteria:**
- - Confirmed games with status = 'confirmed' in database
- - For rating analysis: Players with ≥5 games played
- - For turn order analysis: Games with turn order data recorded
- **Exclusion Criteria:**
- - Unconfirmed or disputed games
- - "Ghost games" (25 games with zero player records, representing database artifacts, commander-game submissions, or submission errors that were redacted)
- - Players with <5 games (insufficient data for stable rating estimates)
- ### 2.2 Data Quality
- **Total Dataset:**
- - 383 games initially recorded in games_master table
- - 25 ghost games excluded (no player records)
- - **358 valid games analyzed** (1,432 player-matches)
- - 81 unique players registered
- - 59 active players with ≥5 games (72.8% retention)
- **Turn Order Data:**
- - 112 games with turn order recorded (31.3% of valid games)
- - 368 player-matches with positional information
- - 96% completeness rate within recorded games
- **Data Limitations:**
- - Self-reported turn order (potential reporting bias)
- - No deck/commander tracking in Season 1
- - Turn count not recorded
- - Selection bias (competitive-oriented player base)
- ### 2.3 Statistical Methods
- **2.3.1 Descriptive Statistics**
- Elo ratings were calculated using the formula specified in section 1.2. We computed mean, median, range, standard deviation, and quartiles for the population of active players (n=59). Win rates were calculated as aggregate (total wins / total matches) rather than average of individual rates to properly weight by activity level.
- **2.3.2 Chi-Square Goodness-of-Fit Test**
- To test H₀ (equal win rates across positions), we employed chi-square goodness-of-fit:
- **χ² = Σ [(Observed - Expected)² / Expected]**
- Where:
- - Observed = Wins at each position (1st, 2nd, 3rd, 4th)
- - Expected = Total wins × (1/4) for each position
- - Degrees of freedom = 4 - 1 = 3
- - Significance level α = 0.05
- We chose chi-square testing over alternative methods (e.g., multinomial regression) because:
- 1. Simple interpretation for community stakeholders
- 2. Robust to moderate sample sizes
- 3. Standard method in game balance analysis
- 4. No assumption of continuous predictors
- **2.3.3 Software and Reproducibility**
- All analyses were conducted in Python 3.12 using:
- - pandas 2.1.0 (data manipulation)
- - scipy 1.11.2 (statistical testing)
- - matplotlib 3.8.0 (visualization)
- - sqlite3 (database queries)
- Analysis scripts were developed with assistance from Claude (Anthropic,
- Claude Sonnet 4.5), an LLM assistant used for code generation and
- statistical consultation.
- ---
- ## 3. Results
- ### 3.1 League Overview
- **Participation Metrics:**
- - 358 confirmed games over 56 days
- - 81 registered players
- - 59 active players (≥5 games) = 72.8% retention
- - 1,432 total player-matches
- - Average games per active player: 23.2 (median: 17, and yes, someone actually played 105 games of cEDH)
- - Data entry: 152 games on launch day of cEDHSkill v 0.02 (September 13)
- **Temporal Patterns:**
- - 73.2% of days had at least one game (41 of 56 days)
- - Classic engagement curve: initial surge (152 games Day 1), sustained week 1 (94 games), moderate mid-season (95 games Sep 21-Oct 15), declining late season (42 games Oct 16-Nov 7)
- ### 3.2 Win-Loss-Draw Distribution
- **Aggregate Outcomes (1,432 matches):**
- - Wins: 315 (22.0%)
- - Losses: 945 (66.0%)
- - Draws: 172 (12.0%)
- **Game-Level Outcomes (358 games):**
- - Games with winner: 315 (88.0%)
- - Games ending in draw: 43 (12.0%)
- A 25% win-rate would assume zero draws. The 12% draw rate reduces available wins, lowering expected performance proportionally.
- ### 3.3 Elo Rating Distribution
- **Summary Statistics (n=59 players with ≥5 games):**
- | Statistic | Value |
- |-----------|-------|
- | Mean | 1008 |
- | Median | 998 |
- | Minimum | 924 |
- | Maximum | 1115 |
- | Range | 191 points |
- | Standard Deviation | 47 points |
- | Q1 (25th percentile) | 978 |
- | Q3 (75th percentile) | 1042 |
- | Interquartile Range | 64 points |
- **Top 10 Players by Elo:**
- | Rank | Player | Elo | W-L-D | GP | Win Rate |
- |------|-----------|-----|-------|----|---------:|
- | 1 | Owl | 1115 | 9-8-2 | 19 | 47.4% |
- | 2 | Amethyst | 1103 | 3-0-2 | 5 | 60.0% |
- | 3 | grenzo | 1101 | 19-24-4 | 47 | 40.4% |
- | 4 | graydog | 1099 | 12-14-3 | 29 | 41.4% |
- | 5 | MrSea | 1090 | 14-17-8 | 39 | 35.9% |
- | 6 | Madi | 1084 | 12-18-7 | 37 | 32.4% |
- | 7/8 | Jaws | 1070 | 7-16-1 | 24 | 29.2% |
- | 7/8 | Ra | 1070 | 8-13-4 | 25 | 32.0% |
- | 9 | padfoot | 1067 | 9-14-4 | 27 | 33.3% |
- | 10 | LegallyAby | 1066 | 6-9-1 | 16 | 37.5% |
- **Interpretation:**
- The 191-point Elo spread represents moderate, healthy skill stratification.
- Most players cluster within one standard deviation of the mean (961-1055 range), with the top 10% separated by approximately 100 points from the median. This degree of compression is appropriate for an inaugural season and will likely expand as skill differences clarify over subsequent seasons.
- ### 3.4 Win Rate Distribution
- **Summary Statistics (win rates for 59 players):**
- - Mean: 20.1%
- - Median: 18.2%
- - Range: 0% to 60%
- - Players above 25%: 20 (33.9%)
- - Players at 20-25%: 11 (18.6%)
- - Players below 20%: 28 (47.5%)
- **Note on Aggregate vs. Individual Averages:**
- The aggregate win rate (22.0%) differs from the mean of individual win rates (20.1%) because aggregate calculation properly weights by activity level. A player with 50 games at 30% win rate contributes more to aggregate statistics than a player with 5 games at 40%. For population-level inferences, aggregate metrics are more appropriate.
- Top performers with 15+ games demonstrate win rates in the 35-40% range, indicating that skill effects are detectable despite multiplayer variance.
- ### 3.5 Turn Order Analysis
- **3.5.1 Data Completeness**
- 112 games (31.3% of total) included turn order data, yielding 368 player-matches with positional information (less than 112 x 4 due to partial-reporting). Within these games, data completeness was excellent:
- - 79 games with all 4 players reporting (71%)
- - 19 games with 2 players reporting (17%)
- - 14 games with 1 player reporting (13%)
- **3.5.2 Positional Win Rates**
- | Position | Wins | Total | Win Rate | vs Expected 25% |
- |----------|------|-------|----------|-----------------|
- | 1st | 26 | 94 | 27.7% | +2.7% |
- | 2nd | 22 | 91 | 24.2% | -0.8% |
- | 3rd | 17 | 87 | 19.5% | -5.5% |
- | 4th | 16 | 96 | 16.7% | -8.3% |
- | **Total** | **81** | **368** | **22.0%** | - |
- **Observed Gradient:** 11.0 percentage points from 1st to 4th position
- **3.5.3 Chi-Square Test Results**
- **Null Hypothesis:** Turn order does not affect win rates (equal distribution expected)
- **Test Statistics:**
- - χ² = 3.20
- - p-value = 0.362
- - Degrees of freedom = 3
- - Critical value at α=0.05: 7.815
- **Conclusion:** FAIL TO REJECT null hypothesis (p = 0.362 > 0.05)
- **Interpretation:**
- While first position shows a 2.7% elevation and fourth position shows an 8.3% deficit relative to the expected 25%, these deviations are statistically consistent with random sampling variation. With 368 observations and p = 0.362, we estimate a 36% probability that the observed pattern could occur by chance alone under the null hypothesis of no positional effects.
- ### 3.6 Complete Outcome Distribution
- **Win-Loss-Draw by Position:**
- | Position | Wins | Losses | Draws | Total |
- |----------|------|--------|-------|-------|
- | 1st | 26 | 55 | 13 | 94 |
- | 2nd | 22 | 55 | 14 | 91 |
- | 3rd | 17 | 59 | 11 | 87 |
- | 4th | 16 | 66 | 14 | 96 |
- **Observations:**
- - Draw rates relatively consistent across positions
- - Loss rates increase with position
- ---
- ## 4. Discussion
- ### 4.1 Rating System Performance
- The Elo conversion successfully discriminates skill levels while maintaining interpretability. The 191-point spread with SD = 47 demonstrates that the rating system is functioning appropriately for an inaugural season. The top 10 players maintain win rates of 30-60%, substantially above the population mean of 22%, confirming that ratings correlate with performance outcomes.
- **Comparison to Other Games:**
- Our Elo distribution is narrower than established competitive games (chess, Go) due to the different formulation of OpenSkill.
- The Elo system used in cEDH cannot be compared to rating metrics in other games. MMR (Matchmaking Rating) is not a value that applies to our league because players arbitrarily and deliberately join games and create their own tables.
- **Validation of OpenSkill:**
- The OpenSkill Bayesian framework appropriately handles:
- 1. **Multiplayer complexity** (4-way outcomes vs. 1v1)
- 2. **Variable player pools** (not everyone plays everyone)
- 3. **Uncertainty quantification** (σ parameter penalizes undersampled players)
- The σ penalty in our Elo formula (−4 × [σ − 8.333]) prevents early-season outliers from dominating the leaderboard. Rank #2 has an impressive 60% win rate but only 5 games, resulting in Elo = 1103 rather than a potentially inflated rating.
- ### 4.2 Turn Order Findings
- **4.2.1 Interpreting Non-Significance**
- The p-value of 0.362 indicates we lack sufficient evidence to conclude turn order systematically affects win rates. This does NOT prove turn order has no effect (absence of evidence ≠ evidence of absence), but rather that:
- 1. If effects exist, they are likely small (<10% impact)
- 2. Current sample size insufficient to detect subtle effects
- 3. Random variation explains observed patterns adequately
- **4.2.2 Mechanistic Considerations**
- Several game-theoretic factors may explain why cEDH shows weaker positional effects than 1v1 formats:
- **Factors Favoring First Position:**
- - Priority on combo attempts
- - Ability to establish board presence before interaction
- - Tempo advantage (untap first, deploy threats first)
- **Factors Favoring Later Positions:**
- - Information advantage (see earlier plays before committing)
- - Hold up interaction with knowledge of threats
- - Political capital ("I went last, so don't target me!")
- **Multi-Opponent Dynamics:**
- - Three opponents can form coalitions against perceived leaders
- - Threat assessment typically targets strongest board state, not position
- - Table politics may override positional advantages
- In cEDH specifically, the prevalence of instant-speed interaction (Force of Will, Pyroblast, Pact of Negation) and turn 3-4 combo attempts may reduce first-player advantage compared to slower Commander variants (Brackets 1-3).
- **4.2.3 Comparison to Existing Literature**
- Academic game theory literature documents first-player advantage in:
- - Chess: ~52-56% win rate for White (various databases)[6]
- - Go: ~51-53% win rate for Black (with komi adjustments)[7]
- - Magic 1v1: ~55-60% for player on the play (format-dependent)[8]
- Our observed 27.7% for first position (vs. 25% expected) represents only a 2.7% elevation, substantially smaller than effects in 1v1 games. This aligns with theoretical predictions that multiplayer formats dilute positional advantages through coalition dynamics.
- **4.2.4 Practical Implications**
- For league organizers:
- - ✅ **Random seating is fair** - no adjustments needed
- - ✅ **No bracketing by position required** - unlike Swiss pairings or other methods
- - ✅ **Current policy validated** - maintain status quo pending additional data
- For players:
- - ✅ **Don't tilt about turn order** - 4th position still wins 16.7%, within variance
- - ✅ **Focus on decision quality** - skill matters more than position
- - ✅ **Sample size matters** - individual game outcomes noisy; trends emerge over multiple games
- ### 4.3 Draw Rate Analysis
- The 12% draw rate (43 of 358 games) is notable and deserves commentary. Common draw scenarios in cEDH include:
- 1. **Mutual destruction:** Multiple players are about to combo off simultaneously, with priority bullying in place
- 2. **Deck depletion:** Game goes long and players run out of resources
- 3. **Stalemate:** Board locks with no clear resolution path
- 4. **Time constraints:** Players may agree to draw rather than play out complex boards (80 min time limit, 20/player)
- The 12% rate of draws could be speculated as higher than typical Commander, but there is not enough data for casual Commander to make this claim.
- ### 4.4 Limitations
- **4.4.1 Internal Validity**
- - **Self-reporting bias:** Turn order is player-reported; no external verification
- - **Selection bias:** Only 31% of games have turn order data; may not represent full population
- - **Unmeasured confounders:** Deck choice, commander, pod composition not tracked
- **4.4.2 External Validity**
- - **Generalizability:** Results specific to this league/meta; may not extend to other communities
- - **Competitive selection:** Player base is competitively oriented; casual metas may differ
- - **Season effects:** Inaugural season data; patterns may change as meta stabilizes
- - **Discord-based:** Tech-savvy population may not represent broader Commander community
- **4.4.3 Statistical Limitations**
- - **Sample size:** 112 games adequate but not definitive; 150-200 games ideal for turn order analysis
- - **Assumptions:** Chi-square assumes independence; pod formation patterns could violate this
- - **Missing data:** 69% of games lack turn order information
- ### 4.5 Future Directions
- **Season 2 Enhancements:**
- Season 2 launches with enhanced cEDHSkill v 0.03.
- Expect revisions to prize structure due to tariffs/external factors.
- Player feedback is needed for improvement.
- ---
- ## 5. Conclusions
- This analysis of cEDH League Season 1 reveals a functioning competitive ecosystem with healthy skill stratification and fair game structure. Key findings include:
- 1. **Rating System Validation:** The OpenSkill/Elo hybrid successfully discriminates skill with a 191-point spread (924-1115), appropriate for an inaugural season. Top players demonstrate 35-40% win rates compared to population mean of 22%, confirming rating validity.
- 2. **Draw-Adjusted Performance:** The aggregate win rate of 22.0% exactly matches theoretical expectations of less than 25% given a 12% draw rate.
- 3. **Turn Order Effects:** Despite an observed 11-percentage-point gradient in positional win rates (27.7% for 1st, 16.7% for 4th), chi-square analysis finds no statistically significant deviation from random variation (p = 0.362). The league's random seating policy is supported by current evidence.
- 4. **Practical Implications:** League organizers can maintain current random seating procedures. Players should focus on skill development rather than positional concerns. Future seasons with larger sample sizes will provide more definitive conclusions.
- 5. **Competitive Viability:** With 72.8% player retention and sustained engagement over 56 days, the league structure demonstrates viability for organized cEDH competition.
- **Final Recommendation:** Continue current policies while enhancing data collection (turn order compliance, deck tracking) for Season 2. Combined analysis across multiple seasons will enable robust conclusions about positional effects and meta trends.
- ---
- ## Acknowledgments
- We thank the cEDH League community for their participation and commitment to data quality. Thank you to MoxMango for taking the lead on running ranked, and thank you to ShakeAndShimmy for allowing ranked to run on their server. Special appreciation to server administrators (Mori, Lerker) for assisting with implementation of the cEDHSkill Discord bot infrastructure and to all players who consistently reported turn order information.
- We would also like to thank Flowwer for providing artwork that was used towards prizing/marketing, as well as Beasts Mark (TFG) for contributing to prize support. Thank you to our league moderators: Anna, sky, JimWolfie.
- Data analysis and statistical computations were performed with assistance
- from Claude (Anthropic), an AI assistant, which helped with Python
- scripting, visualization generation, and statistical methodology.
- ---
- ## Data Availability
- Anonymized data files (CSV format) and analysis scripts (Python) are available from the corresponding author upon reasonable request. Raw database may contain personally identifiable information and is not publicly shared per privacy policy.
- If you would like to watch this report on YouTube: https://www.youtube.com/watch?v=YD3y7A_vnF0
- ---
- ## References
- 1. Gal-Or, Esther. “First Mover and second mover advantages.” International Economic Review, vol. 26, no. 3, Oct. 1985, p. 649, https://doi.org/10.2307/2526710.
- 2. Grepperud, Sverre, and Pål Andreas Pedersen. “First and second mover advantages and the degree of conflicting interests.” Managerial and Decision Economics, vol. 43, no. 6, 25 Nov. 2021, pp. 1861–1873, https://doi.org/10.1002/mde.3494.
- 3. Philihp. “Philihp/Openskill.Js: A Faster, Open-License Alternative to Microsoft TrueSkill.” GitHub, github.com/philihp/openskill.js/. Accessed 14 Nov. 2025.
- 4. Herbrich, Ralf, et al. “TrueSkillTM: A bayesian skill rating system.” Advances in Neural Information Processing Systems 19, 7 Sept. 2007, pp. 569–576, https://doi.org/10.7551/mitpress/7503.003.0076.
- 5. Weng RC, Lin CJ. A Bayesian Approximation Method for Online Ranking. *Journal of Machine Learning Research*. 2011;12:267-300, https://www.csie.ntu.edu.tw/~cjlin/papers/online_ranking/online_journal.pdf
- 6. Chess Statistics, www.chessgames.com/chessstats.html. Accessed 14 Nov. 2025. (Right now, White has won 56.2% of all calculated non-drawn games).
- 7. Counting_Zenist, et al. “Komi and Winrate Statistics.” Online Go Forum, 10 Aug. 2025, forums.online-go.com/t/komi-and-winrate-statistics/57674.
- 8. Multiple sources cite this, however there was one article by Florian Koch that went into great detail about first-player advantage on ChannelFireBall, which unfortunately no longer exists after the COVID Pandemic (https://strategy.channelfireball.com/all-strategy/mtg/channelmagic-articles/play-or-draw/). You can revisit a reference to this now-defunct page on stackexchange: https://boardgames.stackexchange.com/questions/27279/in-magic-is-the-player-going-first-advantaged
- ---
- ## Supplementary Materials
- **Supplementary Figure 1:** Daily activity timeline with engagement metrics https://imgur.com/fnNLkrp
- **Supplementary Figure 2:** Series of graphs showing daily game activity, elo distribution, win-rate distribution, activity distribution, Elo vs Win Rate, and top 10 players elo comparison (coded by arbitrary digits). https://imgur.com/RDzUnxQ
- **Supplementary Figure 3:** Raw turn order data (CSV, 368 observations) https://imgur.com/HwdwrOU
- ---
- **Corresponding Author:**
- Discord: I am isleep2late on https://discord.gg/cedh
- You can also reach AEtheriumSlinky on this server, as well as others mentioned in this paper.
- **Analysis Date:** November 14, 2025
Advertisement
Add Comment
Please, Sign In to add comment