I’ve built a draft assistant for MTGA Powered Cube draft. This is difficult because we have to balance several aspects: 1. The power and win rate of each card individually. 2. How does each card work with other cards I have previously taken. 3. Based on the distribution of cards that we are seeing and what we’ve seen taken (after a pack goes all the way around) what are the open colors, strategies, and lanes.
This post summarizes the current state of the project: the models, the offline results, the first Monte Carlo tree search experiments, a walk-through of a pack experiment, as well as a live example.
A more paper-style writeup is also available here: Policy–Value Learning for Magic: The Gathering Arena Cube Draft.
A draft pick is represented with only information available to the drafter:
The base model is a contextual pick policy trained from public 17lands Powered Cube draft logs. Each card is represented by combining a card-specific learned vector with features extracted from Scryfall metadata and rules text: mana value, color and color-identity bits, card-type bits, rarity, power/toughness fields, keyword indicators, and oracle-text features. In model terms, the representation is
card_repr = learned_card_embedding + alpha * MLP(scryfall_features)
where alpha starts high and is annealed to a nonzero
floor. This lets common cube cards specialize through their learned
embeddings while still giving reasonable representations to changed or
unseen cube cards from their rules text and metadata. The policy scores
every card in the current pack and is trained to imitate human
picks.
There is one extra practical problem for live Arena Cube: the cube
list changes, and some cards in a new live list were not present in the
17lands training/scoring vocabulary. To account for this and prevent a
failure mode of an unknown card breaking the tool, I added a
function-aware k-NN fallback for those cards. Instead of mapping every
missing card to one generic UNK vector, the assistant
builds a pseudo-embedding from similar known cube cards. Similarity
combines structured Scryfall features, face-balanced rules text with
card names removed, and curated role tags like removal, protection,
tutors, rituals, fixing, artifact synergies, and graveyard synergies.
This lets a new protection creature borrow from cards like Mother of
Runes or Giver of Runes, while a new burn spell borrows from cards like
Lightning Bolt or Abrade.
Models trained:
The pick model is a cube-context contextual preference ranker. For a
cube with C cards, every training example has a
C-dimensional current-pack mask, pool count vector,
seen-but-unpicked count vector, and cube mask. The model first computes
card_repr for every card in the cube. It then uses separate
DeepSets encoders for three unordered sets:
pool_repr = DeepSet(cards in my pool)
seen_repr = DeepSet(cards I saw but passed)
cube_repr = DeepSet(cards in the cube list)
The draft context is the concatenation of those three set embeddings, an explicit WUBRG color-commitment vector from the pool, and learned embeddings for pack number and pick number. A small MLP maps that context to the same dimension as a card representation. Each card is scored by a dot product plus a card bias:
context = MLP(pool_repr, seen_repr, cube_repr, color_commit, pack_idx, pick_idx)
score(card | context) = dot(context, card_repr(card)) + bias(card)
The policy is trained over the in-pack support only. The main loss is a contextual preference / triplet ranking loss: the logged human pick should score at least a margin above every other card in the pack.
loss = mean_over_unpicked_cards max(0, margin - score(picked) + score(unpicked))
In the current implementation the representation dimension is 128,
the hidden dimension is 256, and the default margin is 0.2. The Scryfall
feature contribution alpha anneals from 1.0 to a floor of
0.3 over early training. I also use random cube-list masking during
training so the model is robust to partial or changing Arena Cube
lists.
For missing live-cube cards, the k-NN fallback precomputes top neighbors against the trained scoring vocabulary and caches the weighted average. On the Arena Powered Cube 5.0 list, 117 of 540 cards were missing from the current scoring vocabulary, and all were resolved through Arena/Scryfall lookup. In synthetic leave-20-cards-out tests over 20k held-out pick states and 200 trials, overall degradation was small because only a minority of states are affected: top-1 -0.0068, top-3 -0.0036, and MRR -0.0049. On the directly affected states where the picked card was held out, the drops were larger: top-1 -0.1253, top-3 -0.0714, and MRR -0.0902. So k-NN is a useful live fallback, not a substitute for training on those cards once data exists.
The outcome model is separate. It takes a built deck and sideboard,
encodes maindeck and sideboard cards with DeepSets over the same card
representation, concatenates rank/event/color covariates, and predicts
game win probability. I train this on 17lands game data, both per-game
and aggregated by (draft_id, build_index). For deployment,
a simple heuristic deckbuilder maps a draft pool to a 40-card deck
before scoring.
The game-card bonus model is an even simpler outcome model:
logit(win_probability) = intercept + sum(card_in_maindeck * card_bonus) + covariates
I tested pair-interaction variants, but they overfit relative to the card-only model. The learned top bonuses were plausible Powered Cube cards: Ancestral Recall, Time Walk, Black Lotus, Sol Ring, Mana Crypt, and the Moxen.
Finally, I generated static preference pairs for DPO. For a logged
pick state, if one in-pack candidate has a sufficiently higher learned
game-card bonus than another, that pair becomes
(preferred, rejected). DPO fine-tunes the pick policy
toward those game-data preferences while anchoring it to the human-pick
reference policy with a KL penalty. This gives the safe/aggressive
trade-off in the table below.
On a held-out Powered Cube split, the strongest human-imitation model gets roughly:
| Model / setting | Human top-1 | Human top-3 | Game-data pref. acc. |
|---|---|---|---|
| Human continued, no value rerank | 0.6240 | 0.9093 | 0.5494 |
| Human continued + value rerank | 0.6234 | 0.9077 | 0.6062 |
| Safer game-data DPO + value rerank | 0.6161 | 0.9048 | 0.6406 |
| Aggressive game-data + value rerank | 0.5092 | 0.8411 | 0.7262 |
It seems that the safer game-data model keeps most of the human-pick accuracy while improving game-data preference agreement. The aggressive model optimizes game-data preferences much harder, but it no longer behaves like a human drafter and often makes suspicious picks.
A lot. I tested the human-continued model by removing or corrupting the previous-pick context on a 20k held-out sample:
| Context | Top-1 | Top-3 | Same top pick as full context |
|---|---|---|---|
| full pool + seen | 0.6168 | 0.9070 | 1.0000 |
| no pool, keep seen | 0.4268 | 0.7545 | 0.5116 |
| keep pool, no seen | 0.6160 | 0.9065 | 0.9409 |
| no pool, no seen | 0.4070 | 0.7384 | 0.4856 |
| shuffled pool | 0.3507 | 0.6533 | 0.4088 |
The pool matters enormously. The seen-but-unpicked channel matters much less in the current model.
The next experiment was to add Monte Carlo tree search at inference time. This is not training. The model is already trained; MCTS is a search procedure that evaluates each legal pick by simulating possible continuations of the draft and scoring the resulting pools.
For each candidate pick, the search:
Because MTGA logs do not reveal hidden packs or opponent picks, this is an information-set approximation. It knows my pack and my pool, but future packs are sampled rather than perfectly reconstructed.
Latency with the model loaded once, on CPU, using real 17lands pick states:
| Simulations | p50 | p95 | Mean |
|---|---|---|---|
| 10 | 418 ms | 574 ms | 378 ms |
| 25 | 942 ms | 1370 ms | 850 ms |
| 50 | 1825 ms | 2697 ms | 1642 ms |
So 25 simulations is a reasonable default for live search-based reranking.
The current MCTS default is:
--simulations 25
--rank-by final
--root-policy-weight 0.15
--root-value-weight 1.0
--root-value-mode prob
--rollout-temperature 0.0
--cpuct 1.5
The following trace uses real packs from the held-out 17lands eval data. These are not independently sampled fake packs. They are the packs as seen by a real drafter, already depleted by previous picks. Pack sizes go from 14 down to 7.
Caveat: the later packs are real for the human’s actual draft path. If a model makes a different earlier pick, the later packs are not a perfect counterfactual. Still, this is much closer to reality than testing on random packs.
Pick 1, pack size 14
Arena of Glory; Blood Crypt; Containment Priest; Endurance; Godless Shrine; Grief; Omnath, Locus of Creation; Ouroboroid; Skyclave Apparition; Snapcaster Mage; Sundering Titan; The Wandering Emperor; Wishclaw Talisman; Yavimaya, Cradle of Growth
Pick 2, pack size 13
Birds of Paradise; Chain Lightning; Cosmogrand Zenith; Gitaxian Probe; Hymn to Tourach; Life // Death; Multiversal Passage; Soul-Guide Lantern; Talisman of Indulgence; Talisman of Unity; Torsten, Founder of Benalia; Wasteland; Witch Enchanter // Witch-Blessed Meadow
Pick 3, pack size 12
Dark Confidant; Deathrite Shaman; Demonic Tutor; Expressive Iteration; Get Lost; Liliana of the Veil; Questing Druid // Seek the Beast; Razorverge Thicket; Sylvan Caryatid; Underground Mortuary; Virtue of Loyalty // Ardenvale Fealty; Windswept Heath
Pick 4, pack size 11
Copperline Gorge; Elvish Reclaimer; Flickerwisp; Mine Collapse; Savai Triome; Scrubland; Talisman of Progress; Trumpeting Carnosaur; Unexpectedly Absent; Valki, God of Lies // Tibalt, Cosmic Impostor; Zuran Orb
Pick 5, pack size 10
Archon of Cruelty; Bleachbone Verge; Mana Confluence; Questing Beast; Sacred Foundry; Blazing Firesinger // Seething Song; Stomping Ground; Taiga; Tersa Lightshatter; Titania, Protector of Argoth
Pick 6, pack size 9
Crucible of Worlds; Emperor of Bones; Grim Lavamancer; Jetmir’s Garden; Restless Cottage; Restless Fortress; Rofellos, Llanowar Emissary; Sanguine Evangelist; Vampire Hexmage
Pick 7, pack size 8
Bone Shards; Collective Brutality; Exploration; Keen-Eyed Curator; Overgrown Tomb; Restless Vents; Utopia Sprawl; Winds of Abandon
Pick 8, pack size 7
Celestial Colonnade; Deep-Cavern Bat; Gloomlake Verge; Jadar, Ghoulcaller of Nephalia; Prismatic Ending; Sink into Stupor // Soporific Springs; Talisman of Conviction
| Pick | Human log | Human imitation | Safe game-data | Aggressive game-data | MCTS |
|---|---|---|---|---|---|
| 1 | Skyclave Apparition | Snapcaster Mage | Skyclave Apparition | Ouroboroid | Blood Crypt |
| 2 | Gitaxian Probe | Gitaxian Probe | Cosmogrand Zenith | Birds of Paradise | Hymn to Tourach |
| 3 | Windswept Heath | Demonic Tutor | Windswept Heath | Demonic Tutor | Demonic Tutor |
| 4 | Scrubland | Scrubland | Unexpectedly Absent | Zuran Orb | Valki, God of Lies // Tibalt, Cosmic Impostor |
| 5 | Sacred Foundry | Archon of Cruelty | Sacred Foundry | Titania, Protector of Argoth | Archon of Cruelty |
| 6 | Sanguine Evangelist | Emperor of Bones | Sanguine Evangelist | Emperor of Bones | Emperor of Bones |
| 7 | Winds of Abandon | Collective Brutality | Winds of Abandon | Keen-Eyed Curator | Bone Shards |
| 8 | Deep-Cavern Bat | Deep-Cavern Bat | Prismatic Ending | Deep-Cavern Bat | Deep-Cavern Bat |
Human log
Skyclave Apparition; Gitaxian Probe; Windswept Heath; Scrubland; Sacred Foundry; Sanguine Evangelist; Winds of Abandon; Deep-Cavern Bat
Human imitation
Snapcaster Mage; Gitaxian Probe; Demonic Tutor; Scrubland; Archon of Cruelty; Emperor of Bones; Collective Brutality; Deep-Cavern Bat
Safe game-data
Skyclave Apparition; Cosmogrand Zenith; Windswept Heath; Unexpectedly Absent; Sacred Foundry; Sanguine Evangelist; Winds of Abandon; Prismatic Ending
Aggressive game-data
Ouroboroid; Birds of Paradise; Demonic Tutor; Zuran Orb; Titania, Protector of Argoth; Emperor of Bones; Keen-Eyed Curator; Deep-Cavern Bat
MCTS value-conservative
Blood Crypt; Hymn to Tourach; Demonic Tutor; Valki, God of Lies // Tibalt, Cosmic Impostor; Archon of Cruelty; Emperor of Bones; Bone Shards; Deep-Cavern Bat
The safe game-data model most closely follows the human’s white/fixing lane. It takes Skyclave Apparition, Windswept Heath, Sacred Foundry, Sanguine Evangelist, and Winds of Abandon.
The human-imitation model actually chases more raw power here: Snapcaster Mage, Demonic Tutor, Archon of Cruelty. That is interesting because it is the model optimized for human agreement overall, but on this trace it is less committed to the logged human’s fixing-heavy path.
The aggressive game-data model still looks unsafe. Picks like Ouroboroid, Zuran Orb, Titania, and Keen-Eyed Curator are plausible outputs of a model chasing static game-outcome bonuses, but they look contextually dubious.
The most interesting part of the real-pack example is that MCTS appears to choose a lane and strategy early. After taking Blood Crypt, it moves into a black/red, maybe Rakdos-reanimator-ish, power lane: Hymn to Tourach, Demonic Tutor, Valki/Tibalt, Archon of Cruelty, Emperor of Bones, Bone Shards. Whether or not Blood Crypt is the best first pick, the later MCTS picks are at least coherent with that early commitment.
The other models look less strategically consistent in this trace. Human imitation takes individually strong cards, but jumps between Snapcaster, Demonic Tutor, Scrubland, Archon, black interaction, and Deep-Cavern Bat without as clear a deck plan. Safe game-data follows the logged human’s white/fixing lane more closely, but sometimes takes value/removal cards over synergy. Aggressive game-data mixes value and combo cards and produces the least coherent path. This suggest that search can amplify a questionable early lane choice, but it can also make the subsequent picks more internally consistent.
Pack 1, pick 1 shows all strategies rank Sol Ring the highest, because well it is the stringest card by far. Next is Animate Dead, which is a strong card for the reanimate archetype.
At pack 3, pick 1 we’ve taken a bit of a controlling artifacts deck. In which Staff of the Storyteller or Tezzeret would fit nicely into. But Ugin is very strong, providing removal, ramp, and card draw in a single car, and we have a good amount of acceleration and colorless artifacts that he might work well.
The deck ultimately ended up fine, not fantastic. Went 2-2. 2nd loss was due to my misplay (forgot to crew bankbuster and kill Kaito when I had the chance). Both wins were with playing Ugin (once the opponent scooped immediately and the other time I exiled their remaining board the next turn. But was the deck capable of 7 wins? Doubtful. Was it interesting? Very. I hope to get more runs and data from the next round of cube draft.