System at a glance
Simulator
│
├─ asks the Engine for immutable PlayerKnowledge
│ └─ full self-deck and commander definitions, without shuffled order
│
├─ calls Pilot.startGame(PlayerKnowledge) once
│
├─ asks Engine for legal actions
│
├─ asks the Engine decision boundary to build a DecisionPoint
│ ├─ immutable PlayerObservation
│ └─ immutable LegalAction offers with engine-owned identities
│
├─ supplies that DecisionPoint to the selected Pilot
│ └─ PilotDecision references one offered action identity
│
├─ asks Engine to validate and resolve the decision
│
└─ executes the resolved engine Action and records the result
The Engine depends on the shared Pilot contract, not on a particular Pilot implementation. The Simulator selects and supplies the Pilot. The default is currently the heuristic implementation, but callers can inject another Pilot.
Design principles
- The Engine owns rules truth, game mutation, legal-action generation, and validation.
- A Pilot chooses among legal actions; it does not create legality.
- The Simulator owns the game loop, Pilot selection, logging, and aggregation.
- Pilot input is immutable, serializable, and safe for the player whose choice is being made.
- Hidden card instances and shuffled library order never cross the boundary.
- A Pilot normally returns an engine-issued action identity instead of manufacturing an action payload.
- Shared schema types may live in a transport-safe workspace package, while the Engine remains responsible for producing their runtime values.
Responsibilities
| Component | Owns | Must not own |
|---|---|---|
| Engine | GameState, rules evaluation, legal actions, observations, decision validation, action execution |
Heuristic preferences or a dependency on one concrete Pilot |
| Pilot | Policy, scoring, tie-breaking, and selection from offered actions | Rules truth, hidden state, or arbitrary unoffered actions |
| Simulator | Game orchestration, Pilot injection, turns, logs, observations, and result aggregation | A second legality implementation |
| Simulation contracts | Immutable, serializable Engine–Pilot and cross-workspace schemas | Engine runtime state or implementation behavior |
The executable card DSL remains Engine-owned rules truth. At game start, the Engine supplies the Pilot with immutable copies of the complete definitions for the player's own deck and commanders. This lets heuristic and model-backed Pilots understand what every known card does without importing the Engine's global definition catalog.
Decision lifecycle
When a game starts, the Simulator asks the Engine to construct immutable
PlayerKnowledge and calls Pilot.startGame() once. For each subsequent
choice, it gets the current legal Action values from the Engine and calls the
Engine-owned decision boundary with the selected Pilot. The boundary then:
- Constructs a fresh
PlayerObservationfromGameState. - Clones and freezes the legal actions.
- Assigns each offer an identity such as
action-0. - Calls
Pilot.decide()with the resultingDecisionPoint. - Resolves the returned identity to the original engine action.
- Rejects a missing, stale, manufactured, or invalidly specialized choice.
The Simulator executes only the resolved engine action. Empty legal-action
lists receive an Engine-owned END_TURN fallback, so that behavior is not
invented by individual Pilots.
Pilot contract
The public schema lives in @turnzero/simulation-contracts. The minimum
contract consists of:
| Type | Purpose |
|---|---|
PilotCardKnowledge<CardDefinitionPayload> |
One named card, its complete Engine-produced modeled definition, and any separately modeled faces |
PlayerKnowledge<CardDefinitionPayload> |
Stable full self-deck and commander knowledge supplied once per game |
ObservedCard |
Player-visible card identity and state, deliberately smaller than an engine card instance |
OpeningHandObservation |
Player-safe inputs needed by mulligan policy |
PlayerObservation |
Immutable knowledge available for the current decision |
LegalAction<ActionPayload> |
One frozen engine action paired with an engine-issued identity |
DecisionPoint<ActionPayload> |
The observation and complete set of currently offered legal actions |
PilotDecision<ActionPayload> |
A reference to one offer, with an optional validated specialization |
Pilot<ActionPayload, CardDefinitionPayload> |
The replaceable startGame(knowledge) and decide(decisionPoint) policy interface |
The package owns the schema. src/engine/pilotDecision.ts owns construction,
freezing, action identity assignment, and validation of runtime values.
Player knowledge
PlayerKnowledge is stable for the game. It contains every card in the
player's starting non-commander deck, grouped by name and count, plus each
commander. Every entry includes the complete cloned card-definition payload
produced from the Engine DSL. A normal Commander deck therefore exposes 99
non-commander cards; a two-commander deck exposes 98.
The deck list and card rules are player knowledge. The values deliberately do
not contain card-instance IDs or deck order. Definitions are supplied once at
game start rather than repeated inside every DecisionPoint. Future dynamic
knowledge for cards introduced from outside the starting deck remains a
separate extension.
Player observation
PlayerObservation currently exposes:
| Field | Player-facing meaning |
|---|---|
schemaVersion |
Version of the serialized observation shape; currently 1 |
turn, phase, turnContext, turnStep |
Current timing context |
lifeTotal |
The simulated player's life total |
zones.hand |
Cards in hand |
zones.battlefield |
Player-visible battlefield state |
zones.commander |
Command-zone cards |
zones.graveyard and zones.exile |
Player-visible public zones |
library.size |
Number of cards remaining in the library |
library.knownCards |
Individual library cards made known by a modeled look, reveal, or retained permission |
library.knownTopCardId |
Identity of a known top card when the top position is actually known |
openingHand |
Mulligan-specific hand, commander, land-play, and order-independent remaining-deck knowledge |
Observed cards include only explicitly selected player-facing properties, such as identity, name, controller, tapped state, commander/token markers, and counters. Adding another field requires deciding that the information is both available to the player and necessary for policy.
The opening-hand remainingDeck.cardsByName values are aggregate counts used
for probability estimates. They model knowledge of the player's own deck
composition without exposing the identity or order of shuffled card instances.
pendingMulliganBottomCards tells a Pilot how many cards still need to be put
on the bottom after keeping under the London mulligan.
For mulligan planning, the heuristic treats cards in commanderZone as
guaranteed casting options alongside the hand, without counting them toward
physical hand size, land totals, or London-mulligan bottom choices.
Legal actions and decisions
The legal-action list is authoritative. A Pilot selects one of those offers by
returning its actionId:
const firstActionPilot: Pilot<Action> = {
startGame(knowledge) {
// Cache the immutable deck and commander definitions for policy use.
},
decide(decisionPoint) {
const firstAction = decisionPoint.legalActions[0]
if (!firstAction) throw new Error("Expected a legal action")
return { actionId: firstAction.id }
}
}
Most choices need no returned payload. A small number of parameterized offers
represent a bounded family of legal selections without enumerating every
combination. For example, an attacker offer can be specialized into a concrete
declaration. A Pilot may return decision.action only for such an offer, and
the Engine validates the specialization before accepting it.
This pattern keeps large subset or assignment choices proportional to their candidate pools while preserving Engine-owned legality.
Pilot implementations
HeuristicPilot is the default policy today. SimulationOptions.pilot allows
the Simulator to use any implementation of the same interface. Tests already
exercise a trivial first-action Pilot to prove that the boundary is
interchangeable.
The contract is intended to support additional implementations without Engine changes, including deterministic replay, seeded random, alternative heuristic, and model-backed Pilots. Model training and model-specific planning APIs are separate concerns and are not part of the current contract.
Mulligan evaluation harness
The terminal player's mulligan-study mode provides a focused evaluation path
for 25 seeded Baeloth games followed by 25 seeded Quintorius games. Each actor
runs independently from the same 50 seeds. The study records every offered
pregame action and selection through BEGIN_GAME, including London-mulligan
bottom choices and a human unsure marker on close keep or mulligan judgments.
The default human workflow is:
npm run study:mulligans
At each keep-or-mulligan prompt, k and m record emphatic judgments while
ku and mu record the same selection with unsure: true. Every
non-automatic choice immediately asks for its reason while the commander, hand,
and selected decision are still visible. The choice and its non-empty reason
are saved together after every decision to
.data/mulligan-study/baeloth-quintorius-50-v5.json, so q can stop a session and
the same command resumes it. Heuristic results and the comparison summary use
the same file:
Every displayed hand, including the final kept hand, uses one bullet line per card. Hand cards are never collapsed into a pipe- or comma-separated row.
npm run study:mulligans -- --pilot heuristic
npm run study:mulligans -- --report
Each study batch uses a permanent suffix and deterministic seed. Spark and the heuristic are recorded before the human run without showing their choices in the human prompt. The next active batch remains v5 even though earlier local artifacts are no longer available:
npm run study:mulligans -- \
--output .data/mulligan-study/baeloth-quintorius-50-v5.json \
--base-seed baeloth-quintorius-mulligans-v5
Quitting prints the complete command needed to resume the selected batch rather than silently returning to another study.
Historical reports record Spark v2 at 41/50 unseen first hands (82%) versus the heuristic's 39/50 (78%). Its advantage was deck-specific: Spark led 23/25 to 17/25 on Quintorius, while the heuristic led 22/25 to 18/25 on Baeloth. The raw v1-v4 local artifacts were subsequently lost, so those results are historical evidence rather than a reproducible training corpus. Batch numbering remains monotonic; replacement collection begins at v5 rather than reusing an earlier identifier.
Relative .data paths resolve through Git's shared directory to the primary
checkout, and other relative generated-artifact paths are placed beneath the
same shared data root, so a linked worktree never owns the only copy. Before
replacing a study, the writer stores its prior contents as a compressed,
append-only revision below .data/mulligan-study/.history/. Set
TURNZERO_LOCAL_DATA_ROOT to use another durable location. Explicit output
paths inside a linked worktree are rejected.
For older or interrupted studies that have decisions without reasons, fill in only the missing annotations with:
npm run study:mulligans -- --annotate-human
This fallback pass displays the commander, complete hand, mulligan depth,
recorded decision, and whether the keep-or-mulligan judgment was emphatic or
unsure. It also includes London-mulligan bottom-card choices. Each one-line
explanation is saved immediately on that recorded decision; quitting with q
and rerunning the command resumes at the first choice without a reason.
Explanations should use only facts available when the decision was made, never
later draws or outcomes.
The persisted snapshots contain hand names, commanders, aggregate remaining deck counts, mulligan state, and offered action identities. They never contain the shuffled library order or hidden library instance IDs.
DeepSeek is connected by a deliberately asynchronous experiment harness rather
than by changing the synchronous core Pilot.decide() interface. The harness
builds each prompt from Engine-produced PlayerKnowledge and the current
DecisionPoint, supplies the full deck and commander definitions, enriches
them with Oracle facts from the local MTGJSON catalog when available, and
accepts only an offered actionId. Current-hand and commander facts are repeated
next to the decision so the model does not have to recover them from the full
deck payload. The prompt requires an explicit early-mana assessment and
turn-one-through-turn-three plan before the final choice. Forced single-action
steps bypass the provider entirely.
The harness automatically discovers the main checkout's local MTGJSON catalog when running in a worktree. A dry run validates the prompt and reports a conservative request ceiling without making a provider call:
npm run study:mulligans -- --pilot deepseek --dry-run
An actual run requires DEEPSEEK_API_KEY and is guarded by
--max-cost-usd (default $10). Responses, reasons, uncertainty, token usage,
and estimated cost are saved after each decision. Empty or malformed JSON is
retried up to three times, with usage aggregated when a later attempt succeeds.
The prompt is versioned so results from different evaluation instructions cannot
be silently mixed. To discard only older DeepSeek trajectories while preserving
human and heuristic results, run:
npm run study:mulligans -- --pilot deepseek --restart-deepseek
This isolates network latency and provider failure from the normal simulator while still exercising the same Engine-owned observation and legal-action boundary.
Spark mulligan baseline
Spark begins as a mulligan-only learned Pilot rather than a general gameplay model. Its deck profiles are subjective Pilot configuration, not Engine rules: the current Baeloth profile describes a slow graveyard-value build with flexible commander timing, while Quintorius describes a faster spellslinger build where the commander is an early engine. Both encode the reviewed human preference to keep any functional seven once the next mulligan would cost a card.
The local training command builds one player-safe dataset and a logistic-regression model artifact from the active v5 study:
npm run spark:train-mulligans
With no arguments, it reads
.data/mulligan-study/baeloth-quintorius-50-v5.json. Additional or alternative
batches can be supplied by repeating --study; seeded games from different
batches retain distinct grouped-validation identities.
The exporter retains the hand, deck/profile identity, mulligan state, human
reason, original label, and an effective reviewed label when the explanation
contains an explicit statement such as On review MULLIGAN + UNSURE. It never
exports shuffled library order or hidden card-instance IDs. The 19 recorded
bottom-card explanations are retained as deferred choices but are not training
targets for the binary v0 model.
Spark derives deterministic features from Engine-supplied PlayerKnowledge
and the current opening-hand observation. A class-balanced logistic model is
trained locally; uncertain human labels receive half weight. Five-fold
evaluation groups all decisions from one seeded game into the same fold to
avoid trajectory leakage.
Every training invocation creates an immutable
.data/spark/runs/<run-id>/ bundle with exact source-study copies, the dataset,
model, evaluation metrics, and SHA-256 manifest. Convenience outputs may move
forward, but a completed training run is never overwritten.
The resulting artifact can run through the same study harness:
npm run study:mulligans -- --pilot spark
SparkMulliganPilot consumes only player-safe knowledge and a DecisionPoint,
and returns an offered action identity. It uses the learned hand score to choose
keep or mulligan and to rank sequential London-mulligan bottom choices. Model
artifacts are identified by their training timestamp; use --restart-spark to
discard only incompatible Spark trajectories after retraining.
The first model established the dataset, feature, training, artifact, evaluation, and inference boundaries. Its validation numbers are a pipeline baseline, not evidence of product-level generalization. A new unseen cohort is required before claiming that Spark improves on the heuristic.
Spark opening-hand features
The first two reasoned studies contain 157 non-automatic keep-or-mulligan choices and 13 London-mulligan bottom choices. After applying explicit review statements in the reasons, the binary labels are 93 keeps and 64 mulligans; 30 are marked unsure. The reasons establish a hierarchy rather than a universal land-count rule:
- Usable mana is the floor: 36 of the 37 explicitly described zero- or one-land hands were mulligans. A nominal land is not necessarily usable; color, entering tapped, bounce and sacrifice costs, and conditional mana all matter.
- At the free-mulligan decision, a connected resource-development plan distinguishes otherwise mana-safe sevens. The reasons explicitly describe a setup or plan in 65 choices and a concrete early sequence in 65 choices.
- Once the next mulligan costs a card, the standard relaxes sharply: the recorded labels move from 51 keeps / 49 mulligans before a mulligan to 35 keeps / 12 mulligans after one. A functional seven is often preferred to gambling on six.
- Bottom-card choices preserve mana and the executable resource plan, usually removing replaceable interaction, protection, or late payoff.
The original learned feature set overfit correlations from the first cohort.
In particular, the two deck/profile identities supplied a strong prior, raw
Draw, Interaction, and Synergy tags learned misleading negative weights,
and ordinary lands were counted twice through type and zero-mana-value
features. On the second study's 67 shared decision states, 19 of Spark's 26
disagreements were false keeps: six had unusable zero- or one-land mana, two
could not make the required colors, and eleven were mana-safe but lacked a
connected plan. Its seven false mulligans were all Quintorius resource lines
involving conditional ramp, filtering, draw, or ramp into an engine.
Feature schema version 2 therefore replaces raw card, deck, and commander identity features with a bounded, deterministic no-draw analysis. It branches over viable physical land choices through turn five and emits only normalized, player-safe capability and mulligan-state facts:
| Feature family | Meaning |
|---|---|
hand:mana:land-option-count |
Physical cards with a playable land face; modal cards are still one card |
hand:mana:land-drops-through-turn-3 |
Maximum sustainable land drops on one branch, respecting tapped, bounce, sacrifice, and conditional lands |
hand:mana:available:<color>:turn-3 |
Up to three simultaneous mana of each exported Engine mana color, including reachable rocks or dorks |
hand:castable:cards-through-turn-3 |
Distinct nonland cards castable on at least one consistent branch |
hand:realizable:ramp-through-turn-3 |
Castable cards whose mana or land acceleration is actually reachable |
hand:realizable:card-access-through-turn-3 |
Castable cards with immediate draw, selection, or search rather than a raw Draw tag |
hand:opponent-conditional-resource-lines-through-turn-3 |
Reachable resource cards whose payoff depends on an opponent-facing condition |
hand:self-conditional-resource-lines-through-turn-3 |
Reachable resource cards whose payoff still needs a self-controlled prerequisite |
hand:mana:excess-land-options-above-four and five-or-more-land-options |
Nonlinear facts that let five- and six-land hands differ from an ideal three- or four-land start |
plan:max-connected-productive-actions-through-turn-4 |
Best single branch's sequential actions matching the deck profile's productive roles |
plan:excess-lands-without-reachable-resource-development |
Land-heavy hand with no reachable ramp, card access, or credible opponent-dependent resource line |
policy:free-mulligan:excess-lands-without-resource-development |
The same weakness while a replacement seven is still free; paid mulligans remain a separate policy state |
plan:commander-window-* |
Whether an early commander window applies and can be reached; flexible commander profiles do not treat it as a goal |
Card-face definitions come from immutable PlayerKnowledge; the analyzer does
not import the Engine's global card catalog and never receives card-instance
IDs or shuffled order. The Engine remains the owner of rules truth. This
bounded analysis is a Pilot-side forecast used to learn mulligan preferences,
not a legality or full game-tree API.
Feature schema version 3 intentionally invalidates earlier model artifacts. It adds the nonlinear land-surplus and free-mulligan interaction after v3 exposed Spark's tendency to keep land-heavy Baeloth hands with no resource-development plan. The combined 1,500-epoch run uses 241 binary decisions: 143 keeps and 98 mulligans. Five-fold validation, grouped by seeded game and study batch, agrees with 201 decisions (83.4% accuracy and 82.3% balanced accuracy), with 88.1% keep recall and 76.5% mulligan recall. Full-corpus training agreement is 208/241 (86.3%). These are development-corpus results, not evidence of generalization; v4 is the next untouched measurement.
Current migration status
The contract and orchestration boundary are active on the core simulation
paths. Mulligan evaluation and London-mulligan bottom-card selection use only
OpeningHandObservation and the offered legal actions.
HeuristicPilot receives no GameState. Its gameplay chooser and feature
calculations live in src/pilot/heuristic and consume a detached
PlayerRulesSnapshot from PlayerObservation.rules. The engine owns this
versioned payload and the rules queries in src/engine/playerRulesQueries.ts.
Scoring, strategy weights, and action selection remain in the pilot.
The snapshot explicitly lists the public rules facts needed by those queries, including pending choices, continuous effects, public event histories, and zone-object identities. It excludes the live library, RNG, callbacks, and deferred resolution bookkeeping. Library knowledge contains unordered deck counts, cards exposed by current choices, and explicitly visible top positions. Engine query implementations construct a private evaluation game from those facts, using synthetic identities for unknown library cards. That game is never executed or supplied to the pilot. These queries evaluate modeled rules; they do not provide a hidden-outcome rollout or sampling interface.
Existing standalone helpers such as chooseAction(game) remain compatibility
entry points for terminal callers and tests. They obtain legal actions from the
live engine and adapt its public snapshot before entering the heuristic.
HeuristicPilot does not call these adapters. JSON-round-trip gameplay tests
verify that decisions do not depend on a retained live game reference.
The payload currently uses the engine's card and rules types, and the heuristic still uses the local modeled-card catalog. Removing catalog imports or wiring pilot selection through batch workers is separate work. Future observation fields should follow demonstrated player-visible needs.
Extension rules
When a Pilot cannot make a sound decision with the current contract:
- Confirm that the missing fact is legitimately available to the player.
- Add the smallest serializable field to
PlayerKnowledgefor stable facts orPlayerObservationfor changing state. - Have the Engine derive and freeze that value from rules state.
- Test hidden-information exclusion and serialization.
- Update the Pilot to use the new contract field.
- Keep legal-action generation and validation in the Engine.
When introducing a new Pilot, inject it through the Simulator and prove that it
can select an offered action without receiving GameState. When introducing a
new kind of parameterized action, document and test the Engine validator for
the allowed specialization.