Methods

How we train the bots

DokoArena has six established bot generations plus the Paragon beta. Origin, the first, follows the rules and takes a trick whenever it can. Vanguard, the second, added parameters for partner clues, foxes, Karlchen, and saving trump. It still did not know how much those choices mattered for winning.

Ascendant, the third generation, learned those weights through self-play and benchmarks. Zenith handles cheap ace-ruffs and safe partner feeding. Horizon uses the heart ten more carefully and makes sharper announcements. Apex is built for competitive contracts and solos. Paragon is the independent beta successor. Every generation plays with the same incomplete information a person has. Partners stay hidden until the ♣Qs or the contract reveals them. The matrix compares the generations side by side. The rest of the page explains the training loop.

Training volume

bot games simulated across training and evaluation

≈1M

bot games simulated across training and evaluation

card decisions recorded across self-play datasets

2.6M

card decisions recorded across self-play datasets

Head-to-head matrix

Each cell shows the row bot's decided win rate against the column bot. The result uses 5,000 paired deals, competitive Turbo rules, and the game-score metric.

Origin
Vanguard
Ascendant
Zenith
Horizon
Apex
Paragon
Origin
27.7%
15.3%
15.8%
19.3%
24.6%
25.1%
Vanguard
72.3%
37.2%
27.4%
31.1%
33.5%
31.1%
Ascendant
84.7%
62.8%
37.3%
42.0%
36.9%
33.0%
Zenith
84.2%
72.6%
62.7%
61.5%
39.2%
34.8%
Horizon
80.7%
68.9%
58.0%
38.5%
38.7%
34.7%
Apex
75.4%
66.5%
63.1%
60.8%
61.3%
42.0%
Paragon
74.9%
68.9%
67.0%
65.2%
65.3%
58.0%

The self-climb loop

Each bot scores positions with about 29 parameters. Raise fox-keep and it holds the Ace of Diamonds longer. Raise trump-conserve and it stops spending trump early. Raise Karlchen and it fights harder for the last Jack of Clubs. The climb loop changes those parameters, tests the profile, and keeps only versions that beat the previous winner.

The self-climb loop
Mutate parameters
Play against previous winner
Keep winner, discard loser

Generations

The first six generations form a chain. Vanguard added tactical weights to Origin's rule-following base. Ascendant learned those weights, Zenith refined specific card-play mistakes, Horizon tightened restraint and announcements, and Apex added competitive contracts and solo decisions. Each generation kept the previous generation's core and changed the parts we were trying to improve.

Paragon starts over with a different policy. It ranks legal actions using its own public-information model instead of extending the Vanguard and Apex scoring engine. That is why it can make a familiar choice in one position and a completely different choice in another. We kept the lessons that survived the earlier climbs, including protecting partner tricks, handling the heart ten, conserving trump, and judging solos. The rewrite changes how those lessons combine at the table.

What's next

We will keep training stronger bot generations with more self-play, benchmarks against different playstyles, and better adaptation across several games.

We also plan to study player-vs-bot and player-vs-player replays. That will help us identify winning decisions and feed them back to the bots. The long-term goal is a bot that can challenge the strongest players.

How difficulty works

In Ranked vs Computer, the bots play around your rating. In Practice vs Computer, you pick Easy, Medium, Hard, or Terminator yourself. Both work the same way, and only Paragon uses these settings. Older generations always play at that generation's full strength.

The bot looks at every card it is allowed to play and knows which one it likes best. At Terminator it always plays that card. Easier settings sometimes pick a weaker card instead. The easier the setting, the more often that happens, and the worse that other card is allowed to be.

Optional solos follow the same pattern, and so do Re and Kontra. A weaker bot skips those more often, even when the strong one would go for it. A wedding is the exception. If the bot holds both Club Queens, it still announces.

A few mistakes are forbidden at every setting. The bot will not lead an unprotected Fox, and it will still take a partner's Fox with the Heart 10 when that catch is sitting right there.