Methods
How we train the bots
DokoArena has six established bot generations plus the Paragon beta. Origin, the first, follows the rules and takes a trick whenever it can. Vanguard, the second, added parameters for partner clues, foxes, Karlchen, and saving trump. It still did not know how much those choices mattered for winning.
Ascendant, the third generation, learned those weights through self-play and benchmarks. Zenith handles cheap ace-ruffs and safe partner feeding. Horizon uses the heart ten more carefully and makes sharper announcements. Apex is built for competitive contracts and solos. Paragon is the independent beta successor. Every generation plays with the same incomplete information a person has. Partners stay hidden until the ♣Qs or the contract reveals them. The matrix compares the generations side by side. The rest of the page explains the training loop.
Training volume
- bot games simulated across training and evaluation
≈1M
bot games simulated across training and evaluation
- card decisions recorded across self-play datasets
2.6M
card decisions recorded across self-play datasets
Head-to-head matrix
Each cell shows the row bot's decided win rate against the column bot. The result uses 5,000 paired deals, competitive Turbo rules, and the game-score metric.
The self-climb loop
Each bot scores positions with about 29 parameters. Raise fox-keep and it holds the Ace of Diamonds longer. Raise trump-conserve and it stops spending trump early. Raise Karlchen and it fights harder for the last Jack of Clubs. The climb loop changes those parameters, tests the profile, and keeps only versions that beat the previous winner.
Generations
The first six generations form a chain. Vanguard added tactical weights to Origin's rule-following base. Ascendant learned those weights, Zenith refined specific card-play mistakes, Horizon tightened restraint and announcements, and Apex added competitive contracts and solo decisions. Each generation kept the previous generation's core and changed the parts we were trying to improve.
Paragon starts over with a different policy. It ranks legal actions using its own public-information model instead of extending the Vanguard and Apex scoring engine. That is why it can make a familiar choice in one position and a completely different choice in another. We kept the lessons that survived the earlier climbs, including protecting partner tricks, handling the heart ten, conserving trump, and judging solos. The rewrite changes how those lessons combine at the table.
What's next
We will keep training stronger bot generations with more self-play, benchmarks against different playstyles, and better adaptation across several games.
We also plan to study player-vs-bot and player-vs-player replays. That will help us identify winning decisions and feed them back to the bots. The long-term goal is a bot that can challenge the strongest players.
How difficulty works
In Ranked vs Computer, the bots play around your rating. In Practice vs Computer, you pick Easy, Medium, Hard, or Terminator yourself. Both work the same way, and only Paragon uses these settings. Older generations always play at that generation's full strength.
The bot looks at every card it is allowed to play and knows which one it likes best. At Terminator it always plays that card. Easier settings sometimes pick a weaker card instead. The easier the setting, the more often that happens, and the worse that other card is allowed to be.
Optional solos follow the same pattern, and so do Re and Kontra. A weaker bot skips those more often, even when the strong one would go for it. A wedding is the exception. If the bot holds both Club Queens, it still announces.
A few mistakes are forbidden at every setting. The bot will not lead an unprotected Fox, and it will still take a partner's Fox with the Heart 10 when that catch is sitting right there.