Behind the findings

How the Lab tests a question.

We compare specific playing habits in simulated games. Here is what that comparison can tell you, what the numbers mean, and what remains uncertain.

One habit at a time

Each experiment starts with a concrete question about a playing habit: refusing a discard win, calling eagerly, declaring a kong, or folding at a fixed signal. We give one group of computer players that habit and compare them with another group under the same rules. A “policy” means the instructions a computer player follows. The rest of each player’s behavior follows the game AI unless the experiment says otherwise.

The current game model

The Lab uses the game and scoring engines that run Hé Fēn. That means qualifying minimums, claims and scoring conditions are part of the experiment. The software plays complete hands and records outcomes such as wins, calls, blocked win attempts and deal-ins (discarding another player’s winning tile). A habit that completes a shape below the minimum does not get credited with a legal win. Model gaps still matter: the current engine does not enforce Riichi furiten, so the Riichi results are model diagnostics rather than validated real-table advice.

Seats and random deals

The earlier studies alternate seat arrangements to reduce positional bias, with both policies sharing a table. The September 2026 research uses matched pairs instead: each control and treatment replays the same seed, focal seat and round wind against unchanged advanced opponents. Read the design stated in the report; these experiments are not interchangeable.

What the numbers mean

A share of wins compares the wins credited to the policies. It is different from wins divided by all hands played, which also includes drawn hands. Net points include both gains and losses. A blocked-win count can include repeated attempts within one hand. Each report should identify which measure it is discussing, and point totals from different scoring systems are not directly interchangeable.

What a finding cannot tell you

A large number of hands reduces some random variation; it does not remove limitations in the policies, game model or chosen opponents. The Lab measures software players, not human expertise. A result about always declaring versus never declaring does not solve every individual declaration decision. The computer players also estimate how close a hand is to winning. Both groups use the same estimation method, but mistakes in that estimate can affect their decisions.

How to read the published reports

The finding pages retain the reported observations and describe the comparison in plain language. The earlier pages summarize their experiments. The September research notes include paired outcomes, a fixed sample plan, source hashes and approximate uncertainty intervals. Where no uncertainty interval is published, treat the number as a point estimate. A small difference or a failure to detect an advantage is not proof that two choices are equally good.

Null results and corrections belong here

We publish findings that favor a policy, findings that do not show a clear advantage, and corrections when an earlier interpretation was too broad. In particular, Classic and Club kong results must remain attached to their own scoring sheets. If a rule, policy or result is unclear, send us the study title and the point you would like explained.

Read the findings · Ask about the method