Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models
Quick Answer
This study reveals that in a no-limit Hold'em autoregressive model, most recoverable opponent-range information stems from visible betting composition rather than hidden states.
Quick Take
The model achieves 16.5-16.7% top-10 accuracy using action/value and composition, while hidden-state probes yield only 11.4-12.2%. This indicates a need for careful interpretation of belief probes in predictive modeling.
Key Points
- Opponent-range probes show positive results after action/value controls in two of three seeds.
- Visible public betting composition explains more opponent-range signal than hidden states.
- Action/value+composition baselines achieve 16.5-16.7% top-10 accuracy.
- Hidden-state probes yield lower accuracy at 11.4-12.2% top-10.
- Positive belief probes require targeted alternatives for accurate interpretation.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Hidden-state probes often recover latent labels in imperfect-information sequence models, but this alone does not establish that a model maintains a posterior belief distribution over hidden states. This paper studies this ambiguity in a no-range Limit Hold'em autoregressive model trained only on action and value targets, not on an opponent's hand or range. Opponent-range probes are positive after action/value controls in two of three seeds, and the behavior head predicts held-out actions about five percentage points above a baseline using only observable public history. However, visible public betting composition explains more opponent-range signal than residual hidden states, suggesting that most recoverable information comes from betting summaries. Action/value+composition baselines reach 16.5-16.7% top-10 accuracy while composition-residual hidden probes fall to 11.4-12.2%, and matched-composition comparisons are negative in every seed. We call this evidence pattern composition-bounded predictive support: hidden states remain behavior-predictive and opponent-range correlated, but most recoverable range information is explained by visible betting composition rather than residual hidden-state structure. This is a case-study claim about opponent-range representational evidence, not exact Bayesian posterior tracking or a causal belief mechanism. Synthetic control and oracle validations show that the same diagnostics accept posterior-sensitive states and reject raw composition states under matched controls. Thus positive belief probes should be interpreted through targeted alternatives before being treated as evidence of belief tracking.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.19369 [cs.AI] |
| (or arXiv:2607.19369v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.19369 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Qianyu Chen [view email]
[v1]
Sat, 13 Jun 2026 15:56:26 UTC (118 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.