Loaded Dice
When every outcome is legitimate, the attack is in the choosing—and AI now chooses everywhere
In 1980 the Pennsylvania Lottery drew its Daily Number live on television: physical ping-pong balls, rising from machines, in front of the camera. The number was 6-6-6. Beforehand, the host and a few insiders had injected paint into every ball except the 4s and the 6s, so only those could rise, and their friends had bought every combination of 4s and 6s.
What strikes me is that nothing illegitimate came out of the machine. 666 is a valid lottery number, drawn from real balls. The attack forged no outcome; it chose one.
Randomness used to be a quiet input
Attacks on randomness are not new. Most of the digital ones target cryptography: a predictable generator or a reused random value, and a secret key falls out.
But in ordinary software, randomness rarely decided what the program did. A hash seed, a backoff jitter, a load balancer’s coin flip: change any of them and the computation computes the same thing, slightly faster or slower. Randomness was an input to secrecy and to performance, not to behavior.
AI runs on dice
That is no longer true. An AI system draws random numbers at three layers (see Replay Comes Back). At the tensor layer, a reduction is summed in an order that depends on how many requests share the batch. At the token layer, the sampler draws the next token from a distribution. At the trace layer, the agent reads a clock, a process id, a tool’s output, whatever the world happens to say that second.
None of these can simply be removed. Always taking the most likely token produces text the paper that named the problem called “bland and strangely repetitive.” Repeated sampling is how models get better answers: on one coding benchmark, the share of problems solved by at least one sample rises from 15.9% with one sample to 56% with 250. And determinism has a price. Making inference bit-for-bit reproducible required batch-invariant kernels that, in the published measurement, took a run from 26 seconds to 42 at best. OpenAI documents its seed parameter with a sentence worth framing: “Determinism is not guaranteed.”
So randomness is now load-bearing. It decides which words appear, which command runs, which file gets deleted.
A new attack surface
Once randomness decides behavior, steering it is an attack, and the literature already has examples at every layer.
At the token layer, simply changing decoding settings, with no adversarial prompt at all, raised misalignment on eleven open models from 0% to over 95%. Resampling randomly perturbed prompts until one gets through, “best-of-N jailbreaking,” succeeded on 89% of attempts against GPT-4o with ten thousand tries. Work from my group searches the decoding tree directly and finds prompts that look safe under default settings but reach harmful outputs along paths the sampler can legitimately take. And the paint-in-the-balls attack has a literal analogue: a recent preprint shows that a compromised sampler’s random number generator can inject chosen tokens while every logit stays untouched.
At the tensor layer, the batch is a shared random variable. Two preprints show that an attacker who lands in the same batch as a victim can recover the victim’s prompt, through expert routing in one case and through batch-wide quantization in the other. Both leak rather than steer, so far; steering is the obvious next paper. At the trace layer, agent runs can fork on nothing but a timestamp. And nondeterminism is also good cover: an audit that expects outputs to vary has a hard time telling variation from a provider quietly serving a cheaper model.
No ground truth, and no single place to look
Two properties make these attacks different from the ones we know how to handle.
First, there is no wrong answer to catch. A conventional attack produces something checkably bad: a forged signature, a corrupted record, a value no correct execution could produce. An attack on randomness produces an output that some legitimate execution could have produced, so it passes any check of the form “does a legitimate run explain this output?” (see Correctness Without a Reference). The attacker lives inside that existential quantifier. Every candidate is legitimate; the attack is in which one was chosen. 666 passes every check a lottery number can pass.
Second, the choice is hidden by depth. A tie broken differently in a GPU reduction becomes a different token, which becomes a different shell command, which becomes a different state of the machine. By the time the damage is visible at the top, its cause is three layers down, and every layer in between behaved legitimately given its input.
Trust, but replay
Prevention will not get us far. We cannot remove the randomness, and we cannot forbid any outcome, because every outcome is allowed. What we can do is what lotteries learned to do after 1980. In Florida today, for example, the balls are weighed before and after each drawing, an independent accountant witnesses it, and a difference of more than a gram sends the set to investigators. Nobody claims the drawing cannot be rigged. The claim is that a rigged drawing leaves evidence.
For AI the equivalent is a discipline in three steps. Log the draws: every random value at every layer, the batch composition, the sampled tokens and seeds, the clock and tool results. Then replay: given the log, reproduce the run, and check that the randomness was what it claims to be, the same way a randomness beacon publishes signed, hash-chained values anyone can verify after the fact. Then attribute: when something went wrong, replay it with one recorded draw replaced by another legitimate candidate. If the harm disappears, that draw is the root cause. The same record-and-replay machinery serves debugging (see Replay Comes Back); here it faces an adversary instead of a bug.
The definition of root cause matters here, because it fits the attack exactly. You cannot say the chosen draw was invalid; it was valid. You can say that a different valid draw would not have done the harm, and that this draw was the one that mattered. That is a counterfactual, and counterfactuals need a recording.
Trust but verify, then, with one adjustment: for randomness, verification does not mean checking the outcome. It means being able to replay the choice.
References
- Pennsylvania Lottery scandal, 1980: TribLIVE, WESA, Butler Eagle.
- Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The Curious Case of Neural Text Degeneration. ICLR 2020.
- Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V. Le, Christopher Ré, and Azalia Mirhoseini. Large Language Monkeys: Scaling Inference Compute with Repeated Sampling. arXiv:2407.21787, 2024.
- Horace He and Thinking Machines Lab. Defeating Nondeterminism in LLM Inference. Thinking Machines Lab: Connectionism, September 2025.
- Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen. Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation. ICLR 2024.
- John Hughes, Sara Price, Aengus Lynch, Rylan Schaeffer, Fazl Barez, Sanmi Koyejo, Henry Sleight, Erik Jones, Ethan Perez, and Mrinank Sharma. Best-of-N Jailbreaking. NeurIPS 2025.
- Shuyi Lin, Anshuman Suri, Alina Oprea, and Cheng Tan. Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem. MLSys 2026.
- Ziyang You, Xiaoke Yang, Zhanling Fan, Feng Guo, Xiaogen Zhou, and Xuxing Lu. Seed Hijacking of LLM Sampling and Quantum Random Number Defense. arXiv:2605.08313, 2026.
- Itay Yona, Ilia Shumailov, Jamie Hayes, and Nicholas Carlini. Stealing User Prompts from Mixture of Experts. arXiv:2410.22884, 2024.
- Hanna Foerster, Ilia Shumailov, Cheng Zhang, Yiren Zhao, Jamie Hayes, and Robert Mullins. Dynamic Quantization Can Leak Your Data Across the Batch. arXiv:2604.26505, 2026.
- Will Cai, Tianneng Shi, Xuandong Zhao, and Dawn Song. Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs. arXiv:2504.04715, 2025.
- OpenAI Cookbook.
Reproducible outputs with the seed parameter
(page on the
seedparameter). - Florida Lottery. Rule 53ER16-40. Florida Administrative Code, 2016.