Large Language Models and Decision-Making Theory

A new AI-generated tutorial on testing decision-theoretic concepts on LLMs

The next one in the AI tutorial series, and this time purely out of curiosity:

📄 Large Language Models and Decision-Making Theory

I teach decision theory, so the question behind this tutorial was irresistible: if you run the classic experiments on a large language model, what do you get? Does it satisfy the von Neumann–Morgenstern axioms? Does it violate independence the way people do in the Allais paradox? Is it loss averse? Does it update like a Bayesian?

The tutorial covers the theory first — preference axioms, revealed preference and GARP, prospect theory, heuristics and biases, Bayesian updating, games — and then what has actually been found when these are tested on LLMs. The honest answer turns out to be “it depends”: sometimes the model is more normatively consistent than any human, sometimes it reproduces human biases, and quite often it does something that is neither, which may be a genuinely non-human pattern or simply a measurement artefact.

That last part is what I found most interesting. A good chunk of the tutorial is about how to test properly — option-order effects, reading logits versus sampled text, benchmark contamination, personas, reliability — because many published findings in this area are arguably measuring the harness rather than the model.

As always: AI-generated, so treat it as a map rather than gospel. Enjoy.