Quoting at several levels at once
A reinforcement learning market maker places orders across the book rather than at the touch, and is tested against three kinds of counterparty.
2 minAlgorithmic & AI Trading
Patrick Cheridito and Moritz Weiss posted a reinforcement learning approach to market making on 18 August. The problem it takes on is the one that separates a textbook model from a working desk: not where to quote, but how to spread quotes across several price levels at the same time while keeping inventory under control.

The machinery
Order allocations are modelled with multivariate logistic-normal distributions, and a deep-set encoder turns a variable-length set of orders into a fixed-size representation — a sensible choice, since an order book has no natural ordering and no fixed length. Potential-based reward shaping speeds up learning without changing which policy is optimal.
Testing runs against three counterparty models: random noise traders, tactical traders reacting to volume imbalance, and strategic traders acting on a weighted volume signal.
That third case is the one worth watching. A market maker trained against noise learns to harvest spread; one trained against a counterparty that reads the book has to learn not to be read in return, and the paper is explicit that all three are simulations rather than historical data.
The abstract gives no revenue figures against a baseline. Until those appear, this is a description of an architecture rather than a claim about profitability.
Retold from arXiv. This is a summary in our own words; follow the link for the original reporting.