Monitoring, Market Primitives, and the Stability of Algorithmic Collusion (R&R at Theoretical Economics)
This paper develops an analytical framework to study when sophisticated machine learning algorithms may learn to collude. Algorithms observe a state variable and update policies to maximize long-term payoffs; their long-run policies correspond to the stable equilibria of a tractable differential equation. In a repeated Bertrand game, I derive necessary and sufficient conditions under which Nash equilibria are learned. This reveals how the interplay between monitoring technology (state variables) and market conditions determines whether competitive or collusive outcomes emerge. I apply these insights to evaluate two key regulatory policies: limiting algorithmic data inputs and imposing competition in the software provider market.
We analyze strategic communication when advice is generated by a reinforcement-learning algorithm rather than by a fully rational sender. Building on the cheap-talk framework of Crawford and Sobel (1982), an advisor adapts its messages based on payoff feedback, while a decision maker best-responds. We provide a theoretical analysis of the long-run communication outcomes induced by such reward-driven adaptation. With aligned preferences, we establish that learning robustly leads to informative communication even from uninformative initial policies. With misaligned preferences, no stable outcome exists; instead, learning generates cycles that sustain highly informative communication and payoffs exceeding those of any static equilibrium.
We explore the behaviour emerging from learning agents repeatedly interacting strategically for a wide range of learning dynamics including Q-learning, projected gradient, replicator and log-barrier dynamics. Going beyond the better-understood classes of potential games and zero-sum games, we consider the setting of a general repeated game with finite recall, for different forms of monitoring. We obtain a Folk Theorem-like result and characterise the set of payoff vectors that can be obtained by these dynamics, discovering a wide range of possibilities for the emergence of algorithmic collusion.
Strategic Learning: When slow and steady wins the race
(with Galit Ashkenazi-Golan, Edward Plumb, and Yufei Zhang)
Learning agents are increasingly involved in decision making. When this decision making is for a strategic interaction, the question of strategically choosing the learning method emerges. We provide an initial step into understanding the implications of a strategic choice of a parameter - the speed of learning in multiagent gradient learning. We use 2x2 games to map the different considerations that are involved in choosing the speed strategically: the effect on basins of attraction, on cyclic behaviour and on the trajectory in dominance-solvable games. For the latter, we show that, while intuitively learning as fast as possible might seem to be an optimal choice, this is not always the case.
Democratic AI Alignment through Economic Theory
(with Elliot Creager, Kiran Dwivedi, Ali Falahati, Brandon Jaipersaud, Rohit Lamba)
This early-stage research applies economic theory to the socio-technical challenges of large language model (LLM) alignment, advancing along two primary tracks:
Preference Representation under Data Constraints: We are investigating the theoretical limits of aligning models to diverse populations when faced with inherently limited and unevenly distributed human preference data. This track explores the economic and statistical trade-offs inherent in preference aggregation, focusing on how structural bottlenecks in model capacity and data collection impact the representation of under-represented groups.
Axiomatic Approaches to Policy Constraints: Standard reward optimization for fine-tuning often struggles to balance competing heterogeneous preferences without homogenizing outputs. We are formalizing alternative alignment paradigms designed to bound unsafe or strictly prohibited behaviors. By reframing the collective alignment objective, we aim to design mechanisms that safely constrain policies while preserving the maximum possible space for downstream pluralism and personalization.
Zombie Prevalence and Bank Health: Exploring Feedback Effects (R&R at Management Science)
(with Andreea Rotarescu and Kevin Song)
This paper investigates feedback effects between bank health and zombie firms—financially distressed firms receiving subsidized credit. The literature focuses on how banks create zombies, overlooking zombies’ impact on bank health. Using Spanish firm-bank data (2005-2014), we document a vicious cycle: lower bank capital ratios are associated with higher zombie activity in served industries, while higher zombie prevalence is associated with reduced bank capital. We link this to a previously unexplored mechanism where banks respond appropriately to observable financial distress through higher provisioning, but overlook risks from relationship borrowers receiving subsidized rates. Our findings suggest that this feedback stems not from financial distress alone, but from the combination of distress with interest rate subsidies.