Privileged observations enable rapid and reliable policy discovery directly in the physical world
A. Terpin, R. D'Andrea
People refine many high-performance skills using information that is available during practice but unavailable or presented differently during later action: think of a figure skater landing a triple jump, a pitcher throwing a curveball for a strike, or a barista pouring latte art. We ask whether privileged information—information available during training but not during execution—about the state of a physical system can be decisive for discovering a high-performing control policy even when reproducing the resulting behavior does not require that information. For this, we directly interface a generalist reinforcement learning agent with a spinning cylinder in a tabletop water channel to maximize or minimize drag. The agent acts directly on the physical flow and receives dense observations of the cylinder wake. This setup has several desirable properties. First, it is a physical system, with the rich interactions and complex dynamics that only the physical world has: the flow is highly chaotic and extremely difficult, if not impossible, to model or simulate accurately. Second, we can state the objective—drag minimization or maximization—simply and encode it directly in the reward, yet good strategies are not obvious beforehand. Third, decades-old experimental studies provide recipes for simple, high-performance, periodic open-loop policies. Relative to the common no-control reference, these periodic open-loop policies increase drag by 26.6% ± 0.7% and reduce it by 29.7% ± 1.3%. With flow observations, the agent learns within tens of minutes closed-loop policies that increase drag by 25.5% ± 0.9% and reduce it by 32.4% ± 1.6%. We record action trajectories during online policy execution and can subsequently replay them open loop as fixed sequences without flow observations; these replays increase drag by 23.2% ± 2.2% and reduce it by 32.1% ± 3.0%. However, when we withhold flow observations during training, the agent still discovers high-performing drag-minimizing policies (31.4% ± 2.1% drag reduction), but no run learns a high-performing drag-maximizing policy (1.8% ± 4.3% drag increase). The experiments provide a physical demonstration that privileged observations can be decisive for policy discovery even in the extreme scenario when we can subsequently replay the resulting action trajectories successfully. Practically, our findings suggest that rich observations accessible during laboratory training can be powerful even when their continuous acquisition and processing would be impractical at deployment.