2 papers across 2 sessions
We train RL agents directly from high-level specifications, without reward functions or domain-specific oracles.