Physical intelligent systems, represented by autonomous vehicles and embodied intelligent robots, must satisfy strict hard constraints during operation. Safe reinforcement learning is an important method for enabling such systems to perceive, make decisions, and control in complex constrained environments. The key to ensuring the safety of reinforcement learning lies in how to construct constraints with iterative feasibility during the solution process.
To address this challenge, Masayoshi Tomizuka, a member of the U.S. National Academy of Engineering, proposed a method for continuously constructing dynamic safety indicators; A. Ames, an IEEE Fellow, proposed a method for formulating quadratic programming problems using control barrier functions; and C. Tomlin, a member of the U.S. National Academy of Engineering, proposed the Hamilton-Jacobi (HJ) reachability analysis method for safe reinforcement learning. However, the field of safe reinforcement learning still faces the following interrelated challenges: First, how can the relationship between policy safety and constraint feasibility be established? Second, how can the effects of the form and strength of constraints on their iterative feasibility be analyzed? Third, how can different constraint construction methods be compared in terms of their safety assurance performance within a unified theoretical framework?
To address these issues, the research group led by Professor Shengbo Eben Li from the School of Vehicle and Mobility and the School of Artificial Intelligence at Tsinghua University proposed a dual-time-domain analysis framework for safe reinforcement learning. They proved the equivalence between the arbitrariness of constraint satisfaction and the maximality of the feasible region, revealed how the safety of the resulting policy varies with the strength of virtual time-domain constraints, and provided a theoretical framework for safety assurance and algorithm design in reinforcement learning.
The study first notes that the analysis of reinforcement learning safety should be associated with two time domains: the real-time domain and the virtual-time domain. The virtual-time domain is used for problem formulation and algorithm design in reinforcement learning, whereas the real-time domain is used to deploy the trained policy and observe its interaction with the physical system. The hard constraints in the real-time domain are specified by the physical system, while the constraints in the virtual-time domain can take different forms and need not be equivalent to the hard constraints in the real-time domain. Policy safety here refers to the long-term satisfaction of the hard constraints in the real-time domain, which is distinct from whether the virtual-time-domain problem is solvable. This dual-time-domain perspective decouples policy solving from policy application, providing considerable freedom for constraint design in reinforcement learning problems.

A Dual-Time-Domain Analysis Framework for Safe Reinforcement Learning
On this basis, the study develops a definition of feasibility and an iterative feasibility guarantee theory for safe reinforcement learning. It defines “feasibility” as whether the virtual-time-domain problem admits a solution, and uses the “iteratively feasible region” to characterize the set of initial states for which this problem and all subsequent problems admit solutions, indicating that each type of virtual-time-domain constraint has a corresponding iteratively feasible region. Leveraging the controlled invariance of the feasible region, the study proves the relationship between safety and feasibility, namely that for any policy, the following two conditions are equivalent: first, it satisfies the hard constraints over an infinite horizon; second, its iteratively feasible region is maximized. The study further finds that as the strength of the virtual-time-domain constraints increases, the iteratively feasible region first expands and then shrinks, and reaches its maximum when the strength of the virtual-time-domain constraints equals that of the real-time-domain constraints.
The study further classifies existing safe reinforcement learning methods into two categories according to the form of virtual-time-domain constraints: the controlled invariant set method and the constraint aggregation method. The former uses the forward invariance of a set to guarantee the iterative feasibility of the policy, whereas the latter employs an aggregation function to equivalently replace infinite-horizon constraints with one-step constraints. This classification framework incorporates numerous mainstream safe reinforcement learning methods—such as the Safety Index, Control Barrier Function, and Hamilton–Jacobi (HJ) reachability analysis—into a unified theoretical framework. It enables direct comparison of the effects of different methods on policy safety and provides algorithm-design-level guidance for constraint forms conducive to long-term safety.

Membership relationship between safety constraints and the feasible region (schematic diagram)
The research findings were published on April 2 in Foundations and Trends in Systems and Control (FTSYS) under the title "The feasibility theory of constrained reinforcement learning: a tutorial study." The paper comprises 72 pages and 9 chapters, systematically presenting the team's research achievements in the safety and feasibility of reinforcement learning.
FTSYS was founded in 2014 and primarily publishes review and tutorial articles in the fields of artificial intelligence, control theory, and systems science. Its editor-in-chief invites leading research teams worldwide to provide in-depth analysis and systematic discussion on specific topics. The journal publishes only 2–3 articles per year, with a total of 35 articles since its inception.
Yang Yujie, a 2026 PhD graduate of the School of Vehicle and Mobility (currently a Research Assistant Professor in the Department of Electrical and Electronic Engineering at the University of Hong Kong), and Zheng Zhilong, a 2022 PhD student at the School of Vehicle and Mobility, are co-first authors of the paper. Professor Shengbo Eben Li is the corresponding author. Collaborators also include Masayoshi Tomizuka, a member of the U.S. National Academy of Engineering, from the University of California, Berkeley, and Assistant Professor Changliu Liu from Carnegie Mellon University.
This research was supported by the National Natural Science Foundation of China and the Beijing Natural Science Foundation.
Paper Link:https://doi.org/10.1108/FTSYS-03-2026-001