See also

This tutorial is the source of truth for the isaaclab-randomizing-with-events agent skill (skills/user/domain-randomization-events/). When you change this page, update the skill so agent guidance stays in sync. See Agent Skills.

Creating a Manager-Based Base Environment#

Environments bring together different aspects of the simulation such as the scene, observations and actions spaces, reset events etc. to create a coherent interface for various applications. In Isaac Lab, manager-based environments are implemented as envs.ManagerBasedEnv and envs.ManagerBasedRLEnv classes. The two classes are very similar, but envs.ManagerBasedRLEnv is useful for reinforcement learning tasks and contains rewards, terminations, curriculum and command generation. The envs.ManagerBasedEnv class is useful for traditional robot control and doesn’t contain rewards and terminations.

In this tutorial, we will look at the base class envs.ManagerBasedEnv and its corresponding configuration class envs.ManagerBasedEnvCfg for the manager-based workflow. We will use the cartpole environment from earlier to illustrate the different components in creating a new envs.ManagerBasedEnv environment.

The Code#

The tutorial corresponds to the create_cartpole_base_env script in the scripts/tutorials/03_envs directory.

Code for create_cartpole_base_env.py
  1# Copyright (c) 2022-2026, The Isaac Lab Project Developers (https://github.com/isaac-sim/IsaacLab/blob/main/CONTRIBUTORS.md).
  2# All rights reserved.
  3#
  4# SPDX-License-Identifier: BSD-3-Clause
  5
  6"""
  7This script demonstrates how to create a simple environment with a cartpole. It combines the concepts of
  8scene, action, observation and event managers to create an environment.
  9
 10.. code-block:: bash
 11
 12    uv run python scripts/tutorials/03_envs/create_cartpole_base_env.py --num_envs 32
 13
 14"""
 15
 16"""Parse the command-line arguments first."""
 17
 18
 19import argparse
 20
 21from isaaclab.app import add_launcher_args, launch_simulation
 22
 23# add argparse arguments
 24parser = argparse.ArgumentParser(description="Tutorial on creating a cartpole base environment.")
 25parser.add_argument("--num_envs", type=int, default=16, help="Number of environments to spawn.")
 26
 27# append simulation launcher cli args
 28add_launcher_args(parser)
 29# parse the arguments
 30args_cli = parser.parse_args()
 31
 32"""Rest everything follows."""
 33
 34import math
 35
 36import torch
 37
 38import isaaclab.envs.mdp as mdp
 39from isaaclab.envs import ManagerBasedEnvCfg
 40from isaaclab.managers import EventTermCfg as EventTerm
 41from isaaclab.managers import ObservationGroupCfg as ObsGroup
 42from isaaclab.managers import ObservationTermCfg as ObsTerm
 43from isaaclab.managers import SceneEntityCfg
 44from isaaclab.utils import configclass, instantiate
 45from isaaclab.visualizers import VisualizerCfg
 46
 47from isaaclab_tasks.core.cartpole.cartpole_manager_env_cfg import CartpoleSceneCfg
 48
 49
 50@configclass
 51class ActionsCfg:
 52    """Action specifications for the environment."""
 53
 54    joint_efforts = mdp.JointEffortActionCfg(asset_name="robot", joint_names=["slider_to_cart"], scale=5.0)
 55
 56
 57@configclass
 58class ObservationsCfg:
 59    """Observation specifications for the environment."""
 60
 61    @configclass
 62    class PolicyCfg(ObsGroup):
 63        """Observations for policy group."""
 64
 65        # observation terms (order preserved)
 66        joint_pos_rel = ObsTerm(func=mdp.joint_pos_rel)
 67        joint_vel_rel = ObsTerm(func=mdp.joint_vel_rel)
 68
 69        def __post_init__(self) -> None:
 70            self.enable_corruption = False
 71            self.concatenate_terms = True
 72
 73    # observation groups
 74    policy: PolicyCfg = PolicyCfg()
 75
 76
 77@configclass
 78class EventCfg:
 79    """Configuration for events."""
 80
 81    # on startup
 82    add_pole_mass = EventTerm(
 83        func=mdp.randomize_rigid_body_mass,
 84        mode="startup",
 85        params={
 86            "asset_cfg": SceneEntityCfg("robot", body_names=["pole"]),
 87            "mass_distribution_params": (0.1, 0.5),
 88            "operation": "add",
 89        },
 90    )
 91
 92    # on reset
 93    reset_cart_position = EventTerm(
 94        func=mdp.reset_joints_by_offset,
 95        mode="reset",
 96        params={
 97            "asset_cfg": SceneEntityCfg("robot", joint_names=["slider_to_cart"]),
 98            "position_range": (-1.0, 1.0),
 99            "velocity_range": (-0.1, 0.1),
100        },
101    )
102
103    reset_pole_position = EventTerm(
104        func=mdp.reset_joints_by_offset,
105        mode="reset",
106        params={
107            "asset_cfg": SceneEntityCfg("robot", joint_names=["cart_to_pole"]),
108            "position_range": (-0.125 * math.pi, 0.125 * math.pi),
109            "velocity_range": (-0.01 * math.pi, 0.01 * math.pi),
110        },
111    )
112
113
114@configclass
115class CartpoleEnvCfg(ManagerBasedEnvCfg):
116    """Configuration for the cartpole environment."""
117
118    # Scene settings
119    scene = CartpoleSceneCfg(num_envs=1024, env_spacing=2.5)
120    # Basic settings
121    observations = ObservationsCfg()
122    actions = ActionsCfg()
123    events = EventCfg()
124
125    def __post_init__(self):
126        """Post initialization."""
127        # viewer settings
128        self.sim.default_visualizer_cfg = VisualizerCfg(eye=(4.5, 0.0, 6.0), lookat=(0.0, 0.0, 2.0))
129        # step settings
130        self.decimation = 4  # env step every 4 sim steps: 200Hz / 4 = 50Hz
131        # simulation settings
132        self.sim.dt = 0.005  # sim step every 5ms: 200Hz
133
134
135def main():
136    """Main function."""
137    # parse the arguments
138    env_cfg = CartpoleEnvCfg()
139    env_cfg.scene.num_envs = args_cli.num_envs
140    env_cfg.sim.device = args_cli.device
141    # Launch the simulator runtime that the configuration needs
142    with launch_simulation(env_cfg, args_cli):
143        # setup base environment
144        env = instantiate(env_cfg)
145
146        # simulate physics
147        count = 0
148        while env.sim.is_running():
149            with torch.inference_mode():
150                # reset
151                if count % 300 == 0:
152                    count = 0
153                    env.reset()
154                    print("-" * 80)
155                    print("[INFO]: Resetting environment...")
156                # sample random actions
157                joint_efforts = torch.randn_like(env.action_manager.action)
158                # step the environment
159                obs, _ = env.step(joint_efforts)
160                # print current orientation of pole
161                print("[Env 0]: Pole joint: ", obs["policy"][0][1].item())
162                # update counter
163                count += 1
164
165        # close the environment
166        env.close()
167
168
169if __name__ == "__main__":
170    # run the main function
171    main()

The Code Explained#

The base class envs.ManagerBasedEnv wraps around many intricacies of the simulation interaction and provides a simple interface for the user to run the simulation and interact with it. It is composed of the following components:

By configuring these components, the user can create different variations of the same environment with minimal effort. In this tutorial, we will go through the different components of the envs.ManagerBasedEnv class and how to configure them to create a new environment.

Designing the scene#

The first step in creating a new environment is to configure its scene. For the cartpole environment, we will be using the scene from the previous tutorial. Thus, we omit the scene configuration here. For more details on how to configure a scene, see Using the Interactive Scene.

Defining actions#

In the previous tutorial, we directly input the action to the cartpole using robot.actuators.target_command.set_effort_index. In this tutorial, we will use the managers.ActionManager to handle the actions.

The action manager can comprise of multiple managers.ActionTerm. Each action term is responsible for applying control over a specific aspect of the environment. For instance, for robotic arm, we can have two action terms – one for controlling the joints of the arm, and the other for controlling the gripper. This composition allows the user to define different control schemes for different aspects of the environment.

In the cartpole environment, we want to control the force applied to the cart to balance the pole. Thus, we will create an action term that controls the force applied to the cart.

@configclass
class ActionsCfg:
    """Action specifications for the environment."""

    joint_efforts = mdp.JointEffortActionCfg(asset_name="robot", joint_names=["slider_to_cart"], scale=5.0)

Defining observations#

While the scene defines the state of the environment, the observations define the states that are observable by the agent. These observations are used by the agent to make decisions on what actions to take. In Isaac Lab, the observations are computed by the managers.ObservationManager class.

Similar to the action manager, the observation manager can comprise of multiple observation terms. These are further grouped into observation groups which are used to define different observation spaces for the environment. For instance, for hierarchical control, we may want to define two observation groups – one for the low level controller and the other for the high level controller. It is assumed that all the observation terms in a group have the same dimensions.

For this tutorial, we will only define one observation group named "policy". While not completely prescriptive, this group is a necessary requirement for various wrappers in Isaac Lab. We define a group by inheriting from the managers.ObservationGroupCfg class. This class collects different observation terms and help define common properties for the group, such as enabling noise corruption or concatenating the observations into a single tensor.

The individual terms are defined by inheriting from the managers.ObservationTermCfg class. This class takes in the managers.ObservationTermCfg.func that specifies the function or callable class that computes the observation for that term. It includes other parameters for defining the noise model, clipping, scaling, etc. However, we leave these parameters to their default values for this tutorial.

@configclass
class ObservationsCfg:
    """Observation specifications for the environment."""

    @configclass
    class PolicyCfg(ObsGroup):
        """Observations for policy group."""

        # observation terms (order preserved)
        joint_pos_rel = ObsTerm(func=mdp.joint_pos_rel)
        joint_vel_rel = ObsTerm(func=mdp.joint_vel_rel)

        def __post_init__(self) -> None:
            self.enable_corruption = False
            self.concatenate_terms = True

    # observation groups
    policy: PolicyCfg = PolicyCfg()

For a policy that needs a flattened history with all terms from each time step together, set history_length and history_order="time" on its managers.ObservationGroupCfg. The resulting observation has shape (num_envs, history_length * combined_term_dim) and can be reshaped to (num_envs, history_length, combined_term_dim). The default "term" order keeps each term’s history together; use it for policies trained with the existing layout.

Defining events#

At this point, we have defined the scene, actions and observations for the cartpole environment. The general idea for all these components is to define the configuration classes and then pass them to the corresponding managers. The event manager is no different.

The managers.EventManager class is responsible for events corresponding to changes in the simulation state. This includes resetting (or randomizing) the scene, randomizing physical properties (such as mass, friction, etc.), and varying visual properties (such as colors, textures, etc.). Each of these are specified through the managers.EventTermCfg class, which takes in the managers.EventTermCfg.func that specifies the function or callable class that performs the event.

Additionally, it expects the mode of the event. The mode specifies when the event term should be applied. It is possible to specify your own mode. For this, you’ll need to adapt the ManagerBasedEnv class. However, out of the box, Isaac Lab provides three commonly used modes:

  • "startup" - Event that takes place only once at environment startup.

  • "reset" - Event that occurs on environment termination and reset.

  • "interval" - Event that are executed at a given interval, i.e., periodically after a certain number of steps.

For this example, we define events that randomize the pole’s mass on startup. This is done only once since this operation is expensive and we don’t want to do it on every reset. We also create an event to randomize the initial joint state of the cartpole and the pole at every reset.

@configclass
class EventCfg:
    """Configuration for events."""

    # on startup
    add_pole_mass = EventTerm(
        func=mdp.randomize_rigid_body_mass,
        mode="startup",
        params={
            "asset_cfg": SceneEntityCfg("robot", body_names=["pole"]),
            "mass_distribution_params": (0.1, 0.5),
            "operation": "add",
        },
    )

    # on reset
    reset_cart_position = EventTerm(
        func=mdp.reset_joints_by_offset,
        mode="reset",
        params={
            "asset_cfg": SceneEntityCfg("robot", joint_names=["slider_to_cart"]),
            "position_range": (-1.0, 1.0),
            "velocity_range": (-0.1, 0.1),
        },
    )

    reset_pole_position = EventTerm(
        func=mdp.reset_joints_by_offset,
        mode="reset",
        params={
            "asset_cfg": SceneEntityCfg("robot", joint_names=["cart_to_pole"]),
            "position_range": (-0.125 * math.pi, 0.125 * math.pi),
            "velocity_range": (-0.01 * math.pi, 0.01 * math.pi),
        },
    )

Tying it all together#

Having defined the scene and manager configurations, we can now define the environment configuration through the envs.ManagerBasedEnvCfg class. This class takes in the scene, action, observation and event configurations.

In addition to these, it also takes in the envs.ManagerBasedEnvCfg.sim which defines the simulation parameters such as the timestep, gravity, etc. This is initialized to the default values, but can be modified as needed. We recommend doing so by defining the __post_init__() method in the envs.ManagerBasedEnvCfg class, which is called after the configuration is initialized.

@configclass
class CartpoleEnvCfg(ManagerBasedEnvCfg):
    """Configuration for the cartpole environment."""

    # Scene settings
    scene = CartpoleSceneCfg(num_envs=1024, env_spacing=2.5)
    # Basic settings
    observations = ObservationsCfg()
    actions = ActionsCfg()
    events = EventCfg()

    def __post_init__(self):
        """Post initialization."""
        # viewer settings
        self.sim.default_visualizer_cfg = VisualizerCfg(eye=(4.5, 0.0, 6.0), lookat=(0.0, 0.0, 2.0))
        # step settings
        self.decimation = 4  # env step every 4 sim steps: 200Hz / 4 = 50Hz
        # simulation settings
        self.sim.dt = 0.005  # sim step every 5ms: 200Hz

Running the simulation#

Lastly, we revisit the simulation execution loop. This is now much simpler since we have abstracted away most of the details into the environment configuration. We only need to call the envs.ManagerBasedEnv.reset() method to reset the environment and envs.ManagerBasedEnv.step() method to step the environment. Both these functions return the observation and an info dictionary which may contain additional information provided by the environment. These can be used by an agent for decision-making.

Before the environment is created, its configuration is passed to app.launch_simulation(), which launches the simulator runtime that the configuration needs and closes it when the with block exits. Inside this block, instantiate() constructs the environment from the class named by the configuration’s class_type attribute. The class is resolved lazily, so the environment class, which works on the USD stage, is only imported once the simulator runtime is running.

The envs.ManagerBasedEnv class does not have any notion of terminations since that concept is specific for episodic tasks. Thus, the user is responsible for defining the termination condition for the environment. In this tutorial, we reset the simulation at regular intervals.

def main():
    """Main function."""
    # parse the arguments
    env_cfg = CartpoleEnvCfg()
    env_cfg.scene.num_envs = args_cli.num_envs
    env_cfg.sim.device = args_cli.device
    # Launch the simulator runtime that the configuration needs
    with launch_simulation(env_cfg, args_cli):
        # setup base environment
        env = instantiate(env_cfg)

        # simulate physics
        count = 0
        while env.sim.is_running():
            with torch.inference_mode():
                # reset
                if count % 300 == 0:
                    count = 0
                    env.reset()
                    print("-" * 80)
                    print("[INFO]: Resetting environment...")
                # sample random actions
                joint_efforts = torch.randn_like(env.action_manager.action)
                # step the environment
                obs, _ = env.step(joint_efforts)
                # print current orientation of pole
                print("[Env 0]: Pole joint: ", obs["policy"][0][1].item())
                # update counter
                count += 1

        # close the environment
        env.close()

An important thing to note above is that the entire simulation loop is wrapped inside the torch.inference_mode() context manager. This is because the environment uses PyTorch operations under-the-hood and we want to ensure that the simulation is not slowed down by the overhead of PyTorch’s autograd engine and gradients are not computed for the simulation operations.

The Code Execution#

To run the base environment made in this tutorial, you can use the following command:

uv run python scripts/tutorials/03_envs/create_cartpole_base_env.py --num_envs 32 --viz kit
./isaaclab.sh -p scripts/tutorials/03_envs/create_cartpole_base_env.py --num_envs 32 --viz kit

This should open a stage with a ground plane, light source, and cartpoles. The simulation should be playing with random actions on the cartpole. Additionally, it opens a UI window on the bottom right corner of the screen named "Isaac Lab". This window contains different UI elements that can be used for debugging and visualization.

result of create_cartpole_base_env.py

To stop the simulation, you can either close the window, or press Ctrl+C in the terminal where you started the simulation.

In this tutorial, we learned about the different managers that help define a base environment. We include more examples of defining the base environment in the scripts/tutorials/03_envs directory. For completeness, they can be run using the following commands:

# Floating cube environment with custom action term for PD control
uv run python scripts/tutorials/03_envs/create_cube_base_env.py --num_envs 32 --viz kit

# Quadrupedal locomotion environment with a policy that interacts with the environment
uv run python scripts/tutorials/03_envs/create_quadruped_base_env.py --num_envs 32 --viz kit
# Floating cube environment with custom action term for PD control
./isaaclab.sh -p scripts/tutorials/03_envs/create_cube_base_env.py --num_envs 32 --viz kit

# Quadrupedal locomotion environment with a policy that interacts with the environment
./isaaclab.sh -p scripts/tutorials/03_envs/create_quadruped_base_env.py --num_envs 32 --viz kit

In the following tutorial, we will look at the envs.ManagerBasedRLEnv class and how to use it to create a Markovian Decision Process (MDP).