
Learning to control D1 means coordinating leg posture and wheel motion within a single policy. At Direct Drive Tech, we support this work through a public MJLab training package and a separate ROS 2 control workspace. The MJLab package includes the robot model, velocity-tracking tasks, configurable training settings, checkpoints, and playback; related Direct Drive resources document policy export and ONNX execution.
Together, these resources give reinforcement learning robotics teams a staged route from wheel–leg coordination and terrain-control studies to ROS 2 controller integration. The MJLab-to-deployment boundary still requires a validated model-conversion workflow and matching observation, action, and runtime definitions.
Coordinating Leg Positions and Wheel Speeds
The D1 robot platform has a wheel-legged quadruped configuration represented in the training model. Each leg contributes three joint-position actions and one wheel-velocity action. Across the robot, that gives the policy 12 leg-position targets and four wheel-speed targets.
Because the leg-position and wheel-speed channels use separate action scales, researchers can test how commanded speed, turning rate, or terrain changes wheel–leg coordination without treating D1 as a conventional quadruped.
The policy receives body angular velocity, projected gravity, movement commands, leg-joint position information, joint velocities, and the previous action. A 10-frame observation history supplies recent context alongside the current state. That creates a useful basis for studying how movement history affects control, using the robot’s own motion and joint information.
Training exposes separate critic, privileged-information, and terrain-scanner observation groups, while the policy group contains the observations and history used by the actor. This separation distinguishes simulation-only training signals from inputs intended for policy inference.
What D1 Provides for Reinforcement Learning Robotics Research
Our public MJLab training package combines a D1 MuJoCo model with NP3O, a constrained PPO implementation incorporating BarlowTwins representation learning. The training environment targets Python 3.11 and MJLab 1.4.0; the repository recommends a CUDA version above 12.4 for training performance.
Two registered tasks provide the starting points: Mjlab-Velocity-Flat-D1 for flatter terrain mixes and Mjlab-Velocity-Rough-D1 for more varied surfaces. Researchers can select either task within the supplied training workflow according to the terrain conditions they want to study.
Rewards score the movement objective, while separate cost terms penalize selected joint-position, velocity, and torque-limit violations. Researchers can therefore compare command tracking with adherence to configured limits. Observations, commands, rewards, costs, termination conditions, and terrain curricula remain accessible in the code: reward changes shift the training emphasis, while cost changes adjust the penalty assigned to a limit violation.
Custom reward or observation designs can reuse the existing robot model and training loop. Task packages are registered and discovered automatically under the configuration directory, allowing each experiment to have its own task identity without changing the main training script.
Parallel training defaults to 4,096 environments, with the count adjustable through the launch options. Saved checkpoints, TensorBoard logs, and resumed training support comparisons across runs. During evaluation, the playback tool loads a selected checkpoint for visual inspection. Turning, posture changes, and loss of balance can then be examined alongside the training curves. The default play configuration disables observation corruption, push events, gain randomization, and other training-time domain randomization; disturbance testing therefore requires a separate evaluation configuration.
Flat and Rough Tasks Train on Different Terrain Mixes
Our Flat task uses generated terrain containing random-uniform height variations, waves, and slopes. Despite its name, it already exposes the robot to surface changes. Its relatively gentle terrain settings make it a useful baseline for examining velocity tracking while the ground varies beneath the wheels.
The Rough task expands this to seven terrain types: flat ground, pyramid stairs, inverted pyramid stairs, slopes, inverted slopes, random rough ground, and waves. A terrain curriculum changes the difficulty encountered during training. Step dimensions, slope ranges, and surface variation remain editable, allowing a study to focus on a particular transition or obstacle pattern. This supports experiments that distinguish smooth-surface control from movement across steps and uneven ground.
Both configurations also include changes to friction, body mass and inertia, center-of-mass position, and control gains. Observation noise and push events add further variations. These settings let a study isolate a question such as how a policy responds when traction changes or a disturbance interrupts commanded movement.
Evaluate checkpoints on the same terrain mix and report results by surface type. This prevents strong behavior on one surface from hiding failures on another. Record terrain difficulty, command range, and enabled disturbances with each result so the test conditions remain visible.
The D1 Sim-to-Real Path: From MJLab Policy to Robot Controller
A practical five-gate validation workflow is:
- Train under defined command, terrain, and randomization settings.
- Evaluate the selected checkpoint in the default MJLab playback configuration, then use a separate evaluation configuration for disturbance cases.
- Export the selected actor and verify its required inputs and outputs.
- Run the exported ONNX policy through the ROS 2 control workspace in simulation; verify joint order, scaling, history, default pose, and timing.
- Start controlled hardware trials with standing, stopping, and low-speed commands before expanding the test envelope.
The first two gates remain inside MJLab. The later stages require a validated checkpoint-conversion method and consistent input and output definitions across export, ROS 2 integration, and hardware execution.
The public d1_mjlab documentation covers training, checkpoint resume, and playback. A related Direct Drive Isaac Lab workflow documents TorchScript and ONNX export with separate current-observation and history inputs, while the ROS 2 controller consumes ONNX policies. Confirm the selected MJLab checkpoint’s conversion method, network architecture, tensor order, and normalization before treating these repositories as one deployment chain.
Alongside the training package, we provide a ROS 2 control workspace with ONNX inference, finite-state-machine control, and a hardware bridge. D1’s onboard environment uses an NVIDIA Jetson Orin NX 8GB with Ubuntu 22.04 and ROS 2 support. The control workspace supplies the connection between a policy’s output and the robot’s control interfaces.
The public D1 ROS 2 control workspace includes two example policy configurations. Both use 57 observations, a 10-frame history, and 16 action outputs, but they apply different command ranges, action scales, default joint positions, and control gains. This shows why deployment involves pairing a policy with its corresponding controller settings rather than replacing only the model file.
Before deployment, match the joint order, observation scaling, default pose, leg-position and wheel-speed targets, control timing, policy configuration, and installed D1 ROS 2 software branch. A wheel-speed output sent through the wrong joint mapping cannot reproduce the learned motion.
Running the exported network through the ROS 2 simulation path can reveal mismatches in observation construction, history updates, joint mapping, and action handling before hardware testing. This is an integration stage, not a turnkey transfer. Review the current repository instructions before selecting a simulator: they require extra model setup for D1 in MuJoCo and presently note a D1 limitation in the Gazebo path.
Keeping commands and test surfaces consistent helps isolate faults: an issue that appears only after controller integration warrants checking the runtime path before retraining the policy.
Research Questions That Can Start with the Existing Tasks
Velocity tracking and wheel–leg coordination have direct starting points in the registered tasks. A study can compare forward-motion and turning accuracy while changing reward weights or action settings. Tracking error, posture variation, and smoothness during command transitions help explain what changes in the behavior, beyond the aggregate reward.
Terrain adaptation and robustness build on the generated surfaces and randomization settings. One experiment could hold the terrain fixed while changing friction; another could compare performance across the seven Rough terrain types. Completion rate, falls, and tracking error provide different views of the result. Keeping the test conditions fixed makes the comparison more useful.
Constraint-aware control and observation-history studies draw on the cost terms and history-based policy input. Changing cost weights allows a comparison between command tracking and limit violations. A history-length experiment requires matching changes to the environment and policy configuration, then comparing runs under the same conditions. These questions extend the supplied locomotion setup through targeted changes to training settings or model inputs.
The existing package is scoped to wheel-legged quadruped locomotion. Research on wheel-legged bipedal balancing or other objectives requires a corresponding model and training environment, plus new observations, actions, rewards, and controller integration.
Is D1 the Right Platform for Your Research?
Use three questions to decide whether the existing D1 research stack matches the project.
- Does the existing stack cover the research question? D1 is most directly aligned with wheel-and-leg locomotion, velocity-tracking objectives, Flat and Rough terrain studies, constraint tuning, and ONNX-to-ROS 2 integration.
- What requires additional development? If the project depends on an embodiment, sensor set, objective, or runtime outside the supplied resources, budget for the corresponding model, environment, observation and action design, reward or cost structure, and controller work. The public repositories do not guarantee real-robot transfer or benchmark performance.
- What must be confirmed before procurement? Confirm the target D1 configuration, terrain and motion envelope, required sensing, control frequency, the availability of a validated MJLab-to-ONNX conversion for the selected checkpoint, the D1 ROS 2 software version and matching controller branch, and the validation path from simulation to controlled hardware trials.
The D1 tutorials and development documentation connect the public training resources with robot operation and onboard development. Before hardware testing, define the motion-control question and confirm which software, sensing, and validation tasks the research team will handle and which require Direct Drive Tech support.
To discuss a specific setup, contact our team with the motion-control objective and integration requirements. That discussion can establish the appropriate D1 configuration and development-support scope for the project.
FAQs About D1 for Reinforcement Learning Robotics
Is D1 a turnkey sim-to-real reinforcement learning package?
No. The public stack covers MJLab training, Flat and Rough quadruped tasks, checkpoint evaluation, policy export, and ROS 2 inference and control. Teams must still validate observation and action mappings, control timing, software versions, safety procedures, and hardware behavior.
Which research projects can start from the existing D1 MJLab tasks?
The current tasks support wheel–leg coordination, velocity tracking, terrain adaptation, disturbance response, constraint-aware control, and observation-history studies. They cover the wheeled quadruped; biped research, added sensing, or non-locomotion objectives require additional model, task, and controller development.
Does successful playback or ROS 2 simulation mean a policy is ready for hardware?
No. Simulation helps identify observation, history, joint order, scaling, and action-handling mismatches, but it does not guarantee hardware performance. Begin real-robot tests with standing, stopping, and low-speed commands under controlled conditions.
What should a lab evaluate beyond the public repositories before buying D1?
Plan for integration time, controlled hardware testing, safety procedures, instrumentation and logging, and any custom sensing or task logic. Confirm which items are included, which the lab must develop, and which require Direct Drive Tech support.

