Skip to content

Training#

This page tells you how to train, control, watch, evaluate and export a driving policy for the Field Scouting Task. The trainer is PPO (Schulman et al. 2017) in the style of CleanRL (Huang et al. 2022). The file is Learning/acres_learn/adapters/ppo/train.py. The simulator is ACRES Core: 1024 vehicles step in worker processes on the CPU, and the network runs on the GPU.

Status

This page describes the trainer that exists. It trains the flat policy. The flat policy is the negative baseline (Results). The residual training of the Framework is specified, not built. The section Residual Training lists its planned changes.

Configurations#

The configurations are in Learning/acres_learn/configs/.

Configuration State Purpose Observation
ppo_deployable.json Exists; the default A flat policy for the vehicle Only quantities that the vehicle measures; no soil water
ppo_scouting_v1.json Exists A flat policy with the observation of the first specification With the soil-water channel
ppo_residual.json Specified, not built The residual of the Framework The polar_log path encoding

Quick Start#

  1. Build ACRES Core: PYTHON=~/miniconda3/envs/torchenv/bin/python Core/build.sh.
  2. Open a tmux session: tmux new -s ppo. Use tmux attach -t ppo if the session exists.
  3. Go to the repository: cd ~/Codes/ACRES.
  4. Start the run: Learning/Scripts/ppo.sh train scout1.
  5. Detach from tmux with Ctrl-b d. The run continues.

Learning/Scripts/ppo.sh activates the torchenv conda environment. It puts Learning and Core/Build on the Python path. Then it runs python -m acres_learn.adapters.ppo.train --run scout1 in the folder Learning/.

Trainer Options#

Option Function
--run <name> The run name. The run folder is Acres/Saved/Training/<name>
--config <name or file> The configuration of a new run. The default is ppo_deployable.json
--set section.key=value Changes one configuration value. The key must exist. You can repeat the option
--init-from <checkpoint> A new run starts from the model, the optimiser and the normalisers of a checkpoint
--device <cuda or cpu> The device. The default is CUDA if it is available
--no-keys The trainer does not read keys from the terminal
--max-env-steps <n> The trainer stops after this number of environment steps
--max-minutes <n> The trainer stops after this training time
Learning/Scripts/ppo.sh train scout1 --set env.num_envs=512 --set ppo.learning_rate=1e-4
Learning/Scripts/ppo.sh train teacher1 --config ppo_scouting_v1
Learning/Scripts/ppo.sh train fork1 --init-from /abs/path/ckpt_000000400.pt --set curriculum.start_stage=1
  • The first start writes the configuration to the run folder. Later starts of the same run use that copy.
  • --set applies on top of the saved copy. The trainer saves the result.
  • Use an absolute path for --init-from. The trainer runs in the folder Learning/.
  • A fork starts at the stage curriculum.start_stage. Its counters and its episode window start new.

Pause, Resume and Stop#

Use these commands from any terminal.

Learning/Scripts/ppo.sh ctl scout1 pause      # saves a checkpoint, then waits
Learning/Scripts/ppo.sh ctl scout1 resume
Learning/Scripts/ppo.sh ctl scout1 stop       # saves a checkpoint and exits
Learning/Scripts/ppo.sh ctl scout1 status     # stage, steps, rates, recent episodes, last validation
Learning/Scripts/ppo.sh ctl --list            # all runs and their states

In the terminal of the trainer, these keys do the same: p pause, r resume, q save and quit, s status. Ctrl-C saves a checkpoint and exits. A second Ctrl-C exits immediately without a checkpoint.

Pause. A paused trainer waits between two environment steps. The worker processes block on their pipes. The CPU load is 0.05 % of one core.

The trainer releases the cached GPU memory and holds approximately 0.8 GB. The RAM stays allocated (approximately 7.4 GB for 1024 environments). You can run the simulator during a pause.

Start again. After a stop, a crash or a reboot, use the same start command. The trainer resumes from the latest checkpoint. The checkpoint holds the model, the optimiser, the normalisers, the stage, the window of recent episodes, the random generators and the step counters. The environments start new episodes. Only one trainer can own a run folder (lock).

Outputs#

Each run has a folder Acres/Saved/Training/<run>/. The folder is not in git.

Path Content
config.json The configuration of the run
checkpoints/ckpt_<update>.pt Checkpoints: one for each 25 updates, and one at each pause, stop and stage change. The last 5 stay
checkpoints/best.pt The checkpoint with the best validation score: stage, success, coverage, return, in that order
checkpoints/best_stage<k>.pt The best checkpoint of stage k
tb/ TensorBoard events
metrics.csv One row for each update
status.json, control.json, lock The control files (adapters/ppo/control.py)
train.log All lines that the trainer printed
episodes/ Episode logs that watch wrote

The trainer caches the transits of stage 0 in Acres/Saved/Learning/ppo/.

Watch a Run#

TensorBoard. Run Learning/Scripts/ppo.sh tensorboard scout1. Then open http://localhost:6006. Without a run name, TensorBoard shows all runs.

Group Scalars
episode/ Return, success, coverage, crop outside the edge band (m²), edge-band crop (m²), length
episode_terms/ Each reward term, summed for each episode
reward_step/ Each reward term for each step, averaged over the rollout
endings/ The share of episodes for each ending: success, collision, stuck, deep in crop, left the map, lost, stalled, timeout
window/ Success and return over the last 500 episodes
losses/ Policy loss, value loss, entropy, approximate KL, clip fraction, explained variance, gradient norm, KL stop
charts/ Learning rate, action standard deviation, reward scale, mean value, epochs, steps for each second
time/ Seconds for the rollout, the environments and the update
curriculum/ The stage, the steps in the stage, the value warm-up
rehearsal/ Success, collisions and return of the rehearsed episodes of each earlier stage
eval/ The validation missions: success, return, coverage, crop, time

Figure. Run this command in Learning/ with the Python path of ppo.sh: python -m acres_learn.adapters.ppo.plot scout1 --out curves.png. The figure shows the return, the success rate, the steps for each second and the explained variance.

Status. ctl scout1 status prints the stage, the rates, the estimated times, the recent episodes, the last validation and the last checkpoint.

Algorithm#

The values are those of the two configurations that exist.

Setting Value Reason
Environments 1024 in 8 worker processes, 16 Core threads The task bookkeeping in Python runs on 8 cores
Rollout 64 steps (6.4 s), 65,536 samples for each update
Discount, GAE \(\gamma = 0.995\), \(\lambda = 0.95\) A horizon of 200 steps (20 s)
Update 4 epochs of 8 minibatches of 8192 samples, Adam
Learning rate \(3 \times 10^{-4}\) to \(3 \times 10^{-5}\), linear over \(2 \times 10^9\) steps
Losses Clipped surrogate (0.2); clipped value loss (0.2) times 0.5; entropy bonus 0.001; gradient norm clipped at 0.5
Normalisation Advantages for each minibatch. Running mean and variance of the vector inputs, clipped at ±10. Rewards divided by the running deviation of the discounted return, clipped at ±10 The rewards are from −100 to +1 for each step
Policy A Gaussian with a tanh squash and a state-independent log standard deviation
Initial deviation 0.6 on the curvature, 0.25 on the speed See below
Initial mean speed Action 0.46, which is 2.8 m/s See below
ppo.kl_stop 1.0 in ppo_deployable.json An update stops at the first minibatch with an approximate KL above this value
Validation Each 100 updates

Network (network.py).

  • A CNN for the map crop: 4 or 5 channels of 64 × 64 cells, four stride-2 convolutions (32, 64, 64, 64 channels), then 256 features.
  • 1-D convolutions with circular padding for the 360 LiDAR beams (16, 32, 32 channels), then 128 features.
  • An MLP for the proprioception and the path: 130 inputs, two layers of 256.
  • Two trunks (640 to 512 to 256): one for the mean of the actor and one for the value of the critic.
  • The convolutions run in bfloat16 on the GPU. The other layers run in float32.

The environments do not send the map crop. Each environment sends the pose, and the GPU cuts the crop from the ground-class raster (CropRaster). The result is equal to observations.map_crop (tested). The rollout keeps the crop as class codes: 4 KiB for each step and not 80 KiB.

Squashing. The buffer keeps the Gaussian sample \(u\). The environment gets \(a = \tanh u\). The log-probability is exact: \(\log \pi(a) = \log \mathcal{N}(u) - \sum \log(1 - \tanh^2 u)\). The entropy bonus is the entropy of the Gaussian.

Time limits. A time limit truncates an episode. It does not terminate the episode. The trainer bootstraps its last step from the value of its final observation. A success or a failure terminates the episode without a bootstrap (algorithm.gae, tested on a hand-worked case).

Lost episodes end. The tracking term \(-0.5 \max(0, \lvert e_y \rvert - h)^2\) has no lower bound. With env.lost_m (10 m), an episode ends as the failure "lost" when the vehicle is more than 10 m outside its corridor. This is a training aid. The evaluation does not use it.

The speed starts forwards. A speed action below −0.13 is a request for reverse. The gear logic then stops the vehicle, shifts and holds for 1.1 s. The first initial values were a mean of 0 and a deviation of 0.6. Then one quarter of the steps asked for reverse, and the vehicles almost did not move. Thus the mean starts at 2.8 m/s and the deviation at 0.25.

NaN actions. The environment reads a NaN action as a stop. It never stores a NaN as the previous action.

Rewards. The rewards are those of the task (tasks/scouting/rewards.py). The log has each term separately.

Inputs on the Vehicle#

ppo_deployable.json trains a policy that sees only quantities that the vehicle measures.

Input Source on the Polaris
Speed vehicle/vehicle_velocity (ds_dbw VehicleVelocity) or the OxTS velocity (oxts/velocity)
Yaw rate, lateral acceleration, pitch, roll The OxTS RT3000 (oxts/imu)
Steering wheel angle vehicle/steering/report (SteeringReport.steering_wheel_angle)
Four wheel speeds vehicle/wheel_speeds (rad/s) times the tyre radius 0.343 m
Gear vehicle/gear/report
Previous action The last command of the program
Path The plan and the reference path on the map products, and the OxTS pose in the map frame
Map crop: crop, edge band, lane or verge, obstacle The map products (ground_classes.u8) around the pose
LiDAR, 360 beams The Helios points within 1° of the horizontal, the nearest point for each degree

The soil-water channel of the first specification is privileged. The vehicle does not measure soil water. observation.soil_water: false removes it. The crop then has four channels.

Randomisation#

The sensor noise (tasks/scouting/sensor_noise.py) changes only the inputs of the policy. It does not change the physics, the rewards or the task bookkeeping (tested). Each episode draws its errors.

Error Default Reason
GNSS/INS solution: RTK fixed, 80 % Position 2 cm noise and 3 cm bias; heading 0.05° and 0.1° The OxTS holds centimetres with RTK
GNSS/INS solution: float, 15 % Position 5 cm and 0.5 m; heading 0.1° and 0.3° Decimetres without a fixed solution
GNSS/INS solution: none, 5 % Position 10 cm and 1.5 m; heading 0.2° and 1° Metres without corrections. The bias drifts with a time constant of 60 s
Map offset Up to 0.6 m and 0.5°, about the start of the episode The surveyed map and the vehicle frame do not align exactly
Latency 0 to 60 ms Delays of the sensors and the software
Proprioception noise 0.02 m/s speed and wheel speeds, 0.005 rad/s yaw rate, 0.05 m/s² lateral acceleration, 0.5° steering wheel, 0.3° pitch and roll Resolution and noise of the reports
LiDAR 3 cm range noise, 2 % of the beams dropped Range noise and missing returns of the Helios
Drive-by-wire Steering centre ±3°; wheel speed ±10 %; each fitted parameter of polaris.json within its fitted deviation, clipped at 2 sd The vehicle on a different day
Soil water (last stage only) Uniform from 0.5 to 1.3 Dry to almost saturated soil
Start heading (last stage only) Within 5° of the heading of H

A stage sets its share of the sensor errors with noise_scale. It sets its share of the drive-by-wire errors with dbw_scale.

Flat-Policy Curriculum#

The flat policy trains in stages. A stage moves on when its criterion is true. A stage can also end on a step budget (max_env_steps).

ppo_scouting_v1.json:

Stage Missions Condition to Move On
0 goto 2000 transits of 50 to 300 m on lanes, dry. One fifth starts at H 95 % of the last 2000 episodes reach the target, after 5 M steps
1 scout_lane_bound Single-field missions of the lane-bound training fields 95 % complete and mean crop outside the edge band below 1 m², over 200 episodes, after 20 M steps
2 scout_all Single-field missions of all 43 training fields The same over 400 episodes, after 50 M steps
3 short_missions 2 to 3 training fields with the optimal plan 90 % complete over 400 episodes, after 50 M steps
4 missions 1 to 8 training fields; dry, wet and hardware shift; randomised Last stage

The lane-bound training fields are F11, F13, F15, F17, F18, F33, F58 and F59. At least 90 % of their loops is on lane or verge.

ppo_deployable.json:

Stage Missions Condition to Move On Budget Noise, Drive-by-Wire Scale
0 goto Transits as above; 40 % start at H 95 % over 2000 episodes, after 5 M steps None 0.5, 0
1 scout_loops The loops of the lane-bound fields alone, from the straight next to the entry 70 % over 400 episodes, after 10 M steps 150 M 1.0, 0.5
2 scout_lane_bound Single-field missions of the lane-bound fields 70 % and crop below 10 m², after 20 M steps 300 M 1.0, 0.5
3 scout_all Single-field missions of all training fields 70 % and crop below 15 m², after 50 M steps 400 M 1.0, 0.5
4 short_missions 2 to 3 training fields 90 % over 400 episodes, after 50 M steps None 1.0, 1.0
5 missions 1 to 8 training fields; dry, wet and hardware shift Last stage None 1.0, 1.0

Stages 1 to 3 rehearse earlier stages. Stage 1 draws 20 % of its episodes as transits from H. Stage 2 draws 15 % loops and 10 % transits. Stage 3 draws 10 % loops and 5 % transits.

Plans. A single-field mission uses the frozen plan of configs/scouting_reference_v1.json. A mission with more fields uses the exact dynamic programme on the dry cost tables. The training draws only fields of the training split (configs/splits_v1.json).

Validation. Each 100 updates, the policy drives the validation missions of the stage with its mean action.

Stage Type Validation Missions
Transits 64 transits from a different seed
Loops The loop of each of the 6 validation fields
Single field Each of the 6 validation fields alone
2 to 3 fields The validation fields in pairs and triples
Last stage Singles and pairs under each condition set

Curriculum Options#

These options came from the diagnoses of the flat runs (Results).

Option Function ppo_deployable.json
env.launch_hold Holds the speed command from a decrease while the vehicle stands (spec.LaunchHold). The exported policy carries the option On
env.path_through The path input continues into the path of the next phase On
env.path_lookahead With path_through: the path input of a lone transit continues along the leg that follows Off
env.stall_s, env.stall_penalty Ends an episode as "stalled" after the vehicle stands for this time Off
env.lost_m Ends an episode as "lost" at this distance outside the corridor 10 m
env.rewards "v1" (first version) or "bounded" (failures never pay) "v1"
missions: loop, loop_start: straight A stage of loops alone. The episode starts on the straight next to the entry corner Stage 1
Stage rehearsal A share of the episodes comes from an earlier stage. The move-on test does not count them Stages 1 to 3
Stage entry_share A share of the transits is the GoTo from H to a field entry 0
Stage home_share A share of the transits starts at H 0.4
curriculum.value_warmup_env_steps The first steps of a stage train only the critic 1 M
ppo.kl_stop Stops an update at a large KL 1.0

The configuration check rejects env.rewards = "bounded" when a failure costs less than the floor divided by \(1 - \gamma\). It rejects env.path_lookahead without env.path_through.

Watch an Episode in Unreal#

Learning/Scripts/ppo.sh watch scout1                               # latest checkpoint, first validation mission
Learning/Scripts/ppo.sh watch scout1 --checkpoint best --field F42
Learning/Scripts/ppo.sh watch scout1 --stage 0 --index 3           # the fourth validation transit
  1. watch drives the episode in ACRES Core with the mean action.
  2. It writes an MCAP episode log to Acres/Saved/Training/<run>/episodes/. --out selects a different folder.
  3. It prints the replay command for the game: Packaged/Linux/Acres.sh -VehicleDemo -Vehicle=polaris -EpisodeReplay=<file.mcap>.
  4. With --render, the packaged game renders an MP4 next to the log. Run only one game at a time.

Pause the trainer before you start the game. The game then has the CPU and the GPU.

Evaluation#

A checkpoint drives the evaluation suite in place of the reference driver. The plans and the scores are the same (eval.run with --driver). Thus the tables agree with Scouting Baselines.

Learning/Scripts/ppo.sh eval scout1 --checkpoint best --suite --out ~/ppo-eval/suite          # the 600 runs
Learning/Scripts/ppo.sh eval scout1 --single-field --split validation --out ~/ppo-eval/val    # six fields alone

\(J^*\) and \(T_\text{ref}\) are the frozen values of the reference driver (configs/scouting_reference_v1.json). The output folder gets report.md, report.json and results.json.

GoTo from H. This command drives the GoTo from H to the entry of each field. Run it in Learning/ with the Python path of ppo.sh.

python -m acres_learn.eval.goto --driver reference --driver ppo:goto2@best --runs 4 --noise both --out <dir>
Option Function
--driver reference, or ppo:<run or checkpoint>[@best, @latest or @best_stage<k>]. You can repeat the option
--fields all, train, val, test or a list
--runs Runs for each field
--noise off, on or both. With noise, the policy samples its actions
--lookahead policy, on or off: the path input past the entry

Export#

The vehicle PC has no torch. The export writes the actor and its normaliser to one .npz file.

Learning/Scripts/ppo.sh export scout1 --checkpoint best --out Acres/Saved/Training/scout1/policy.npz
Piece (adapters/ppo/deploy.py) Function
NumpyPolicy The actor in numpy. It reproduces the float32 actions of the torch network to \(10^{-5}\) (tested)
VehicleObservation The observation from the vehicle quantities, with the same code as training
planar_scan The Helios cloud to 360 beams of one degree
CommandMapper An action to the ds_dbw SteeringCmd, UlcCmd and GearCmd

| make_node | A minimal rclpy node at 10 Hz. It is a skeleton. The deployment program is the full front end |

The file carries the options of the training: launch_hold, path_through, path_lookahead and soil_water. Deployment uses the file.

This example shows the call sequence for one mission:

from acres_learn.tasks.scouting.map_products import ScoutingMap
from acres_learn.tasks.scouting import routes
from acres_learn.adapters.ppo.deploy import NumpyPolicy, VehicleObservation, CommandMapper, planar_scan

m = ScoutingMap.load()
policy = NumpyPolicy.load("scout1.npz", m=m)                    # the actor and its normaliser
legs = routes.mission_legs(routes.Network(m), plan)             # the plan, e.g. [("F39", 0, True)]
obs = VehicleObservation(m, legs, ["F39"], t_ref_s, path_through=policy.path_through)
mapper = CommandMapper(policy.launch_hold)                      # the gear logic of the vehicle
# each 0.1 s:
o = obs.observe(pose=(east_m, north_m, yaw_rad), speed_mps=..., yaw_rate_radps=..., lateral_accel_mps2=...,
                steering_wheel_rad=..., pitch_rad=..., roll_rad=..., wheel_speeds_mps=[fl, fr, rl, rr], gear=...,
                lidar_ranges=planar_scan(points_xyz))           # a batch of one
action = policy(o)[0]                                           # in [-1, 1]^2, the mean action
obs.acted(action)
cmds = mapper.commands(action, speed_mps)   # {"steering_cmd": {...}, "ulc_cmd": {...}, "gear_cmd"?: {...}}

Training runs the convolutions in bfloat16. This moves the actions by a median of \(5 \times 10^{-4}\). Use --set network.amp_bf16=false to train in float32. That is approximately 30 % slower.

Throughput#

Measured on the development machine (16 threads, RTX 5060 Ti 16 GB) in stage 0.

Environments Steps for Each Second Environments / Update Time
256 11,500 0.6 s / 0.7 s
512 12,000 1.1 s / 1.4 s
1024 (default) 12,300 2.0 s / 2.7 s
2048 (16 workers) 12,800 3.6 s / 5.5 s

The environments alone step 35,500 times a second with 1024 environments in 8 workers. The update is approximately half of the time.

Residual Training#

State: specified, not built. The Framework gives the rules. The planned changes of the trainer are:

Change Rule in the Framework
One stage of segment episodes; no curriculum Segment Episodes
The policy output is a residual on the base command Residual Law
The path input uses the polar_log encoding Path Encoding
The critic gets two more inputs than the actor Critic Group
The launch hold runs before the shield. The shield filters each command in the environment, before the gear logic Order in a Control Step, Shield
The reward variant residual. It has a shield term of 0.5 after the bounded map, a shaping term with \(c_\Phi = 0.55\), failures at −1200 and a collision at −1500 Reward
A truncation bootstraps with the critic of the shaped problem. The potential at a truncation is not 0 Shaping Term
rewards.gamma is equal to ppo.gamma Reward
The code of the residual law, the shield, the reward and the shaping term passes the golden vectors Golden-Vector Contract
A critic-only warm-up at the start of the run Reasons
A whole-mission evaluator for each gain \(\alpha\) Evaluation Protocol

The trainer loop, the control tools, the encoders and the export stay as they are.

Limitations#

  • A resume starts new episodes. The state of ACRES Core in the middle of an episode is not in the checkpoint.
  • The update is half of the time. The GPU update and the rollout alternate.
  • Workers and parked vehicles are not in ACRES Core. The training does not randomise them.
  • The soil water is one value for each episode.
  • The planar LiDAR on the vehicle is the Helios rings within 1° of the horizontal. ACRES Core casts one level ray for each degree. A comparison on a recorded drive is open.