Training#
This page tells you how to train, control, watch, evaluate and export a driving policy for the Field Scouting
Task. The trainer is PPO (Schulman et al. 2017) in the style of CleanRL (Huang et al. 2022). The
file is Learning/acres_learn/adapters/ppo/train.py. The simulator is ACRES Core: 1024 vehicles step in worker
processes on the CPU, and the network runs on the GPU.
Status
This page describes the trainer that exists. It trains the flat policy. The flat policy is the negative baseline (Results). The residual training of the Framework is specified, not built. The section Residual Training lists its planned changes.
Configurations#
The configurations are in Learning/acres_learn/configs/.
| Configuration | State | Purpose | Observation |
|---|---|---|---|
ppo_deployable.json |
Exists; the default | A flat policy for the vehicle | Only quantities that the vehicle measures; no soil water |
ppo_scouting_v1.json |
Exists | A flat policy with the observation of the first specification | With the soil-water channel |
ppo_residual.json |
Specified, not built | The residual of the Framework | The polar_log path encoding |
Quick Start#
- Build ACRES Core:
PYTHON=~/miniconda3/envs/torchenv/bin/python Core/build.sh. - Open a tmux session:
tmux new -s ppo. Usetmux attach -t ppoif the session exists. - Go to the repository:
cd ~/Codes/ACRES. - Start the run:
Learning/Scripts/ppo.sh train scout1. - Detach from tmux with Ctrl-b d. The run continues.
Learning/Scripts/ppo.sh activates the torchenv conda environment. It puts Learning and Core/Build on the Python
path. Then it runs python -m acres_learn.adapters.ppo.train --run scout1 in the folder Learning/.
Trainer Options#
| Option | Function |
|---|---|
--run <name> |
The run name. The run folder is Acres/Saved/Training/<name> |
--config <name or file> |
The configuration of a new run. The default is ppo_deployable.json |
--set section.key=value |
Changes one configuration value. The key must exist. You can repeat the option |
--init-from <checkpoint> |
A new run starts from the model, the optimiser and the normalisers of a checkpoint |
--device <cuda or cpu> |
The device. The default is CUDA if it is available |
--no-keys |
The trainer does not read keys from the terminal |
--max-env-steps <n> |
The trainer stops after this number of environment steps |
--max-minutes <n> |
The trainer stops after this training time |
Learning/Scripts/ppo.sh train scout1 --set env.num_envs=512 --set ppo.learning_rate=1e-4
Learning/Scripts/ppo.sh train teacher1 --config ppo_scouting_v1
Learning/Scripts/ppo.sh train fork1 --init-from /abs/path/ckpt_000000400.pt --set curriculum.start_stage=1
- The first start writes the configuration to the run folder. Later starts of the same run use that copy.
--setapplies on top of the saved copy. The trainer saves the result.- Use an absolute path for
--init-from. The trainer runs in the folderLearning/. - A fork starts at the stage
curriculum.start_stage. Its counters and its episode window start new.
Pause, Resume and Stop#
Use these commands from any terminal.
Learning/Scripts/ppo.sh ctl scout1 pause # saves a checkpoint, then waits
Learning/Scripts/ppo.sh ctl scout1 resume
Learning/Scripts/ppo.sh ctl scout1 stop # saves a checkpoint and exits
Learning/Scripts/ppo.sh ctl scout1 status # stage, steps, rates, recent episodes, last validation
Learning/Scripts/ppo.sh ctl --list # all runs and their states
In the terminal of the trainer, these keys do the same: p pause, r resume, q save and quit, s status. Ctrl-C saves a checkpoint and exits. A second Ctrl-C exits immediately without a checkpoint.
Pause. A paused trainer waits between two environment steps. The worker processes block on their pipes. The CPU load is 0.05 % of one core.
The trainer releases the cached GPU memory and holds approximately 0.8 GB. The RAM stays allocated (approximately 7.4 GB for 1024 environments). You can run the simulator during a pause.
Start again. After a stop, a crash or a reboot, use the same start command. The trainer resumes from the latest
checkpoint. The checkpoint holds the model, the optimiser, the normalisers, the stage, the window of recent episodes,
the random generators and the step counters. The environments start new episodes. Only one trainer can own a run folder
(lock).
Outputs#
Each run has a folder Acres/Saved/Training/<run>/. The folder is not in git.
| Path | Content |
|---|---|
config.json |
The configuration of the run |
checkpoints/ckpt_<update>.pt |
Checkpoints: one for each 25 updates, and one at each pause, stop and stage change. The last 5 stay |
checkpoints/best.pt |
The checkpoint with the best validation score: stage, success, coverage, return, in that order |
checkpoints/best_stage<k>.pt |
The best checkpoint of stage k |
tb/ |
TensorBoard events |
metrics.csv |
One row for each update |
status.json, control.json, lock |
The control files (adapters/ppo/control.py) |
train.log |
All lines that the trainer printed |
episodes/ |
Episode logs that watch wrote |
The trainer caches the transits of stage 0 in Acres/Saved/Learning/ppo/.
Watch a Run#
TensorBoard. Run Learning/Scripts/ppo.sh tensorboard scout1. Then open http://localhost:6006. Without a run
name, TensorBoard shows all runs.
| Group | Scalars |
|---|---|
episode/ |
Return, success, coverage, crop outside the edge band (m²), edge-band crop (m²), length |
episode_terms/ |
Each reward term, summed for each episode |
reward_step/ |
Each reward term for each step, averaged over the rollout |
endings/ |
The share of episodes for each ending: success, collision, stuck, deep in crop, left the map, lost, stalled, timeout |
window/ |
Success and return over the last 500 episodes |
losses/ |
Policy loss, value loss, entropy, approximate KL, clip fraction, explained variance, gradient norm, KL stop |
charts/ |
Learning rate, action standard deviation, reward scale, mean value, epochs, steps for each second |
time/ |
Seconds for the rollout, the environments and the update |
curriculum/ |
The stage, the steps in the stage, the value warm-up |
rehearsal/ |
Success, collisions and return of the rehearsed episodes of each earlier stage |
eval/ |
The validation missions: success, return, coverage, crop, time |
Figure. Run this command in Learning/ with the Python path of ppo.sh: python -m acres_learn.adapters.ppo.plot
scout1 --out curves.png. The figure shows the return, the success rate, the steps for each second and the explained
variance.
Status. ctl scout1 status prints the stage, the rates, the estimated times, the recent episodes, the last
validation and the last checkpoint.
Algorithm#
The values are those of the two configurations that exist.
| Setting | Value | Reason |
|---|---|---|
| Environments | 1024 in 8 worker processes, 16 Core threads | The task bookkeeping in Python runs on 8 cores |
| Rollout | 64 steps (6.4 s), 65,536 samples for each update | |
| Discount, GAE | \(\gamma = 0.995\), \(\lambda = 0.95\) | A horizon of 200 steps (20 s) |
| Update | 4 epochs of 8 minibatches of 8192 samples, Adam | |
| Learning rate | \(3 \times 10^{-4}\) to \(3 \times 10^{-5}\), linear over \(2 \times 10^9\) steps | |
| Losses | Clipped surrogate (0.2); clipped value loss (0.2) times 0.5; entropy bonus 0.001; gradient norm clipped at 0.5 | |
| Normalisation | Advantages for each minibatch. Running mean and variance of the vector inputs, clipped at ±10. Rewards divided by the running deviation of the discounted return, clipped at ±10 | The rewards are from −100 to +1 for each step |
| Policy | A Gaussian with a tanh squash and a state-independent log standard deviation | |
| Initial deviation | 0.6 on the curvature, 0.25 on the speed | See below |
| Initial mean speed | Action 0.46, which is 2.8 m/s | See below |
ppo.kl_stop |
1.0 in ppo_deployable.json |
An update stops at the first minibatch with an approximate KL above this value |
| Validation | Each 100 updates |
Network (network.py).
- A CNN for the map crop: 4 or 5 channels of 64 × 64 cells, four stride-2 convolutions (32, 64, 64, 64 channels), then 256 features.
- 1-D convolutions with circular padding for the 360 LiDAR beams (16, 32, 32 channels), then 128 features.
- An MLP for the proprioception and the path: 130 inputs, two layers of 256.
- Two trunks (640 to 512 to 256): one for the mean of the actor and one for the value of the critic.
- The convolutions run in bfloat16 on the GPU. The other layers run in float32.
The environments do not send the map crop. Each environment sends the pose, and the GPU cuts the crop from the
ground-class raster (CropRaster). The result is equal to observations.map_crop (tested). The rollout keeps the crop
as class codes: 4 KiB for each step and not 80 KiB.
Squashing. The buffer keeps the Gaussian sample \(u\). The environment gets \(a = \tanh u\). The log-probability is exact: \(\log \pi(a) = \log \mathcal{N}(u) - \sum \log(1 - \tanh^2 u)\). The entropy bonus is the entropy of the Gaussian.
Time limits. A time limit truncates an episode. It does not terminate the episode. The trainer bootstraps its last
step from the value of its final observation. A success or a failure terminates the episode without a bootstrap
(algorithm.gae, tested on a hand-worked case).
Lost episodes end. The tracking term \(-0.5 \max(0, \lvert e_y \rvert - h)^2\) has no lower bound. With env.lost_m
(10 m), an episode ends as the failure "lost" when the vehicle is more than 10 m outside its corridor. This is a
training aid. The evaluation does not use it.
The speed starts forwards. A speed action below −0.13 is a request for reverse. The gear logic then stops the vehicle, shifts and holds for 1.1 s. The first initial values were a mean of 0 and a deviation of 0.6. Then one quarter of the steps asked for reverse, and the vehicles almost did not move. Thus the mean starts at 2.8 m/s and the deviation at 0.25.
NaN actions. The environment reads a NaN action as a stop. It never stores a NaN as the previous action.
Rewards. The rewards are those of the task (tasks/scouting/rewards.py). The log has each term separately.
Inputs on the Vehicle#
ppo_deployable.json trains a policy that sees only quantities that the vehicle measures.
| Input | Source on the Polaris |
|---|---|
| Speed | vehicle/vehicle_velocity (ds_dbw VehicleVelocity) or the OxTS velocity (oxts/velocity) |
| Yaw rate, lateral acceleration, pitch, roll | The OxTS RT3000 (oxts/imu) |
| Steering wheel angle | vehicle/steering/report (SteeringReport.steering_wheel_angle) |
| Four wheel speeds | vehicle/wheel_speeds (rad/s) times the tyre radius 0.343 m |
| Gear | vehicle/gear/report |
| Previous action | The last command of the program |
| Path | The plan and the reference path on the map products, and the OxTS pose in the map frame |
| Map crop: crop, edge band, lane or verge, obstacle | The map products (ground_classes.u8) around the pose |
| LiDAR, 360 beams | The Helios points within 1° of the horizontal, the nearest point for each degree |
The soil-water channel of the first specification is privileged. The vehicle does not measure soil water.
observation.soil_water: false removes it. The crop then has four channels.
Randomisation#
The sensor noise (tasks/scouting/sensor_noise.py) changes only the inputs of the policy. It does not change the
physics, the rewards or the task bookkeeping (tested). Each episode draws its errors.
| Error | Default | Reason |
|---|---|---|
| GNSS/INS solution: RTK fixed, 80 % | Position 2 cm noise and 3 cm bias; heading 0.05° and 0.1° | The OxTS holds centimetres with RTK |
| GNSS/INS solution: float, 15 % | Position 5 cm and 0.5 m; heading 0.1° and 0.3° | Decimetres without a fixed solution |
| GNSS/INS solution: none, 5 % | Position 10 cm and 1.5 m; heading 0.2° and 1° | Metres without corrections. The bias drifts with a time constant of 60 s |
| Map offset | Up to 0.6 m and 0.5°, about the start of the episode | The surveyed map and the vehicle frame do not align exactly |
| Latency | 0 to 60 ms | Delays of the sensors and the software |
| Proprioception noise | 0.02 m/s speed and wheel speeds, 0.005 rad/s yaw rate, 0.05 m/s² lateral acceleration, 0.5° steering wheel, 0.3° pitch and roll | Resolution and noise of the reports |
| LiDAR | 3 cm range noise, 2 % of the beams dropped | Range noise and missing returns of the Helios |
| Drive-by-wire | Steering centre ±3°; wheel speed ±10 %; each fitted parameter of polaris.json within its fitted deviation, clipped at 2 sd |
The vehicle on a different day |
| Soil water (last stage only) | Uniform from 0.5 to 1.3 | Dry to almost saturated soil |
| Start heading (last stage only) | Within 5° of the heading of H |
A stage sets its share of the sensor errors with noise_scale. It sets its share of the drive-by-wire errors with
dbw_scale.
Flat-Policy Curriculum#
The flat policy trains in stages. A stage moves on when its criterion is true. A stage can also end on a step budget
(max_env_steps).
ppo_scouting_v1.json:
| Stage | Missions | Condition to Move On |
|---|---|---|
0 goto |
2000 transits of 50 to 300 m on lanes, dry. One fifth starts at H | 95 % of the last 2000 episodes reach the target, after 5 M steps |
1 scout_lane_bound |
Single-field missions of the lane-bound training fields | 95 % complete and mean crop outside the edge band below 1 m², over 200 episodes, after 20 M steps |
2 scout_all |
Single-field missions of all 43 training fields | The same over 400 episodes, after 50 M steps |
3 short_missions |
2 to 3 training fields with the optimal plan | 90 % complete over 400 episodes, after 50 M steps |
4 missions |
1 to 8 training fields; dry, wet and hardware shift; randomised | Last stage |
The lane-bound training fields are F11, F13, F15, F17, F18, F33, F58 and F59. At least 90 % of their loops is on lane or verge.
ppo_deployable.json:
| Stage | Missions | Condition to Move On | Budget | Noise, Drive-by-Wire Scale |
|---|---|---|---|---|
0 goto |
Transits as above; 40 % start at H | 95 % over 2000 episodes, after 5 M steps | None | 0.5, 0 |
1 scout_loops |
The loops of the lane-bound fields alone, from the straight next to the entry | 70 % over 400 episodes, after 10 M steps | 150 M | 1.0, 0.5 |
2 scout_lane_bound |
Single-field missions of the lane-bound fields | 70 % and crop below 10 m², after 20 M steps | 300 M | 1.0, 0.5 |
3 scout_all |
Single-field missions of all training fields | 70 % and crop below 15 m², after 50 M steps | 400 M | 1.0, 0.5 |
4 short_missions |
2 to 3 training fields | 90 % over 400 episodes, after 50 M steps | None | 1.0, 1.0 |
5 missions |
1 to 8 training fields; dry, wet and hardware shift | Last stage | None | 1.0, 1.0 |
Stages 1 to 3 rehearse earlier stages. Stage 1 draws 20 % of its episodes as transits from H. Stage 2 draws 15 % loops and 10 % transits. Stage 3 draws 10 % loops and 5 % transits.
Plans. A single-field mission uses the frozen plan of configs/scouting_reference_v1.json. A mission with more
fields uses the exact dynamic programme on the dry cost tables. The training draws only fields of the training split
(configs/splits_v1.json).
Validation. Each 100 updates, the policy drives the validation missions of the stage with its mean action.
| Stage Type | Validation Missions |
|---|---|
| Transits | 64 transits from a different seed |
| Loops | The loop of each of the 6 validation fields |
| Single field | Each of the 6 validation fields alone |
| 2 to 3 fields | The validation fields in pairs and triples |
| Last stage | Singles and pairs under each condition set |
Curriculum Options#
These options came from the diagnoses of the flat runs (Results).
| Option | Function | ppo_deployable.json |
|---|---|---|
env.launch_hold |
Holds the speed command from a decrease while the vehicle stands (spec.LaunchHold). The exported policy carries the option |
On |
env.path_through |
The path input continues into the path of the next phase | On |
env.path_lookahead |
With path_through: the path input of a lone transit continues along the leg that follows |
Off |
env.stall_s, env.stall_penalty |
Ends an episode as "stalled" after the vehicle stands for this time | Off |
env.lost_m |
Ends an episode as "lost" at this distance outside the corridor | 10 m |
env.rewards |
"v1" (first version) or "bounded" (failures never pay) |
"v1" |
missions: loop, loop_start: straight |
A stage of loops alone. The episode starts on the straight next to the entry corner | Stage 1 |
Stage rehearsal |
A share of the episodes comes from an earlier stage. The move-on test does not count them | Stages 1 to 3 |
Stage entry_share |
A share of the transits is the GoTo from H to a field entry | 0 |
Stage home_share |
A share of the transits starts at H | 0.4 |
curriculum.value_warmup_env_steps |
The first steps of a stage train only the critic | 1 M |
ppo.kl_stop |
Stops an update at a large KL | 1.0 |
The configuration check rejects env.rewards = "bounded" when a failure costs less than the floor divided by \(1 -
\gamma\). It rejects env.path_lookahead without env.path_through.
Watch an Episode in Unreal#
Learning/Scripts/ppo.sh watch scout1 # latest checkpoint, first validation mission
Learning/Scripts/ppo.sh watch scout1 --checkpoint best --field F42
Learning/Scripts/ppo.sh watch scout1 --stage 0 --index 3 # the fourth validation transit
watchdrives the episode in ACRES Core with the mean action.- It writes an MCAP episode log to
Acres/Saved/Training/<run>/episodes/.--outselects a different folder. - It prints the replay command for the game:
Packaged/Linux/Acres.sh -VehicleDemo -Vehicle=polaris -EpisodeReplay=<file.mcap>. - With
--render, the packaged game renders an MP4 next to the log. Run only one game at a time.
Pause the trainer before you start the game. The game then has the CPU and the GPU.
Evaluation#
A checkpoint drives the evaluation suite in place of the reference driver. The plans and the scores are the same
(eval.run with --driver). Thus the tables agree with Scouting Baselines.
Learning/Scripts/ppo.sh eval scout1 --checkpoint best --suite --out ~/ppo-eval/suite # the 600 runs
Learning/Scripts/ppo.sh eval scout1 --single-field --split validation --out ~/ppo-eval/val # six fields alone
\(J^*\) and \(T_\text{ref}\) are the frozen values of the reference driver (configs/scouting_reference_v1.json). The
output folder gets report.md, report.json and results.json.
GoTo from H. This command drives the GoTo from H to the entry of each field. Run it in Learning/ with the Python
path of ppo.sh.
python -m acres_learn.eval.goto --driver reference --driver ppo:goto2@best --runs 4 --noise both --out <dir>
| Option | Function |
|---|---|
--driver |
reference, or ppo:<run or checkpoint>[@best, @latest or @best_stage<k>]. You can repeat the option |
--fields |
all, train, val, test or a list |
--runs |
Runs for each field |
--noise |
off, on or both. With noise, the policy samples its actions |
--lookahead |
policy, on or off: the path input past the entry |
Export#
The vehicle PC has no torch. The export writes the actor and its normaliser to one .npz file.
Learning/Scripts/ppo.sh export scout1 --checkpoint best --out Acres/Saved/Training/scout1/policy.npz
Piece (adapters/ppo/deploy.py) |
Function |
|---|---|
NumpyPolicy |
The actor in numpy. It reproduces the float32 actions of the torch network to \(10^{-5}\) (tested) |
VehicleObservation |
The observation from the vehicle quantities, with the same code as training |
planar_scan |
The Helios cloud to 360 beams of one degree |
CommandMapper |
An action to the ds_dbw SteeringCmd, UlcCmd and GearCmd |
| make_node | A minimal rclpy node at 10 Hz. It is a skeleton. The deployment program is the full front end |
The file carries the options of the training: launch_hold, path_through, path_lookahead and soil_water.
Deployment uses the file.
This example shows the call sequence for one mission:
from acres_learn.tasks.scouting.map_products import ScoutingMap
from acres_learn.tasks.scouting import routes
from acres_learn.adapters.ppo.deploy import NumpyPolicy, VehicleObservation, CommandMapper, planar_scan
m = ScoutingMap.load()
policy = NumpyPolicy.load("scout1.npz", m=m) # the actor and its normaliser
legs = routes.mission_legs(routes.Network(m), plan) # the plan, e.g. [("F39", 0, True)]
obs = VehicleObservation(m, legs, ["F39"], t_ref_s, path_through=policy.path_through)
mapper = CommandMapper(policy.launch_hold) # the gear logic of the vehicle
# each 0.1 s:
o = obs.observe(pose=(east_m, north_m, yaw_rad), speed_mps=..., yaw_rate_radps=..., lateral_accel_mps2=...,
steering_wheel_rad=..., pitch_rad=..., roll_rad=..., wheel_speeds_mps=[fl, fr, rl, rr], gear=...,
lidar_ranges=planar_scan(points_xyz)) # a batch of one
action = policy(o)[0] # in [-1, 1]^2, the mean action
obs.acted(action)
cmds = mapper.commands(action, speed_mps) # {"steering_cmd": {...}, "ulc_cmd": {...}, "gear_cmd"?: {...}}
Training runs the convolutions in bfloat16. This moves the actions by a median of \(5 \times 10^{-4}\). Use --set
network.amp_bf16=false to train in float32. That is approximately 30 % slower.
Throughput#
Measured on the development machine (16 threads, RTX 5060 Ti 16 GB) in stage 0.
| Environments | Steps for Each Second | Environments / Update Time |
|---|---|---|
| 256 | 11,500 | 0.6 s / 0.7 s |
| 512 | 12,000 | 1.1 s / 1.4 s |
| 1024 (default) | 12,300 | 2.0 s / 2.7 s |
| 2048 (16 workers) | 12,800 | 3.6 s / 5.5 s |
The environments alone step 35,500 times a second with 1024 environments in 8 workers. The update is approximately half of the time.
Residual Training#
State: specified, not built. The Framework gives the rules. The planned changes of the trainer are:
| Change | Rule in the Framework |
|---|---|
| One stage of segment episodes; no curriculum | Segment Episodes |
| The policy output is a residual on the base command | Residual Law |
The path input uses the polar_log encoding |
Path Encoding |
| The critic gets two more inputs than the actor | Critic Group |
| The launch hold runs before the shield. The shield filters each command in the environment, before the gear logic | Order in a Control Step, Shield |
The reward variant residual. It has a shield term of 0.5 after the bounded map, a shaping term with \(c_\Phi = 0.55\), failures at −1200 and a collision at −1500 |
Reward |
| A truncation bootstraps with the critic of the shaped problem. The potential at a truncation is not 0 | Shaping Term |
rewards.gamma is equal to ppo.gamma |
Reward |
| The code of the residual law, the shield, the reward and the shaping term passes the golden vectors | Golden-Vector Contract |
| A critic-only warm-up at the start of the run | Reasons |
| A whole-mission evaluator for each gain \(\alpha\) | Evaluation Protocol |
The trainer loop, the control tools, the encoders and the export stay as they are.
Limitations#
- A resume starts new episodes. The state of ACRES Core in the middle of an episode is not in the checkpoint.
- The update is half of the time. The GPU update and the rollout alternate.
- Workers and parked vehicles are not in ACRES Core. The training does not randomise them.
- The soil water is one value for each episode.
- The planar LiDAR on the vehicle is the Helios rings within 1° of the horizontal. ACRES Core casts one level ray for each degree. A comparison on a recorded drive is open.