Skip to content

Learning Stack#

The learning stack trains and runs learned drivers for the ACRES vehicles. The Python package is acres_learn in Learning/. The first task is the Field Scouting Task for the Polaris Ranger.

The stack has one rule. A classical planner and a scripted mission automaton carry the long horizon. A learned policy corrects the classical driver for one leg at a time. A shield with Lean 4 proofs has the last word on each command. The page Framework gives the reasons for this rule.

The ACRES learning stack

Blue: classical. Orange: learned. Green: classical with Lean proofs. Grey: simulator and tools.

The diagram shows the target of the Framework. The planner, the reference driver, ACRES Core, the PPO trainer, the deployment program and the Lean suite exist. The residual policy, the shield, the automaton module and the whole-mission evaluator are specified, not built. Their Lean models and proofs exist. Results gives the state.

Pages#

Page Content
Framework The layers, the rates, the design choices and the rejected alternatives
Training How to train, control, watch, evaluate and export a policy; all options
Deployment The modes SIM, SHADOW and DRIVE, the drivers, the field-day kit
Verification The Lean 4 suite: the theorems, the findings, the golden vectors, the trusted base
Results The flat-policy negative baseline and the results of the framework
Tasks The task specifications, the map products and the reference baselines

Goals#

  1. Train fast. Many simulations run in parallel, faster than real time, without rendering.
  2. Use ROS 2. The simulators publish the topics of the real vehicles. Software moves between them unchanged.
  3. Keep the classical floor. With the gain \(\alpha = 0\) the system is the reference driver, which completes 59 of 59 missions.
  4. Prove the parts that permit a proof. The drive-by-wire command path, the gear logic, the action limits and the reward properties have Lean models and proofs. The residual law, the shield, the automaton, the shaping term and the planner also have them.
  5. State the parts without a proof. No proof covers the neural network, the perception inputs or the localisation. No proof covers the difference between the model and the real vehicle.
  6. Show each episode. Each episode has a log. The game can replay the log from all cameras.

Package Layout#

Path Content
Learning/acres_learn/tasks/ Task definitions. scouting/ holds the map products, routes, mission bookkeeping, reference driver, observations, rewards and terminations
Learning/acres_learn/planners/ Cost tables from ACRES Core, the option graph and the solvers
Learning/acres_learn/envs/ The Gymnasium vector environment on ACRES Core; the drive-by-wire client of the game
Learning/acres_learn/adapters/ppo/ The PPO trainer, its control tools, the network, the numpy policy and the export
Learning/acres_learn/eval/ The evaluation suite, the scorers and the reference measurements
Learning/acres_learn/deploy/ The deployment program for the simulated and the real vehicle
Learning/acres_learn/configs/ The task configuration, the splits, the frozen reference values and the trainer configurations
Learning/Scripts/ ppo.sh (training tools) and deploy.sh (deployment tools)
Learning/tests/ The tests. run_all.py runs all of them
Core/ ACRES Core: the simulator without Unreal, with the acres_core Python module
Verification/ The Lean 4 models, proofs and differential tests
ROS/acres_interfaces, ROS/acres_sim, ROS/acres_core_sim Messages, the bridge to the game, the same interfaces on ACRES Core

An adapter is the only code that a new algorithm needs. The task, the environments, the data and the evaluation stay the same.

Simulator Tiers#

Tier Simulator Sensors Speed Use
0 Core ACRES Core, batched State, ray-cast LiDAR 650 times real time for each thread Reinforcement learning, cost tables, evaluation
1 Headless Unreal The game with -nullrhi Collision-trace LiDAR, no camera Real time or faster Checks near the full world
2 Training render The game with a low-cost render profile Camera at 224 or 384 pixels, LiDAR Not measured Closed-loop evaluation of vision models
3 Calibrated The game with the calibrated sensors Calibrated camera and LiDAR Real time Datasets and final checks

ACRES Core compiles the same model sources as the game. Learners that need images do not render during training. The game renders selected episodes later from their logs.

Interfaces#

ROS 2 Humble is the interface of record. ACRES Core and the game publish the same topics, services and actions. The one exception is batched training. The trainer calls the batched Python API of ACRES Core in the same process.

Vehicle Topics
Polaris vehicle/{steering,throttle,brake,gear,ulc}/cmd and /report (ds_dbw_msgs 2.3.11), vehicle/enable, vehicle/disable, vehicle/odom, oxts/{imu,velocity,fix}, lidar/points, camera/image_raw, camera/camera_info, /tf, /tf_static
Maxxum can/tx, can/rx (can_msgs/Frame with SAE J1939 and ISO 11783 frames), odom, imu/data, gnss/fix, lidar/points, camera/*, implement/state
Simulator Interface Kind Purpose
step_simulation, simulate_steps Service, action Advance the requested physics steps (lockstep)
set_simulation_state Service Play, pause, stop
reset_simulation Service Reset the world or only the agents
spawn_entity, delete_entity, get_entity_state, set_entity_state Services Agents and props
get_named_poses Service The places of places.json
acres/set_conditions Service Soil water, weather, time of day
acres/set_vehicle_shift Service Hardware shifts with labels
acres/record Service Start and stop an MCAP episode log
acres/farm_state, acres/field_query Services Ground truth for tasks and scores
/sim/ground_truth/odom, /sim/episode Topics Exact pose; episode identifier and time
/clock Topic The physics clock

With lockstep on, the simulator advances only when a client asks. A slow model then does not make a slow vehicle. The same seed and the same commands give the same run.

Data#

  • Episode log. Each run writes an MCAP file. The file holds the poses, the commands, the farm events, the conditions and the reward terms.
  • Replay. The game loads an MCAP file and puts the agents at the logged poses. It renders from all cameras.
  • Datasets. A converter from MCAP to the LeRobot format (v3.0, and v2.1 for FastWAM) is planned for the vision models.

Roadmap#

Phase Content State
1 Groundwork Map products of the scouting task, the Learning/ package Done: each field has a closed loop with curvature ≤ 0.15 m⁻¹
2 ROS 2 and lockstep acres_interfaces, acres_sim, lockstep, MCAP record and replay Done: the field-day kit test passes 10 of 10 on acres_sim
3 ACRES Core The library, the Polaris model, ray-cast LiDAR, the batched Python API Done: Core and game agree within two times the run-to-run spread of the game
4 Task, planner, reference The scouting task, cost tables, solvers, reference driver, evaluation suite Done: 59 of 59 single-field missions; regret 0 against brute force
5 Reinforcement learning Flat PPO on whole missions Stopped: 0 of 64 missions. It is the negative baseline (Results)
5b Framework pivot Residual on pure pursuit, shield, shaped reward, automaton, whole-mission evaluator The specification and the Lean proofs exist. The code is not built (Results gives the state)
6 Cost model Energy and slip cost of each segment as a function of soil moisture, for the exact planner Next
7 Vision and language models ACT, Diffusion Policy, SmolVLA, pi0.5, FastWAM as segment drivers behind the same shield Later
8 Harness Skills as ROS 2 actions, a plan checker, a VLM as task sequencer above the exact planner Later
9 Real vehicle SHADOW drives, then DRIVE with the gain and the shield state shown to the safety driver The deployment program exists. Only the laboratory does drives

Compute#

  • Local machine (RTX 5060 Ti, 16 GB): ACRES Core, the trainer, the ROS 2 stack, the game, small vision models.
  • Large GPU (L40S class, later): fine-tuning of pi0.5 and FastWAM.

Risks#

  • Core and game can disagree. The checks of phase 3 put them within two times the run-to-run spread of the game.
  • The vehicle has no wet-soil measurement. The slip on soft ground comes from the soil model, not from ACRE data.
  • The planar LiDAR in ACRES Core has no crop. The real LiDAR sees tall crop as a wall beside the lane.
  • No measurement of the render-later speed exists.