Results#
This page gives the measured results of the learning stack. It has two parts. The first part is the flat-policy baseline, which is a negative result. The second part is the state of the Framework.
All runs are in ACRES Core on the map products of 1 October 2026 (hash dc4af5f51f316987), unless a row gives a
different basis.
Summary#
| Subject | Result | Date |
|---|---|---|
| Reference driver, single-field missions | 59 of 59 complete | 2 October 2026 (run again for this page) |
| Reference driver, GoTo from H | 236 of 236, with and without noise | 1 October 2026 |
| Exact planner against brute force | Regret 0 on 60 subsets (largest difference \(3.7 \times 10^{-16}\)) | 2 October 2026 (run again for this page) |
| Flat PPO, GoTo from H | 69 % without noise, 63 % with noise (best checkpoint) | 1 October 2026 |
| Flat PPO, whole single-field missions | 0 of 64 | 1 October 2026 |
| Residual framework | No result. The code is specified, not built. The Lean models and proofs exist | |
| Lean suite | 554 theorems pass the axiom audit (56 + 241 + 257); code comparisons pass within their stated tolerances; the golden vectors agree | 2 October 2026 |
Flat-Policy Baseline#
A flat policy is one PPO network that outputs the full command for the whole mission. Six runs trained it with the
configuration ppo_deployable.json and 1024 environments. No run completed a whole mission.
Runs#
| Run | Start | Steps | Result |
|---|---|---|---|
scout1 |
New | 257 M | Stage 0 passed in 14 minutes (10 M steps). Whole missions failed: the vehicles stopped and stayed stopped |
scout2 |
Best stage-0 checkpoint of scout1, with the loops stage |
167 M | The loops stage ended on its budget at 0.65 success. Whole missions: 98 to 100 % collisions at the garage exit |
scout3 |
Best stage-0 checkpoint of scout1, loops from the straight |
70 M | The loops stage passed at 0.71 after 13 M steps. Whole missions: success 0 |
scout4 |
Best loops checkpoint of scout3, with rehearsal, value warm-up and KL stop |
66 M | Whole missions: success 0 for 62 M steps |
goto1 |
Best checkpoint of scout4; GoTo legs only |
56 M | GoTo from H: 44 % |
goto2 |
Best checkpoint of scout4; GoTo legs with the path look-ahead |
318 M | GoTo from H: 69 % at 141 M steps. A plateau near 0.6 follows |
GoTo from H#
The test is the GoTo leg from H to the entry of each of the 59 fields, four runs for each field (236 runs). The legs are 307 to 2545 m long (median 1430 m). With noise, the policy samples its actions.
| Driver | Without Noise | With Noise | Fields with 8 of 8 | Main Failures without Noise |
|---|---|---|---|---|
| Reference driver | 236 of 236 | 236 of 236 | 59 | None |
scout1, best stage 0 |
0 of 236 | 19 of 236 (8 %) | 0 | 184 collisions, 44 lost |
scout4, best |
148 of 236 (63 %) | 131 of 236 (56 %) | 4 | 60 collisions; 45 failures at the entrance sign |
goto1, best |
103 of 236 (44 %) | 96 of 236 (41 %) | 4 | 89 timeouts a few metres before the entry |
goto2, update 400 (26 M steps) |
154 of 236 (65 %) | 147 of 236 (62 %) | 17 | 31 timeouts, 27 lost |
goto2, update 2150 (141 M steps) |
162 of 236 (69 %) | 148 of 236 (63 %) | 22 | 32 timeouts, 17 deep in crop |
The stage-0 policy of scout1 trained on transits of 50 to 300 m. On the long legs, its distance-to-go input was more
than 10 standard deviations outside the training range.
Whole Missions#
The test is the eight lane-bound single-field missions, eight runs each (64 runs), with the training noise.
| Policy | Complete | Failures by Phase |
|---|---|---|
scout4, best |
0 of 64 | GoTo 21, Scout 36, Return 7 |
A GoTo at 63 to 69 % and a loop at approximately 70 % give approximately one third of the missions, if the phases are independent. The result of 0 of 64 shows that the policy also fails at the joins between the phases.
Diagnosis 1: The Vehicles Stopped#
scout1 passed stage 0 and then failed on whole missions. After 85 M steps no mission was complete. The coverage
decreased from 0.8 to 0.2. 60 % of the episodes ran out of time after approximately 11,000 steps.
| Cause | Evidence |
|---|---|
| The vehicles stopped and stayed stopped | The checkpoint of update 2075 stood in 6 of 8 missions until the time limit, 8 to 12 m before the end of a GoTo |
| A stop is a trap in the ULC | The ULC brakes at standstill when its speed reference decreases. A noisy reverse command of 0.45 m/s never moved the vehicle in 60 s. A constant command moved it in 4.7 s |
| A stop looked cheap | A stop costs approximately \(-0.25 / (1 - \gamma) = -50\) in discounted value. A failure costs −50 to −100. A timeout is a truncation |
| The end of a leg looked like a stop | The path input repeated the last point of the leg. All transits of stage 0 ended there with a success |
The fixes are options of the trainer (Curriculum Options): env.launch_hold,
env.path_through, the loops stage and env.stall_s.
Side runs from the best stage-0 checkpoint of scout1 (256 environments, approximately 30 minutes each):
| Run | Steps | Success | Coverage | Endings |
|---|---|---|---|---|
| Whole missions, no fix | 3.7 M | 0 | 0.45 | Lost 56 %, deep in crop 25 %, collision 15 % |
| Whole missions, launch hold and path through | 4.5 M | 0 | 0.54 | Lost 49 %, collision 28 % |
| Whole missions, launch hold and stall end | 3.3 M | 0 | 0.38 | Stalled 46 % |
| Whole missions, all three | 1.5 M | 0 | 0.00 | Stalled 99 %, at H |
| Loops alone, no other fix | 3.0 M | 0.51 | 0.83 | Lost 16 %, left the map 11 % |
| Loops alone, all three | 5.4 M | 0.54 | 0.86 | Lost 21 %, stalled 11 % |
The loops stage made the loops learnable. The stall end did not help: a stop that costs as much as a crash became the choice of the policy immediately.
Diagnosis 2: Hard Places on the Map#
The best loops checkpoint of scout2 drove each lane-bound loop eight times with the training noise. The table gives
the number of scouted loops out of 8, before and after the map products were rebuilt with larger margins (Scouting Map
Products).
| Field | Before | After, from the Entry | After, from the Straight | Failure and Cause |
|---|---|---|---|---|
| F11 | 8 | 8 | 7 | The 4.5 m lane shared with F13: the loop was on one side and is now in the middle |
| F13 | 8 | 8 | 6 | A corner entry passed 3 to 6 m inside |
| F15 | 6 | 0 | 8 | The closure: each run passed 3.0 to 4.6 m inside the corner entry |
| F17 | 7 | 4 | 4 | Deep in crop and lost at a corner at the west end |
| F18 | 0 | 0 | 0 | Left the map: the policy drives straight on to the edge of the tile |
| F33 | 0 | 0 | 2 | Lost: the policy drives on along the lane past the hairpin through the crop strip |
| F58 | 7 | 6 | 6 | The shed adjacent to the lane, now 4.0 m away |
| F59 | 6 | 6 | 1 | From the straight the vehicle does not move |
The new products removed the geometric causes. F18 and F33 stayed unsolved. They are behaviours that the policy must learn, not geometry.
Diagnosis 3: The Collapse on Whole Missions#
scout2 entered the whole-mission stage and never completed a mission. Its last checkpoints drive straight out of the
garage at 6 m/s into the tree row 22 m north of H.
| Policy | Collisions at the Garage Exit | At the Entrance Sign | Other |
|---|---|---|---|
scout1, best stage 0 (update 100) |
32 | 14 | 5 lost near H, 13 still driving |
scout2, best loops stage (update 800) |
2 | 61 | 1 still driving |
scout2, whole missions (update 2544) |
64 | 0 |
Three causes came together.
| Cause | Evidence |
|---|---|
| The garage exit was partly forgotten | Only stage 0 drove it. The loops stage never starts at H. Its LiDAR view is 9 standard deviations outside the loops statistics |
| The first update of the new stage damaged the policy | Approximate KL 8.9, clip fraction 0.83, gradient norm 816 against approximately 5. The explained variance of the critic was −0.2 |
| The entrance sign stops each mission | The sign is in the 12 m lane, 227 m from H. The policies cut the corner at 5.5 to 6.5 m/s where the path limit is 2.1 m/s. 61 of 64 runs hit the sign |
The deeper cause is the reward balance. A failure ends an episode for a fixed −100. A mission that continues and fails later costs much more: −3047 for deep in crop, −7024 for lost, −14,613 for a timeout, on average. Thus an early collision (−140) was the best return available to the policy.
Side runs (forks into the whole-mission stage, 128 environments). "Fixes" are rehearsal, value warm-up and KL stop.
| Fork Of | Fixes | Steps | Crashes at the Garage Exit (Training / Evaluation of 64) | At the Sign | Past the Sign |
|---|---|---|---|---|---|
scout2 update 800 |
None | 3 M | 5.3 % / 5 | 48 | 8 |
scout2 update 800 |
All | 4 M | 0.6 % / 0 | 63 | 1 |
scout2 update 800 |
None | 10 M | 0.2 % / 0 | 35 | 12 |
scout2 update 800 |
All | 10 M | 6.2 % / 6 | 17 | 34 |
scout3 best loops stage |
None | 3 M | 18.1 % / 34 | 0 | 30 |
scout3 best loops stage |
All | 4 M | 2.5 % / 0 | 3 | 38 |
The fixes keep the policy intact in the first millions of steps. They do not remove the attraction of an early crash.
Diagnosis 4: The Bounded Reward#
The bounded reward makes each failure cost at least the worst continuation (Failures Never
Pay). Two forks of the best loops checkpoint of scout3 compared the two
reward versions (128 environments, the three fixes on).
| Reward | Steps | Early Crashes | Evaluation: Garage / Sign / Other / Still Driving | Coverage | Action Deviation |
|---|---|---|---|---|---|
| First version | 6 M | 0.7 % | 0 / 12 / 16 / 36 | 0.30 | 0.20 |
| Bounded | 5 M | 1.4 % | 0 / 7 / 13 / 44 | 0.39 | 0.15 |
| First version | 10 M | 0.4 % | 0 / 32 / 19 / 13 | 0.28 | 0.17 |
| Bounded | 10 M | 0.9 % | 0 / 53 / 2 / 9 | 0.37 | 0.11 |
No run completed a mission. The bounded run ended fewer episodes deep in crop (12 % against 19 %) and covered more (0.37 against 0.28). It hit the sign more frequently (53 of 64 against 32) and ran out of time more frequently. The bounded reward gave no clear gain. The default stayed the first version.
Diagnosis 5: The GoTo Runs#
goto1 trained only on GoTo legs. Its best checkpoint reached the entry in 44 % of the GoTo legs from H. Of its
failures, 89 were timeouts a few metres before the entry. A lone GoTo has no next leg. Thus the path input had no
continuation and the points collected at the end.
goto2 used env.path_lookahead: the path input continues along the Scout leg of the field. The failures at the
entrance sign decreased from 45 (scout4) to 6. The success reached 65 % at 26 M steps and 69 % at 141 M steps. Later
checkpoints were not better.
Smoke Test of the Trainer#
On 30 September 2026 a run with ppo_deployable.json trained for 14 minutes (9.8 M steps) in three sessions.
- The training success of stage 0 increased from 0 to 0.86 in 2.0 M steps.
- The criterion of stage 0 was true at 5.44 M steps (7.5 minutes): success 0.952.
- The 64 validation transits gave 0.83 at 1.3 M steps and 0.94 at 5.2 M steps.
- Each restart resumed at the update where the run stopped (69 and 109).
metrics.csvhas updates 1 to 149 without a gap. - The rate stayed at 12,300 to 13,400 steps for each second.
Conclusion#
The flat policy learns a leg and does not learn a mission. The literature reports the same result for flat PPO on sparse-waypoint driving (Framework). The runs stay in the repository as the negative baseline. The configurations stay runnable.
Framework#
State#
| Item | State |
|---|---|
| Specification (Framework) | Written on 2 October 2026. Corrected to the Lean models on the same date |
| Classical floor: 59 of 59 single-field missions | Measured again on 2 October 2026 |
| Residual law, shield, segment episodes, shaped reward, automaton module | Specified, not built |
Whole-mission evaluator, deployment driver residual |
Specified, not built |
| Lean models and proofs of the residual law, the shield, the automaton, the shaping term, the residual reward and the planner | Exist: 498 theorems (Verification) |
| Golden vectors for the code of the framework | Exist: 12 files (Golden-Vector Contract) |
Constants of the residual reward in configs/scouting_v1.json (rewards_residual) |
Exist |
| Corrections of the code that the proofs found (Known Problems) | Not started |
| Training run of the residual | Not started |
| Evaluation of the gains \(\alpha\) = 0, 0.25, 0.5, 1.0 | Not started |
| Headroom study of the field order (energy-optimal against distance-optimal tours) | Not started |
Results#
There are no results of the framework. This section gets the tables when the runs are complete:
- The check of \(\alpha = 0\) against the reference driver.
- The whole-mission table for each gain, noise setting and soil condition.
- The shield interventions for each rule.
- The headroom study.
The project set the acceptance rules before the runs.
Checks Run for This Page#
These commands ran on 2 October 2026. Each command runs in Learning/ with the Python path of
Learning/Scripts/ppo.sh.
| Command | Result |
|---|---|
python -m acres_learn.eval.run --single-field --out <dir> |
59 of 59 complete, 81 s |
python -m acres_learn.eval.run --regret-check 60 --out <dir> |
60 subsets, largest regret \(3.69 \times 10^{-16}\) |
python -m acres_learn.eval.goto --driver reference --fields F53,F57 --runs 1 --out <dir> |
2 of 2; routes 1446 m and 307 m |
python -m acres_learn.eval.reference --energy-ref |
\(E_\text{ref}\) = 2689.77 J at 4.27 m/s |
python -m acres_learn.eval.reward_check --out <dir> |
Lane +0.517; corn −5.583; edge band +0.004 for each step |
Verification/check.sh |
All checks passed: 554 audited theorems, 25 s |
python Learning/tests/run_all.py |
107 tests pass in torchenv; 86 pass in ros2 (3 files skipped there) |