Skip to content

Field Scouting Task#

The field scouting task is the benchmark task of the learning stack. The vehicle is the Polaris Ranger. For the instruction "scout F39", the vehicle drives from the ICSC garage to the field. It drives around the field perimeter and does not go into the crop. Then it returns to the garage.

A mission can name more than one field. A planner then selects the order of the fields. It also selects the entry and the direction for each field. The plan must use the minimum fuel energy and the minimum energy lost to wheel slip. One model must do a mission with any number of fields, on any of the 59 fields of the map.

Status

The task exists in code. This includes the map products, the mission, the plan, the phases and the actions. It also includes the flat-policy observation, the rewards v1 and bounded, the terminations and the evaluation suite. This is the baseline formulation.

The section Framework Formulation gives the solution with a planner, a mission sequence, a residual driver and a shield. Those parts are specified, not built. Their Lean models and proofs exist (Verification).

A classical controller, reinforcement learning, imitation learning, vision-language-action models and a world-action model use this task. The evaluation scores all of them with the same code.

Version 1 keeps the soil and the weather constant during a mission. Version 2 lets them change. The last section describes version 2.

Purpose#

  • The task is the work of the vehicle. The Polaris Ranger of the laboratory drives on the lanes of the farm. Its drive-by-wire, LiDAR and camera models are calibrated against its logs (Polaris Ranger).
  • One behaviour applies to many fields. The drive around a field is the same behaviour for each field. Each field has a different shape, size and set of adjacent fields. Thus a field that is not in the training set is a clear test of generalisation.
  • The task includes a decision with a known optimum. The order of the fields is a routing problem. An exact solver gives the optimum. Thus the evaluation can measure the distance of each planner from the best plan.
  • The simulator gives the ground truth. The simulator knows the fuel, the slip energy, each crushed plant and each contact. No person must label the results.

Terms#

Symbol Meaning
\(\mathcal{F}\) The 59 fields of the ACRE map (fields.json). Each field is a polygon \(B_f\).
\(H\) The home pose: the ICSC garage spawn (survey cell u 474.6, v 592.2, heading north).
\(M = (H, F, \omega, \theta)\) A mission. It has home and the set \(F \subseteq \mathcal{F}\) of \(n = \lvert F \rvert\) fields. It has the conditions \(\omega\) (soil water, weather, time of day). It has the vehicle parameters \(\theta\) (with hardware shifts).
\(L_f\) The loop of field \(f\): the closed path that the vehicle drives around the field.
\(K_f\) The entries of \(f\): points on \(L_f\) where the lane network joins the loop.
\(E_\text{fuel}\) Fuel energy burnt (lower heating value), J. Litres = \(E_\text{fuel} / (43.4 \text{ MJ/kg} \times 0.745 \text{ kg/l})\).
\(E_\text{ground}\) Energy lost at the ground, J: tyre slip loss + side-slip scrub + soil rutting work (Slip + Lateral + Soil of the energy ledger, AcresPowerModel.h).
\(\lambda_g\) Weight of the ground losses in the objective (default 1).
\(\Delta t\) Control step, 0.1 s (10 Hz). The physics runs at 120 Hz.

Distances are in metres. Positions are ENU metres about the tile centre (the enu_m of places.json) in each file that the task writes. Energies are in joules.

These terms have one meaning on the task pages and on the learning pages:

Term Meaning
Phase One step of a mission: GoTo, Scout or Return. The code uses the name "skill" for a phase.
Leg The reference path of one phase.
Segment A part of a leg. A segment is one training episode of the residual.
Reference driver The pure-pursuit driver with a speed profile.
Base command The command of the reference driver in one control step.
Residual The bounded correction that a learned policy adds to the base command.
Shield The deterministic function that replaces an unsafe command.
Intervention A control step in which the output of the shield is different from its input.
Flat policy The first learned driver: one PPO network that outputs the full command for the whole mission.

Map Products#

A script builds the map products one time for each map version. Tests check them and a figure shows them. Each simulator tier and each model uses the same geometry. Scouting Map Products gives the files, the build steps and the measured statistics. This section gives the rules.

Ground classes. Each point of the tile is in one class:

Class Content Driving
Lane Grass lanes, roads, gravel, yards, dirt tracks (surface_polygons.json classes grass_lane, asphalt, concrete, gravel, dirt). Permitted
Verge Mown grass outside the fields (class grass). Permitted
Edge band The crop within 1.0 m of the loop where the loop must be inside a field. Permitted. The crop damage counts as unavoidable.
Crop The remaining area of each field, planted or fallow. Not permitted
Obstacle Buildings, bins, trees, fences, parked vehicles (collision geometry). Not permitted
Other Ground that is in no other class: the woodland floor south of US 52 and in the F52 woods, bare strips (57 ha). Not used by loops or transit

The edge band has two sources. The first source is the outer 2 m of a field along stretches that have no lane or verge outside. The second source is the crop that a curvature-limited corner cuts.

The loops are on lane or verge for a median 67 % of their length. Many fields share an edge with an adjacent field and have no lane between them. For 11 fields, at least 90 % of the loop is on lane or verge. For 17 fields, less than 50 % is on lane or verge (F05 17 %, F40 27 %, F41 33 %, F46 38 %).

These 17 fields are the plots F33-F45 and the fields on the outer edge of the tile. Each field needs the edge band at some point. Without the edge band, these fields cannot be scouted. The raster ground_classes.u8 holds the class of each 0.5 m cell.

Loops. For each field \(f\):

  1. The build simplifies the outline \(B_f\) with a 0.5 m tolerance. The survey outline is a staircase of 1.524 m steps. A closing with a 7.5 m radius bridges notches narrower than 15 m. The vehicle cannot drive into these notches. The result is the scouting outline. The loop and the checkpoints follow the scouting outline. The outline has no holes.
  2. The centre line of the loop is offset up to 2.5 m outside the outline. The offset is the largest value between 1.0 and 2.5 m that satisfies three conditions. The whole vehicle must be off the crop. The vehicle is 1.59 m wide over the tyres (1.31 m track, 0.28 m rear tyres). The centre line must be 4 m from obstacles where the ground permits it, else 3 m, else 2 m. The centre line must be 6 m inside the edge of the tile. On a narrow lane the loop keeps a smaller offset along the middle of the lane. It does not go into the crop.
  3. Where no such offset exists, the loop moves inside the field onto the edge band. The causes are an adjacent field, a building or a tree line. The centre line is then 1.0 m inside \(B_f\). The wheels are between 0.2 and 1.8 m inside. The build smooths the joints. The curvature is never more than 0.15 m⁻¹. The drive-by-wire permits 0.2 m⁻¹, a turning radius of 5 m.
  4. Each loop point has a class (lane, verge or edge band) and a speed limit. The limit is 5 m/s on lane and verge and 3 m/s on the edge band. The limit is lower in tight corners: the lateral acceleration is at most 1.5 m/s².
  5. The unavoidable crop area of the loop is \(A^\text{edge}_f\). It is 1.59 m for each metre of the edge-band stretches. The body of the vehicle pushes down each crop taller than 0.33 m. Corn, soybean and potato are all taller. Thus the crushed width is the full vehicle width, not the tyre tracks. The game and ACRES Core give the same measurement. The sum for all loops is 16,862 m². The median field has 183 m².
  6. Smoothing can move a loop across an obstacle. The loop is then bent around the obstacle with an arc of 7.5 m radius. A hairpin through crop at the end of a narrow strip is one turn of up to 11 m radius. It is not two corners of 7.5 m.

Checkpoints and coverage. Checkpoints \(c_{f,j}\) are at intervals of 10 m along the scouting outline. A checkpoint is covered when the reference point of the vehicle passes within 8 m of it. At 8 m, the camera and the LiDAR see the field edge well.

The coverage of a field is the share of its covered checkpoints. A field is scouted when two conditions are true. The coverage is at least 0.95. The vehicle is again at the entry where it joined the loop. The coverage refers to the field, not to the path. Thus a driver that finds a better line gets the same credit.

Entries and the transit graph. \(K_f\) holds the points where \(L_f\) meets the lane network. These are lane junctions within 6 m of the loop. A field has at most 6 entries, the most spread-out ones, and at least one. The transit graph \(G\) has nodes at lane junctions, at entries and at \(H\). Its edges follow the centre lines of lanes and roads (the NPC road graph and the grass-lane layer).

The lane network has 96 pieces. Least-cost routes over lane and verge join the pieces. Where no other connection exists, a connector can cross the edge band.

There are 13 such connectors with a total length of 643 m, among the plots F33-F45. Transit never uses crop. The graph has 505 nodes and 516 edges (30.9 km). It connects each field to \(H\).

Mission#

A mission \(M\) is complete when all these conditions are true:

  • Each field of \(F\) is scouted.
  • The vehicle is stopped (speed < 0.1 m/s).
  • The vehicle is within 3 m and 20° of \(H\). It can face away from the garage, as \(H\) does, or towards the garage.
  • The time is less than the time limit.

The time limit is

\[ T_\text{max} = 1.5 \times T_\text{ref}(M), \]

where \(T_\text{ref}\) is the time that the reference driver needs for the optimal plan.

A mission fails on each of these events:

  • A collision.
  • The centre of the vehicle is more than 6 m into the crop.
  • The vehicle is stuck (see below).
  • The vehicle leaves the map: a wheel is off the tile, where the terrain ends.
  • The time is more than the time limit.

The crop depth is the distance to the nearest cell of ground_classes.u8 that is not crop. Thus the edge band and the corner cuts do not count. They are up to 8.4 m inside a field (F09, along the edge of the tile).

Heading at home. The first version of this specification accepted only the heading of \(H\). The vehicle turns no tighter than 5 m and the garage is 3 m behind \(H\). Thus the vehicle comes back with its front towards the garage.

No forward path of that radius ends within 20° of the heading of \(H\) without contact between the body and the building (Scouting Baselines). With the reverse gear, the vehicle can drive backwards into the garage. The opposite heading stays accepted (decision of September 2026). A reverse manoeuvre adjacent to a building gives no benefit to the task.

Stuck. The vehicle is stuck when two conditions are true for 3 s. The mean longitudinal slip of all driven wheels is above 0.4. The speed is below 0.3 m/s. The planner uses a stricter limit.

It removes a lane, loop or edge from the options of the plan when the reference drive on it has one of three results. The vehicle is stuck. The sustained slip is above 0.3. The rut depth is above 8 cm.

Scale. The field areas are 0.1 to 26 ha (median 0.8 ha). The perimeters are 135-2995 m (median 416 m). The straight-line distances from \(H\) are 24-956 m (median 434 m). The loops are 145-2202 m long (median 387 m), 34.9 km in total.

A mission with five fields has approximately 2 km of loops and 2 km of transit. This is approximately 17 minutes at 4 m/s, or 10,000 control steps. A policy cannot learn the credit assignment for 10,000 steps from one signal at the end of the mission. Thus the task has a plan and phases.

Plan#

The plan sets three items for each field \(f \in F\): the order \(\sigma\), the entry \(k_f \in K_f\) and the direction \(d_f \in \{\text{clockwise}, \text{anticlockwise}\}\). The mission path is then

\[ H \to k_{\sigma_1} \xrightarrow{L_{\sigma_1}} k_{\sigma_1} \to k_{\sigma_2} \xrightarrow{L_{\sigma_2}} k_{\sigma_2} \to \cdots \to k_{\sigma_n} \xrightarrow{L_{\sigma_n}} k_{\sigma_n} \to H. \]

The vehicle drives each loop \(L_{\sigma_i}\) in its direction \(d_{\sigma_i}\).

A transit leg can use the loops at its ends. The vehicle can join a loop before its entry and drive along the loop to the entry. After it comes back to the entry, it can drive on along the loop to the point where it leaves. This is the method to turn the vehicle without the reverse gear.

Some lanes are dead ends, and some of them have no loop that gives a way around. On such a lane, a leg turns with the reverse gear. The manoeuvre is a three-point turn: the shortest Reeds–Shepp path (Reeds and Shepp 1990) of 6 m radius on permitted ground. Each backward metre counts as two metres. Each change of direction counts as 10 m. The planner prices such a U-turn as 60 m of lane (Scouting Baselines).

Objective.

\[ J(\text{plan}) = E_\text{fuel} + \lambda_g E_\text{ground} \qquad \text{subject to no stuck-limited edge, and } T \le T_\text{max}. \]

The fuel pays for the ground losses one time, because the engine delivers that work. \(\lambda_g\) counts the ground losses a second time. The reasons are ruts in the lanes, soil compaction and the risk of a stuck vehicle. The default is \(\lambda_g = 1\). The evaluation also reports \(E_\text{fuel}\) and \(E_\text{ground}\) separately. Thus a later analysis can use a different weight.

Costs come from simulation. The reference driver drives each transit-graph edge one time in each direction in ACRES Core. It also drives each loop one time in each direction, from a canonical start. The drives use the conditions \(\omega\) of the mission. Each drive records \(E_\text{fuel}\), \(E_\text{ground}\), the time and the peak sustained slip.

The cost of a transit leg is the cheapest path through \(G\) (Dijkstra on the edge costs). The cost of a loop depends only on its direction (slopes, the side of the turns). It does not depend on the start point. For 59 fields, one condition set needs some thousands of short simulations, which is minutes of CPU time.

Structure. In version 1 the conditions do not change during the mission. Thus a loop has the same cost in each order, and \(\sum_f J(L_f, d_f)\) depends only on the directions. The order and the entries change only the transit legs. The problem is an asymmetric generalised travelling-salesman problem. It is asymmetric because uphill and downhill costs are different.

It is a generalised problem because each field is a cluster of entries. The problem is small. Exact dynamic programming over subsets solves \(n \le 12\) in seconds, with up to 6 entries and 2 directions for each field. The routing solver of OR-Tools with disjunctions solves larger \(n\). A lower bound then gives the gap.

solvers.exact_dp needs at least one field. For a mission without fields it raises ValueError. The callers handle that mission before the call. Lean proves that the dynamic programme is exact (Verification).

Planners compared:

Planner Function
Optimal Exact DP (\(n \le 12\)) or OR-Tools with its gap. It defines the regret of the other planners.
Nearest first Greedy: the unvisited field with the nearest cheapest entry, anticlockwise.
As listed The order of the fields in the instruction, nearest entry, anticlockwise.
Learned An attention-based routing policy (Kool et al. 2019), trained with REINFORCE on the cost matrices.
VLM A vision-language model that gets the map, the fields and cost queries for each edge as tools (not the solver).

Regret is the cost of a planner above the optimum. The same driver drives both plans: \(\text{regret} = (J - J^*) / J^*\).

Phases#

A mission has three types of phase. Each phase has a leg, which is its reference path.

Phase Goal Ends When
GoTo Follow the transit path to an entry (or to \(H\)). The vehicle is within 2 m of the target. Or its projection is on the last 2 m of the path, inside the corridor. Or the vehicle is in the last 12 m and within 2 m of the path of the next phase. That point must be in the first 12 m of the next path.
Scout Drive \(L_f\) from \(k_f\) in direction \(d_f\) back to \(k_f\). The vehicle is again within 2 m of the entry, with its progress at the end of the loop. The field is scouted if its coverage is then at least 0.95.
Return GoTo \(H\) and stop. The completion test of the mission is true.

Phase sequence. The episode bookkeeping (mission.EpisodeTracker) moves to the next phase at each milestone: an entry reached, a loop closed, the stop at \(H\). A learned policy never selects the phase. A mission with \(n\) fields has the sequence GoTo, Scout for each field, then Return. This code exists.

Mission automaton. The Framework specifies the same sequence as a mission automaton with the states IDLE, GOTO(i), SCOUT(i), RETURN, DONE and FAILED. The Lean model and its proofs exist. The automaton as a separate module is specified, not built. These rules apply.

  • The time limit is not an event of the automaton.
  • When a failure and a milestone occur in one control step, the failure has priority.
  • HOME_STOPPED is the completion test of the mission: each requested field is scouted and the vehicle is stopped at \(H\). A field is scouted when its coverage is at least 0.95.

The same driver for each field. A driver gets the reference path and the map around the vehicle. It never gets the name of the field. A new field is a new path through new surroundings. The flat policy drives all three phases with one network.

The residual driver of the framework formulation uses the reference driver on the mission path and adds its correction. The phase is an input of the residual policy.

Actions#

A driver sends two commands in each control step (0.1 s). They are the drive-by-wire commands of the real Polaris. Thus the same output drives the ds_dbw interface of the vehicle. The flat policy outputs two numbers in [−1, 1], which map to the commands:

Index Command Range
0 SteeringCmd, curvature mode −0.2 to 0.2 m⁻¹
1 UlcCmd, velocity mode, signed (negative is backwards) −1.5 to 6 m/s: −1 → −1.5 m/s, 0 → stop, 1 → 6 m/s (two linear pieces)

The SteeringCmd contains the rate and acceleration limits of the steering wheel. The values are 171.9 deg/s and 1000 deg/s², as the path follower of the laboratory sends them. The default rate of the firmware is 100 deg/s. With that rate, the steering needs 3.6 s from straight to full lock.

Reverse. The sign of the velocity selects the direction. The gear follows the sign, as the Dataspeed ULC does with enable_shift ("allow ULC to issue gear shift commands"). The command mapping of the task does this (spec.GearManager). Thus each simulator and the vehicle get the same GearCmd and UlcCmd messages. The sequence is:

  1. A velocity against the engaged gear by more than 0.2 m/s is a request for the other gear. This is a backward velocity in L or a forward velocity in R.
  2. The ULC gets 0 until the vehicle is stopped (below 0.1 m/s). The gear actuator shifts only near standstill, through neutral, in approximately 1 s.
  3. The task sends one GearCmd (ds_dbw Gear R = 2 or L = 5).
  4. The ULC gets 0 for 1.1 s more while the shift occurs.
  5. The ULC gets the velocity.

A command inside the band of ±0.2 m/s keeps the gear. Thus a policy with a velocity near zero does not shift again and again. The backward velocity is at most 1.5 m/s. This is walking speed, because the view behind the vehicle is limited.

One signed axis. The action space stays a box of two continuous numbers. Each method uses this box without change: the Gaussian head of PPO, the chunked imitation models and the action experts of the VLAs. A separate gear action makes the action space hybrid. Such a discrete action is important only at standstill, but the policy samples it in each step. It is almost always "keep the gear".

Thus it is hard to explore and easy to change by accident. The two linear pieces keep 0 at the centre of the axis. The forward speeds keep 80 % of the axis. The gear logic adds a stop and a pause. The policy sees this result: the gear is in the proprioception and the path gives the direction of travel. The policy pays for it through the time term and the gear term.

Models that predict chunks (ACT, Diffusion Policy, SmolVLA, π0.5, FastWAM) predict 8 steps (0.8 s). They execute the steps at 10 Hz.

The residual driver does not output the full command. It outputs a bounded correction of the base command (Residual Action).

Observations#

This section gives the observation of the flat policy. Residual-Policy Observation gives the differences for the residual policy.

Flat-Policy Observation#

Group Content Receiver
Proprioception Speed, yaw rate, lateral acceleration, measured steering wheel angle, pitch, roll, the four wheel speeds, the gear, the previous action. All methods (the real vehicle has all of these).
Path The next 16 reference-path points at 2 m intervals in the body frame. Each point has its class (lane / verge / edge band), speed limit and direction of travel (+1 forwards, −1 backwards). Then: the distance that remains in the phase, the phase (one-hot), the remaining mission time as a fraction. All methods.
Map crop An egocentric raster of 32 × 32 m at 0.5 m. The vehicle is 8 m from the bottom and points up. Channels: crop, edge band, lane/verge, obstacle, soil water. The RL teacher gets channels. Vision models get a rendered colour image with the route on it.
LiDAR 360 ranges (1° apart, 30 m) from the ray-cast LiDAR at the Helios mount. The RL teacher, and models that use LiDAR.
Camera The front camera, resized to 224 × 224 (FastWAM on RoboTwin also uses 384 × 384). Vision models.
Instruction "Scout F39 clockwise from its north-west entry, then return to the ICSC garage." Language models. The plan gives the text.

The soil-water channel is the only privileged input. The real vehicle does not measure soil water. The student models (the vision models) see the soil only through the camera, the wheel speeds and the response of the vehicle.

Rewards#

This section gives the rewards v1 and bounded of the baseline formulation. The residual policy trains on a third variant (Residual Reward).

The reward of a driving policy in one control step is a sum of terms. The log records each term separately. Thus a diagnosis of the training can use each term. The scales follow four rules:

  1. Driving the path at the reference speed on a dry lane gives a clearly positive reward (+0.45 for each step). Thus the policy never learns that a stop is best.
  2. A shortcut through crop costs more than it gains (−5 for each step against +1 for progress).
  3. Energy differences decide between behaviours that are otherwise equal. Faster driving and wetter lines cost more. The objective of the planner counts them in the same way.
  4. Each signal arrives within some seconds of the behaviour that caused it. The progress term and the checkpoints give this. The sparse bonuses only mark milestones.

Reward Terms#

\[ r_t = r^\text{prog}_t + r^\text{cover}_t + r^\text{time}_t + r^\text{energy}_t + r^\text{crop}_t + r^\text{presence}_t + r^\text{track}_t + r^\text{smooth}_t + r^\text{safe}_t + r^\text{gear}_t + r^\text{event}_t \]
Term Definition Weight Scale at 4 m/s on Dry Lane
Progress \(w_p \, \Delta s^*_t\). \(s^*\) is the furthest arc length reached along the leg (monotone). The direction of travel does not matter: a backward stretch of a three-point turn counts. \(w_p = 2.5\) m⁻¹ +1.0 per step
Coverage \(b_c\) for each newly covered checkpoint (Scout only). A step with the vehicle partly in crop outside the edge band pays nothing, but the checkpoint counts as covered. \(b_c = 0.5\) +0.5 each 10 m
Time \(-c_t\) \(c_t = 0.25\) −0.25 per step
Energy \(-w_E \, (\Delta E_\text{fuel} + \lambda_g \Delta E_\text{ground}) / E_\text{ref}\). \(E_\text{ref}\) is the energy of the reference driver per step at 4 m/s on dry grass lane. \(w_E = 0.3\) −0.3 per step
Crop \(-w_c \, \Delta A^\text{crop}_t / A_\text{ref} - w_{ce} \, \Delta A^\text{edge}_t / A_\text{ref}\): newly crushed crop area outside and inside the edge band, from the farm model. \(A_\text{ref} = 1.59 \text{ m} \times 0.4 \text{ m} = 0.636 \text{ m}^2\). \(w_c = 5\), \(w_{ce} = 0.2\) 0 on lane; −0.2 on edge band; −5 in crop
Crop presence \(-P\) on each step on which a part of the shrunk plan box is on crop outside the edge band (ground class crop). The plan box is 0.5 m behind the rear axle to 3.42 m ahead, 1.6 m wide. It is shrunk by 0.5 m on each side. Crushed crop counts too. \(P = 1\), shrink 0.5 m 0 on lane and on the edge band; −1 in crop
Tracking \(-w_e \, \max(0, \lvert e_y \rvert - h)^2\). \(e_y\) is the cross-track error. \(h\) is the half-width of the drivable corridor at that point minus 0.8 m. \(w_e = 0.5\) m⁻² 0 inside the corridor
Smoothness \(-w_a \lVert a_t - a_{t-1} \rVert^2\) \(w_a = 0.05\) ≈ 0 when steady
Safety \(-w_o \, v \max(0, 1 - (d_\text{min} - 2)/8)\). \(d_\text{min}\) is the distance to the nearest vehicle, worker or obstacle. \(w_o = 0.2\) s/m 0 with nothing within 10 m
Gear \(-w_g\) for each GearCmd sent (a change between L and R). \(w_g = 0.5\) 0 without a change; −1 for the two changes of a three-point turn
Events An entry reached +10. A field scouted +50. Stopped at home +50. Failures: see below.

Failures end the episode with a penalty: collision −100, stuck −50, more than 6 m inside crop −50, leaving the map −50. The penalties of the bounded variant are larger (below). A timeout is a truncation, not a failure. PPO bootstraps the value. The remaining time is an observation of the flat policy. Thus the flat policy can see that a timeout is near.

Discount. \(\gamma = 0.995\). This is an effective horizon of 200 steps (20 s). GAE uses \(\lambda = 0.95\). The dense terms carry credit further than 20 s. PPO normalises the rewards by a running estimate of their scale.

Protection against exploits. Progress is monotone. Coverage depends on the field edge. Crop and energy come from the physics and the farm model, not from a report of the policy. A circle on the spot earns nothing and costs time. A stop earns −0.25 for each step.

Why These Numbers#

  • Progress is a potential. \(r^\text{prog}\) is the difference of the potential \(\Phi = w_p s^*\) (Ng, Harada and Russell 1999). It guides the policy. \(s^*\) is the furthest point reached. Thus a drive back and forth earns nothing. A corner cut is worth only the arc length that it skips. The crop term and the corridor outweigh that.
  • Time and energy balance the speed. Progress pays for each metre, not for each second. Without the time cost, a crawl is the best policy. The time cost rewards speed. The energy cost increases with speed and slip, and it penalises speed. The balance is near the speed limits of the loop. For the behaviour of the reference driver, the step reward is \(1.0 - 0.25 - 0.3 = +0.45\).
  • Crop is dominant but not absolute. A step fully in crop costs \(5 \times 1.59 \times 0.4 / A_\text{ref} = 5\) before the other terms. Thus a shortcut loses approximately 4 for each step net. The penalty is dense: it is per crushed area, from the record of the farm model. Thus the gradient leads away from the crop immediately. The 6 m termination only stops runs that are clearly lost.
  • Presence in crop costs, not only crushing. The crop term pays for newly crushed area. Thus, without a second term, a slow creep through a field costs little more than the energy. A second drive over flattened crop costs nothing more than the energy. The presence term charges \(P = 1\) for each step with a part of the vehicle in crop outside the edge band. The term is of September 2026. A formal check of the reward (the Lean proof presence_term_fix) needs \(P > w_E k\). \(k\) is the largest energy per step over \(E_\text{ref}\) (approximately 1.2 measured). Then the energy that a path through crop saves can never pay for the presence. \(P = 1\) is sufficient up to \(k \approx 3.3\). Such a step also pays no coverage bonus. Thus a checkpoint is never cheaper to reach through the crop.
  • The plan box is shrunk for the presence test. The edge band has no presence charge. The band is only 0.2 m wider than the vehicle on each side. Thus the plan box is shrunk by 0.5 m on each side before the test (crop_presence_shrink_m). Without the shrink, the reference driver had a part of its box in crop on 41 % of the steps of the longest edge-band stretch (F14). The edge band then gave −0.39 for each step, which is worse than a stop. With the shrink, no point of a planned loop triggers the term. The drift of the reference driver past the band costs −0.015 for each step on the edge band. A position more than 0.5 m past the band still pays the full \(P\). A path through a field also pays the full \(P\). Thus the Lean claim is unchanged (Scouting Baselines).
  • The edge band is almost free. Its crop is the unavoidable cost of the loop. It counts at 1/25 of real crop. Thus the policy prefers a lane where both exist, but it does not avoid the edge band where the edge band is the loop. At the speeds of the reference driver there, the crop terms leave the step reward slightly positive (approximately +0.02 for each step, Scouting Baselines). This stays unchanged (decision of September 2026). The edge band is the loop, not an alternative to it. A larger margin is possible only with a smaller price for crop. With the shrunk presence term, the step of the reference driver there is +0.009. A policy still gets a reward when it holds the middle of the band.
  • Reverse is permitted, not encouraged. Progress also counts along backward stretches. Thus a three-point turn is worth its metres where the path has one. The gear term (−0.5 for each change) and the stops for the shifts prevent shifts where the path does not need them.
  • The reference normalises the energy. \(E_\text{ref}\) makes the term approximately −0.3 for each step in normal driving. A wet patch that doubles the slip loss gives a difference that the policy can compare with progress.
  • Coverage keeps the scout honest. The checkpoints are on the field edge, not on the path. Thus the policy earns reward when it sees the field. A loop that cuts inside the edge band loses coverage and pays for crop.
  • Milestones are small compared with the dense terms. A 416 m loop earns approximately 1040 in progress and 21 in coverage. The +50 at its end is a marker. It is not the signal that the policy learns from.

Failures Never Pay#

In the first version, a failure costs less than a continued drive in a bad state. Thus a crash was the cheapest exit from a mission that went wrong. The stage-2 training runs of the flat policy learned to take this exit (Results).

The measurement used the best stage-1 policy of scout3. It drove the eight lane-bound missions eight times each with the training noise:

State Cost
A step that crushes crop above 4 m/s −6.5 on average (1 % of them below −20)
A step more than 5 m outside the corridor −29 (the tracking term is quadratic: −50 for each step at the 10 m lost_m limit)
The worst 2 s of a run −25 for each step on average
The worst 20 s of a run −8.3 for each step on average
Discounted value of a continued drive while crushing at speed (\(\gamma\) = 0.995) −450 (5 % below −1140)
Discounted value of a continued drive off the corridor −277
A collision −100
Each other failure −50

The discounted values are the returns-to-go of the runs.

The bounds below use exact numbers and raw rewards. Floating-point rounding can reverse adjacent values or reach the floor. PPO reward scaling and clipping are outside these proofs. See Review Findings to Resolve.

The option env.rewards = "bounded" of a training run selects the bounded variant (configs/scouting_v1.json, rewards_bounded). The default "v1" keeps the first version. Thus earlier runs stay reproducible. The terms are the same. Two items change:

  • The dense reward of a step has a lower bound. \(r\) is the sum of each term but the events. At or above the knee −2, \(r' = r\). Below the knee, \(r' = -2 - 3x / (x + 3)\) with \(x = -2 - r\). \(r'\) approaches the floor −5 but never reaches it over exact numbers (rewards.bound; the log has the term bound, \(r' - r\)). The exact-number map has slope 1 at the knee and is strictly increasing. Thus each comparison of two steps has the same result as before. Normal driving, the edge band, a stop and a step partly in crop are unchanged. The map compresses only errors at speed. Crushing crop at 4 m/s goes from approximately −6.4 to −3.8. A position 5 m off the corridor goes from −29 to −4.7. Crop still never pays: −3.6 for each step through a corn field at 4 m/s against +0.52 on the lane (eval.reward_check).
  • Failure penalties. A collision costs −1250. Stuck, deep in crop, leaving the map, "lost" and "stalled" cost −1000. A continuation without failures has a lower bound of \(-5 / (1 - \gamma) = -1000\) over exact numbers.

The Lean proofs failure_never_pays and failure_never_pays_v2 compare a penalty with a continuation that has non-negative events and a terminal value at least that penalty. A later, larger failure can violate this assumption. See Verification. The PPO configuration check rejects the variant when the discount makes the failures cheaper than the floor over \(1 - \gamma\). It does not establish these properties after the historical PPO reward scaling and clipping.

Doubles can reach the floor at extreme inputs and reverse adjacent reward values through rounding. The exact-number monotonicity claim does not apply to every floating-point input.

These proofs compare the penalty alone. The dense reward of the step with the failure is not in the comparison. A later proof found the limit of the statement. A failure on a step with a positive dense reward can be worth more than a continuation at the floor (failure_pays_on_a_good_step). Such a continuation needs a position 95 m outside the corridor, which the "lost" limit of 10 m excludes in training. The residual reward has larger penalties and a proof that counts each dense reward.

The bound keeps the penalties finite. With the terms of the first version, the worst step is the tracking term at the lost_m limit, −50. The failures then need −10,000. A penalty scaled to the remaining time does not help. A time limit is a truncation, and PPO bootstraps its value. Thus the value sees the horizon of the discount (200 steps), not the remaining time.

A stop stays a bad choice. A stop on the lane costs −0.31 for each step (time and idle fuel). This is −62 in discounted value. A drive along the lane at the reference speed is worth +103. The cheapest failure is −1000.

A stop is better than a failure, which is correct. A stop does not end the episode, because the stall end stays off. Thus no failure competes with it. Driving is the best of the three. The shaping term of the residual reward keeps this order.

Planner Rewards#

A learned planner or a VLM planner works over whole phases. This is a semi-Markov decision process with one decision for each field. The reward for field \(f\) with entry \(k\) and direction \(d\) is the realised cost:

\[ R = -\frac{E_\text{fuel} + \lambda_g E_\text{ground}}{E_\text{norm}} + 1_\text{scouted} \cdot 1. \]

\(E_\text{norm}\) is the cost of the optimal plan divided by \(n\). Thus one field of an optimal mission costs approximately −1. A completed mission gives a final +1. A failure gives −1 for each field that is not scouted. The evaluation scores a planner by its regret, not by this reward.

Terminations#

Event Type Applies To
The mission is complete (stopped at \(H\), each field scouted). Termination (success) All episodes that end at \(H\)
Collision, stuck, more than 6 m inside crop, a wheel off the tile. Termination (failure) All episodes
"Lost": more than env.lost_m (10 m) outside the corridor. Termination (failure), training aid Training
"Stalled": stopped for env.stall_s (20 s). Termination (failure), training aid, off by default Training
The time limit \(1.5 \, T_\text{ref}\). Truncation All episodes
The end of the segment. Truncation Segment episodes of the residual (specified, not built)

A termination ends the return: the value after it is zero. A truncation does not end the task. PPO bootstraps the value from the last observation. The failure tests have this order when two occur in one step: collision, stuck, deep in crop, left the map.

The training aids are not part of the evaluation. In the first version, "lost" costs −50 and "stalled" costs env.stall_penalty (−100 in the two PPO configurations). In the bounded variant, "lost" and "stalled" cost −1000. In the residual variant, they cost −1200.

When a failure and a milestone occur in one control step, the failure has priority. The episode ends as a failure.

Segment episodes. One training episode of the residual is a segment of a leg. The Framework gives the rules. This part is specified, not built.

Framework Formulation#

State: specified, not built. The Lean models and their proofs exist. The Framework is the specification. If this section and the framework page disagree, the framework page is correct.

The sections above give the baseline formulation: one flat policy outputs the full command for the whole mission. That formulation completed no whole mission (Results). The framework formulation gives the long horizon to classical parts. A learned policy only corrects the command of the reference driver.

Solution in Four Parts#

Part Function Decision State
Planner Exact dynamic programme over subsets of fields The order of the fields, and the entry and the direction of each field Exists. Lean proves that it is exact
Mission sequence Mission automaton The phase and the leg The tracker has the transitions. Lean proves the automaton
Residual driver The reference driver and a bounded residual The curvature and the speed in each control step The reference driver exists. The residual is specified
Shield Action replacement The command that the vehicle receives Specified. Lean proves the rules

A mission has this sequence.

  1. The planner solves the mission. The result is the plan and the reference path of each leg.
  2. The automaton gets the event START. It goes from IDLE to GOTO(0).
  3. In each control step, the reference driver calculates the base command for the current leg.
  4. The residual law adds the bounded correction of the policy to the base command.
  5. The launch hold, then the shield, then the gear logic process the command. The drive-by-wire gets the result.
  6. A milestone test of the task raises an event. The automaton goes to the next phase.
  7. HOME_STOPPED ends the mission in DONE. A failure ends the mission in FAILED.

The policy has no decision about the order, the phase or the path. With the gain \(\alpha = 0\) the driver is the reference driver. The reference driver completes 59 of 59 single-field missions without the shield (Scouting Baselines). The acceptance rule of the framework requires the same result with the shield.

Residual Action#

The reference driver calculates the base command \((\kappa_\text{PP}, v_\text{PP})\) in each control step. The policy outputs \(a = (a_0, a_1)\). The command is the base command plus a bounded correction:

\[ \kappa = \kappa_\text{PP} + g \, a_0 \, \Delta\kappa, \qquad v = v_\text{PP} + g \, a_1 \, \Delta v. \]
Symbol Value
\(g\) The gain \(\alpha\) after the clamp to \([0, 1]\). A NaN gain gives 0
\(a_0, a_1\) The policy output: a NaN becomes 0, then the clip to \([-1, 1]\)
\(\Delta\kappa\) 0.05 m⁻¹
\(\Delta v\) \(\min(1.0 \text{ m/s},\; 0.3 \, \lvert v_\text{PP} \rvert)\). It is 0 while the measured speed is below 0.3 m/s. It is 0 if the measured speed or the base speed is not finite
  • With a gain that is not above 0, the command is the base command, bit for bit.
  • The speed keeps the sign of the base speed. Thus the residual cannot ask for a gear change.
  • If the base speed is 0, the speed is 0. Thus the residual cannot move a vehicle that the base command stops.
  • The law does not clip the command. The command goes to the shield.

Residual-Policy Observation#

Group Size Difference from the Flat-Policy Observation
Proprioception 13 The previous action is the final command after the shield
Path 136 16 waypoints as a direction and \(\log(1 + d)\); then the cross-track error, the heading error, the base command, the speed cap and the phase
Map crop 4 × 64 × 64 No soil-water channel
LiDAR 360 The same
Critic 2 Training only, for the critic only: \(N_\text{ref}(s^*)\) in minutes and the end-of-mission flag

The actor gets no distance to the goal and no remaining-time fraction. The observation has no Lean model.

Residual Reward#

The option env.rewards = "residual" selects this variant. The constants are in configs/scouting_v1.json, section rewards_residual.

\[ r'_t = B(d_t) - w_s \, 1_\text{intervention} + e_t + F_t \]
Symbol Definition Value
\(d_t\) The dense sum: each reward term but the events
\(B\) The bounded map of Failures Never Pay Knee −2, floor −5
\(w_s\) The shield term: the cost of a step with an intervention 0.5
\(e_t\) The events See below
\(F_t\) The shaping term \(c_\Phi = 0.55\), \(\gamma = 0.995\)

The reward subtracts the shield term after the bound \(B\). Thus it has its full value in each state. Inside the bound it is worth only 0.005 on a step 5 m outside the corridor.

Event Value
Stopped at \(H\) +50
Entry reached, field scouted (whole-mission episodes only) +10, +50
Collision −1500
Stuck, deep in crop, left the map, lost, stalled −1200 each

The failures of this variant cost more than the failures of the bounded variant (−1000 and −1250). The shield term moves the floor of a step from −5 to −5.5. The failure values satisfy two inequalities.

  1. \(F \le (-5 - 0.5) / (1 - 0.995) = -1100\). The comparison needs a continuation value of at least \(F - 5.5\). It does not cover a later, more severe failure. The value −1200 gives a margin of 100 against the bound −1100.
  2. \(D_\text{max} + 50 + F < -1100\), with \(D_\text{max} = 2.5 \times 14 + 0.5 \times 4 - 0.25 = 36.75\). Thus a step with a failure is worth less than each continuation without a failure. The step has its dense reward and one milestone. \(36.75 + 50 - 1200 = -1113.25\). The margin is 13.25. The value −1150 does not satisfy the inequality.

\(D_\text{max}\) is the largest dense reward of one control step. It uses two limits from the code. The progress of one step is at most 14 m. One step covers at most 4 new checkpoints. These limits are hypotheses of the proof.

The verification gate checks the configuration and the checkpoint geometry of the current map. Runtime guards for the residual variant reject inputs that exceed the limits. A new map must pass the geometry check.

Residual PPO must use raw rewards, with gamma=0.995, normalize_rewards=false and reward_clip=null. The configuration validator rejects residual training until the production framework exists. The cast to float32 still introduces rounding. Lean does not prove the effects of that rounding on training.

At a failure the shaping term gives back the remaining cost: the step with the failure carries \(-1200 + 0.55 \, N_\text{ref}\). This value is positive above 2182 reference steps. The statements about returns in Failures Never Pay stay true.

Shaping Term#

\[ F_t = \gamma \, \Phi(s^*_{t+1}) - \Phi(s^*_t), \qquad \Phi(s^*) = -c_\Phi \, N_\text{ref}(s^*), \qquad c_\Phi = 0.55 \]

The shaping term is potential-based (Ng, Harada and Russell 1999). \(s^*\) is the furthest arc length that the vehicle reached on the episode path. \(N_\text{ref}(s^*)\) is the remaining time of the reference speed profile from \(s^*\) to the end of the episode path, in control steps.

The table \(N_\text{ref}\). A table gives \(N_\text{ref}\) at each point of the episode path.

Item Definition
Value at the last point 0
Value at point \(i\) The sum of the times of the path segments after \(i\), divided by the control step (0.1 s). The sum starts at the end of the path
Segment time A term of deploy.mission.path_time_s: the length of the segment divided by the mean reference speed of the segment
Between the points numpy.interp(s*, arcs, N): linear between the points, the total before the first point, 0 at and after the last point
Rounding None. \(N_\text{ref}\) is a real number

The rules at the end of an episode.

Case Rule
Termination (a failure, or the stop at \(H\)) \(\Phi\) of the next state is 0. Thus \(F_t = -\Phi(s^*_t) = 0.55 \, N_\text{ref}(s^*_t)\)
Truncation (the end of a segment, the time limit) \(F_t\) uses the real next state. PPO bootstraps with the critic of the shaped problem. The exact value of that critic is \(V(s_T) + 0.55 \, N_\text{ref}(s^*_T)\)
End of a segment The potential is not 0 there. Do not set \(\Phi(s_T) = 0\) at a truncation
Validation and evaluation \(F_t = 0\). The reports give the unshaped return

With these rules the shaped return is the unshaped return minus \(\Phi(s_0)\). This difference does not depend on the policy. Thus the shaped task and the unshaped task have the same optimal policies. The potential is a function of the progress \(s^*\), not of the vehicle position. Thus the critic gets \(N_\text{ref}(s^*)\) as an input.

A stop in the shaped reward. The shaping term adds \((1 - \gamma) \, c_\Phi \, N_\text{ref}(s^*) = 0.00275 \, N_\text{ref}\) to a step without progress. Thus the shaped step of a stop is \(-0.31 + 0.00275 \, N_\text{ref}\), which is positive above 112.7 reference steps. The driving step from the same progress point has the same term.

Thus the term has no effect on the difference of the two steps. The shaping adds \(\gamma \, c_\Phi \, \Delta N_\text{ref} = 0.547\) for each reference step of progress to the driving step. The order of the returns (failure < stop < lane) is the same with and without the shaping term.

Shield#

The shield is a deterministic function without state. It gets the proposed command and returns the command that the vehicle receives. It is after the launch hold and before the gear logic. Training, evaluation and deployment use the same function.

Rule Function
0. Unusable inputs If the base command, the cross-track error, a half width, the direction or the speed cap is not finite, the output is a stop. A NaN obstacle distance also gives a stop
1. Proposal that is not finite A proposed value that is NaN or infinite becomes the value of the base command
2. Geofence Outside the fence, the driver can only steer back relative to pure pursuit. The speed is between 0 and the base speed. The fence is 1.1 m inside the edge of the drivable corridor. A backward stretch exchanges the two sides
3. Corner speed \(\lvert v \rvert \le \max(\min(c, \sqrt{1.5 / \lvert \kappa \rvert}), \lvert v_b \rvert)\), with \(\kappa\) the output curvature and \(c\) the speed cap of the path
4. Obstacle stop distance \(\lvert v \rvert \le \text{stopCap}(d_\text{obs} - 1.5 \text{ m})\): the largest speed that can stop before the margin, with a latency of 0.5 s and a deceleration of 0.8 m/s²
5. Actuator limits Curvature in ±0.2 m⁻¹, speed in −1.5 to 6 m/s

An intervention is a control step in which the output curvature or the output speed is not equal to the proposal. The backup action is the pure-pursuit command inside the actuator limits.

The reward of the flat policy has no term for the corner speed. In the framework, the corner speed is a limit of the shield, not a reward term.

Design limit. On 30.7 % of the Scout path points, at least one side of the corridor has a fence of 0. On 18.4 % of the points, the two sides have a fence of 0. There the residual can steer to one side only and cannot add speed.

Proved Task Properties#

Lean proves these properties of the models. Verification gives the theorems and the assumptions.

Task Property Statement Proof
The plan is optimal The dynamic programme returns a feasible plan of minimum cost, or no plan if none is feasible Optimality
A table error has a bounded cost If each cost is wrong by at most \(\varepsilon\), the plan costs at most \(2 (2n + 1) \varepsilon\) more than the optimum Regret
A plan from an untrusted source The checker accepts a proposal if and only if it is a feasible plan Checker
The phase sequence ends At most \(2n + 2\) events change the phase. No state is a deadlock Mission Automaton
DONE has a meaning DONE occurs only after a closed loop for each field and the stop at \(H\) Mission Automaton
The residual is bounded The command is within \(g \, \Delta\) of the base command. With \(\alpha = 0\) it is the base command Residual Law
The residual cannot reverse The speed keeps the sign of the base speed Residual Law
The shield output obeys each rule The output satisfies each rule. A safe proposal passes without a change Shield
The shield never adds speed The output speed is between 0 and the proposed speed Shield
Failures never pay A failure is worth less than each continuation without a failure Residual Reward
Crop never pays A step in crop earns less than the lane step of the same progress Residual Reward
A stop is not the best policy The returns have the order failure < stop < lane Residual Reward
The shaping term is neutral The shaped task and the unshaped task have the same optimal policies Reward Shaping

Lean does not prove these properties.

  • The policy network has no proof.
  • The proofs of the shield are about the command. They do not prove that the vehicle stays out of the crop or stops in time. The map, the LiDAR and the pose are trusted inputs.
  • The production code of the residual, the shield and the residual reward does not exist. It must pass the golden vectors.
  • The completion test HOME_STOPPED and the other milestone tests have no proof.

Splits#

The fields are split one time. The split uses strata of area and of the share of the loop on lane or verge.

Split Fields Use
Train 43 All training missions
Validation 6 Model selection
Test 10, with at least two under 50 % lane-bound Never seen in training

Test missions have three types. In the first type, all fields are from the training split. This tests the basic function. In the second type, one held-out field is among training fields.

This tests generalisation. In the third type, no field is from the training split. Scouting Baselines lists the fields of each split.

Evaluation and Metrics#

The evaluation uses a fixed, seeded suite of 200 missions: \(n \in \{1, 2, 3, 5, 8\}\), 40 missions for each size. Each mission runs under three condition sets: dry; wet and constant; hardware shift. State-based policies run in ACRES Core.

A subset is cross-checked in Unreal. Vision models run in Unreal with rendered cameras, in lockstep. Thus a slow model has no penalty for its inference time. The report gives the inference time separately.

Mission success is a constraint, not the headline. The classical stack completes 59 of 59 single-field missions. A learned driver cannot improve this number. The headline metrics of the framework are the quantities that can improve:

Headline Metric Definition
Cross-track error Distance from the reference path: mean, RMS, 95th percentile and maximum, in metres.
Crop margin Crop area crushed outside the edge band, in m².
Ground (slip) energy \(E_\text{ground}\), kJ for each mission.
Shield interventions Interventions for each 1000 control steps, for each rule.
Time over \(T_\text{ref}\) Mission time divided by \(T_\text{ref}\).
Success for each phase The share of GoTo, Scout and Return phases that end correctly.

The evaluation protocol of the framework reports these metrics with and without sensor noise, on dry and wet soil, for each gain \(\alpha\). With \(\alpha = 0\) it must reproduce the reference driver. The whole-mission evaluator is specified, not built. The scores of the table below exist (eval/scorers.py).

The full score table is:

Score Definition
Mission success Complete missions / all missions
Coverage Mean coverage of the requested fields
Fuel Litres for each mission, and for each km
Ground loss \(E_\text{ground}\), kJ for each mission
Objective and regret \(J\), and \((J - J^*)/J^*\), with \(J^*\) the optimal plan driven by the reference driver
Crop damage Crushed area outside the edge band (m²), and inside it against \(A^\text{edge}_f\)
Time Minutes for each mission, against \(T_\text{ref}\)
Stuck and collisions Counts for each 100 missions
Safety Minimum distance to workers; time within 5 m of a worker at more than 2 m/s
Smoothness RMS change of the commands for each step
Latency Model inference time for each chunk (vision models)
Imagination For FastWAM: agreement of its predicted frames with the rendered frames (reported, not ranked)

Each score is a mean with a 95 % confidence interval. The report gives it for each condition set and each mission size. It gives it separately for the seen fields and the held-out fields.

Tracks. The driving track gives the optimal plan to each driver. The planning track uses the reference driver for the plan of each planner. The end-to-end track combines both.

A real-data check runs open loop. The models get the frames and states of a perimeter loop that a person drove on a field day. The check compares their commands with the commands of the driver.

Methods#

Method Level Inputs Notes
Reference driver Driving Path, proprioception Pure pursuit with a speed profile. It also drives backwards: the same arc through a point behind the vehicle, at most 1.5 m/s, with a stop at each change of direction. It defines \(E_\text{ref}\), \(T_\text{ref}\) and \(J^*\).
Residual driver Driving The residual-policy observation Specified, not built. The primary learned method of the framework. PPO trains a bounded residual on the reference driver for one segment. The shield is always active. With \(\alpha = 0\) it is the reference driver.
Flat PPO Driving All state groups, LiDAR The negative baseline: one network for the whole mission. It completed no whole mission (Results).
ACT, Diffusion Policy Driving Camera, map image, proprioception LeRobot; imitation baselines. Later: segment skills behind the same interface.
SmolVLA Driving Camera, map image, proprioception, instruction Fits a 16 GB GPU. Later, as a segment skill.
π0.5 (openpi, LoRA) Driving The same Needs a larger GPU for fine-tuning. Later, as a segment skill.
FastWAM Driving The same; it also predicts frames LeRobot v2.1/3.0 data. Needs a larger GPU for fine-tuning. Later, as a segment skill.
Optimal, nearest first, as listed Planning Cost matrix Solver baselines
Attention routing policy Planning Field positions, costs Learned planner
VLM planner Planning Map, instruction, cost tools The start of the agentic harness

The later driving methods use the same segment interface, the same mission automaton and the same shield as the residual driver. Framework gives the reasons for this structure.

Version 2: Changing Weather#

Version 2 is the same task with conditions that change during the mission. Rain arrives at a forecast time. Fields dry after rain. The soil water of each field follows the water model of the farm. The costs become time-dependent, \(J(e, t)\). The planner sees the forecast and the current soil water.

It can plan again when a field is complete (a receding-horizon plan). The rewards do not change, because they count the energy that is spent. Version 2 adds two scores. The first is the regret against a planner that knows the weather in advance. The second is the regret against the version-1 plan from the start of the mission, without a new plan.