Tasks
A task is one definition of a job for a vehicle. Each simulator tier uses the same definition. Each learner and each
classical controller also uses it. Thus the scores of different methods are comparable.
Parts of a Task
| Part |
Content |
| Map products |
The geometry of the task. A script builds the products one time for each map version. |
| Mission |
The job, its completion test and its failure tests. |
| Plan |
The decisions that a planner makes before the vehicle moves. |
| Phases |
The sequence of steps in a mission. The task controls the sequence. A learned policy does not. |
| Actions |
The commands that a driver sends to the vehicle in each control step. |
| Observations |
The inputs that a driver gets in each control step. |
| Rewards |
The reward terms, their weights and the reasons for the weights. |
| Terminations |
The events that end an episode. |
| Metrics |
The scores of the evaluation. |
| Baselines |
The reference driver, the planners and their measured scores. |
Location of the Code
- The code of a task is in
Learning/acres_learn/tasks/<task>/.
- The constants of the field scouting task are in
Learning/acres_learn/configs/scouting_v1.json.
- ACRES Core and the game read the same constants and the same map products.
- The learners, the training and the deployment are on the Learning pages.
List of Tasks
Field Scouting Pages
- Field Scouting Task is the specification. It gives the mission, the plan, the phases, the actions, the
observations, the rewards, the terminations and the metrics. It has two formulations of the driving problem.
- Scouting Map Products describes the geometry: ground classes, loops, checkpoints, entries and the
transit graph.
- Scouting Baselines describes the reference driver, the planners and their measured scores.
The mission, the map products, the plan and the metrics are the same in the two formulations.
| Item |
Baseline Formulation |
Framework Formulation |
| Driver |
One flat PPO policy outputs the full command for the whole mission |
The reference driver gives the base command. A learned policy adds a bounded residual |
| Order of the fields |
The exact planner |
The exact planner. Lean proves that it is exact |
| Phase sequence |
The episode tracker |
The mission automaton. Lean proves its properties |
| Safety layer |
None. The reward has penalties |
A shield replaces each command that breaks a rule. Lean proves its rules |
| Reward |
v1 or bounded |
residual: the bounded reward, a shield term of 0.5 after the bound, a shaping term, failures at −1200 and a collision at −1500 |
| Episode |
A stage of a curriculum |
One segment of a leg |
| State |
Exists. It is the negative baseline: 0 of 64 whole missions |
Specified, not built. The Lean models and proofs exist |