Fine-Tuning¶
With fine_tune=True and reference_boundaries, the pipeline trains the
engine on the reference polygons over the composite, then runs inference with
the new checkpoint (passed to the engine as engine_params["checkpoint_path"]
and recorded in the provenance record as facts.fine_tuned_checkpoint).
gdf = agribound.delineate(
study_area="area.geojson",
source="naip",
year=2022,
engine="dinov3",
reference_boundaries="reference_fields.gpkg",
fine_tune=True,
fine_tune_epochs=20,
gee_project="my-project",
)
agribound delineate --study-area area.geojson --source naip --year 2022 --engine dinov3 \
--reference reference_fields.gpkg --fine-tune --fine-tune-epochs 20 --gee-project my-project
Which engines¶
| Engine | Fine-tunable | What is trained |
|---|---|---|
delineate-anything |
yes | Ultralytics YOLO11-seg from the selected Delineate-Anything weights (da_model, default large_v2) |
geoai |
yes (required: no published field weights) | Mask R-CNN ResNet50-FPN from the COCO weights |
dinov3 |
yes (required: no published field weights) | DINOv3 + DPT; full fine-tuning by default, LoRA or decoder-only optional |
prithvi |
yes | Prithvi-EO-2.0 + UPerNet (terratorch), full fine-tuning by default, LoRA optional |
ftw |
no | Train with ftw-baselines (ftw model fit -c <config.yaml>) and pass engine_params={"checkpoint_path": ...} |
embedding |
no | No trainable weights; the reference is used for evaluation only |
ensemble |
no | Fine-tune each member in its own run and pass each member its checkpoint |
fine_tune=True with a non-fine-tunable engine raises ValueError with these
instructions; the engine is never replaced by another one. When
fine_tune=True the output is not evaluated against the same reference
(train on one area, evaluate on another).
Training data¶
agribound.engines.finetune._data cuts chips and masks from the composite:
- RGB engines (Delineate-Anything, GeoAI, DINOv3) read the canonical R, G, B bands with the same scene-level 1-99 percentile stretch as at inference (uint8 rasters unchanged). Prithvi reads Blue, Green, Red, narrow NIR, SWIR 1, SWIR 2 as reflectance × 10000.
- Masks: 0 background, 1 field interior, 2 field boundary (reference pixels
within
boundary_erosionpixels of another polygon or background), 255 where the image has no valid data. Instance masks give each reference polygon its own id, so touching fields stay separate. - The reference must be complete inside the chips that contain its polygons: pixels outside every reference polygon are labelled background.
engine_params for the chips:
| Parameter | Default | Meaning |
|---|---|---|
chip_size |
224 (Prithvi), 256 (DINOv3), 512 below 4 m GSD else 256 (Delineate-Anything); GeoAI: sized from the reference fields (1.25 × the 90th percentile of the fields' longer bounding-box sides, in pixels, rounded up to a multiple of 32 and clamped to 256–1024 px) | Chip size in pixels. GeoAI infers on windows of the chip size, and no window sees the whole of a field larger than a chip. The engine rejoins such a field's pieces only where they meet at a window edge (merge_window_seams, see Engines), so agribound logs a WARNING when more than 10 % of the reference fields are larger than the chip. |
boundary_erosion |
2 | Boundary width in pixels. |
min_label_fraction |
0.01 | Minimum fraction of chip pixels inside reference polygons. |
min_valid_fraction |
0.5 | Minimum fraction of valid image pixels. |
value_scale |
- | Required for Prithvi on local rasters. |
Train/validation split¶
assign_splits is used by every trainer and is seeded from config.seed
(agribound._repro.get_rng):
fine_tune_split="block"(default): chips are grouped into square blocks offine_tune_block_size_m(default 5000 m) in an equal-area CRS, and whole blocks go to validation until aboutfine_tune_val_split(default 0.2) of the chips; this reduces the spatial leakage of a random chip split."random": a seeded permutation of the chips."column": groups fromfine_tune_split_column, a column of the reference layer (majority value among the reference polygons in each chip).
At least one training and one validation chip are guaranteed when there are
two or more chips; otherwise a ValueError is raised.
Engine-specific settings¶
Delineate-Anything. Labels are the reference polygons clipped to each
chip. Chips are upsampled by the same super-resolution factor as at inference
(2 at 4 m GSD or coarser), so imgsz is 512 when chip size × factor is 512
(a WARNING is logged and imgsz_matches_model_input: false recorded
otherwise). Settings: seed=config.seed, deterministic=True, mosaic=0,
AdamW with lr0 0.002 (yolo_lr0) and a bias warmup learning rate of 0
(yolo_warmup_bias_lr; what Ultralytics' optimizer="auto", whose AdamW
choice the recipe follows, sets; agribound 1.0.1 used Ultralytics' default of
0.1, which applies only to a named optimizer), flips, batch 16 (yolo_batch),
epochs=fine_tune_epochs; Ultralytics keeps best.pt. The test suite runs
this trainer against a stub of Ultralytics. Example 12's NAIP runs fine-tuned
large_v2 with agribound 1.0.1's recipe and Ultralytics 8.4.163 (650 chips of
512 px, 10 epochs); their in-sample scores, and those of the GeoAI and DINOv3
models fine-tuned on the same reference, are in the gallery.
Small training sets
Ultralytics accumulates gradients over 64 images, so a few dozen chips
give about one optimizer step per epoch after the warmup, and the default
lr0 can move the pretrained weights further than the run can recover.
On the 62 training chips (512 px, SPOT 6/7 panchromatic) of an oil palm
estate in example 23, the validation mask mAP50 at the default lr0 was
0.02 after the first epoch (the checkpoint Ultralytics kept) and 0.00 from
the third epoch on; with engine_params={"yolo_lr0": 1e-4} it was 0.47
after 20 epochs, against 0.17 for the released weights (agribound 1.0.1's
bias warmup scored 0.00 at the default lr0). With a small training set, set a smaller
yolo_lr0 and compare the validation scores (in the Ultralytics
results.csv of the run) with those of the released weights. A better
validation score does not guarantee a better result elsewhere: in example
23 the model fine-tuned on that estate did worse than the released weights
on another estate, 69 km from the training blocks (centres 81 km apart).
GeoAI. geoai's Mask R-CNN recipe run on agribound's own train/validation
chips (geoai's own trainer would re-split the chips randomly): SGD (lr
learning_rate 0.005, momentum 0.9, weight decay 5e-4), StepLR, batch 4,
the checkpoint with the best validation mask IoU (best_model.pth), early
stopping after early_stopping_patience (10) epochs. The chip size is saved
next to the checkpoint and becomes the default inference window. On Apple MPS
the model trains on CPU (WARNING), as at inference.
A WARNING is logged when the best validation IoU is below 0.1 (the model has
probably not learned the field boundaries, and its output may be artefacts)
or when there are fewer than 10 training chips (treat the model as a smoke
test). The warnings are also written to the warnings list of
best_model.pth.agribound.json and to the run's provenance record. They are
a floor for "learned something", not a quality target: a Namoi test run with
one training and one validation chip peaked at a validation IoU of 0.033 and
delineated a regular lattice of ovals.
DINOv3. geoai.dinov3_finetune.train_dinov3_segmentation: cross-entropy
with ignore_index=255, AdamW (learning_rate 1e-4, weight_decay 1e-4)
with cosine annealing, batch 4, early stopping on val_loss
(early_stopping_patience 10), best checkpoint by val_loss. Full
fine-tuning by default (use_lora=False, freeze_backbone=False; about 303 M
backbone parameters for ViT-L/16); use_lora=True (rank lora_rank, default
4; about 0.39 M adapter parameters) or freeze_backbone=True (decoder only).
trainer_kwargs is passed to the Lightning trainer. The trainable-parameter
counts of each run are written to <checkpoint>.agribound.json.
The returned checkpoint is the best one by val_loss. geoai-py also saves a
last.ckpt, which agribound deletes once the best checkpoint is known.
Checkpoints hold the optimizer state: a ViT-L/16 full fine-tuning checkpoint
is about 3.7 GB (the ~303 M float32 backbone weights plus AdamW's two moment
buffers). While training runs, both checkpoints exist, so the cache directory
needs about twice that in free space.
Prithvi. terratorch SemanticSegmentationTask with a Prithvi-EO-2.0
backbone (model_name, default Prithvi-EO-2.0-300M-TL), a UPerNet decoder
and three classes; AdamW (learning_rate 1e-4), batch 8, best checkpoint by
val/loss. use_lora=True adds LoRA on the query and value projections
(lora_rank default 16); use_lora together with freeze_backbone raises
ValueError (it would freeze the adapters too). early_stopping_patience
(no early stopping by default) adds a Lightning EarlyStopping callback on
val/loss, and trainer_kwargs is passed to the Lightning trainer. Needs the GFM environment. On
Apple MPS the default 224 px chips run on CPU; chip_size=192 (and
tile_size=192 at inference) runs on MPS.
Caching¶
Each fine-tuning run gets its own cache directory keyed by the study area,
source, year/date range, compositing settings, engine, base model, a
fingerprint of the reference file (path, modification time, size), epochs,
split settings, seed, bands, the training-related engine_params (all but
checkpoint_path and the sam_* keys) and the engine's default chip-size
rule, so a changed default does not reuse a checkpoint trained on chips of
another size; for Delineate-Anything it also includes the version of the
training recipe (agribound.engines.finetune._yolo.RECIPE_VERSION), so a run
cached by an earlier recipe is trained again. A second call with the same
inputs returns the cached checkpoint without retraining
(finetune_manifest.json).
Large areas¶
agribound tiles make refuses fine_tune: true unless
--allow-fine-tune-per-tile is given (it would train one model per tile). Fine-tune once over the reference area and pass the checkpoint to the
tiles with --engine-param checkpoint_path=<path>; see
HPC.