Configuration Reference¶
Every run is described by one AgriboundConfig dataclass. It can be built in
Python, loaded from YAML (AgriboundConfig.from_yaml), or assembled from CLI
flags; every path runs the same validation. Unknown keys are rejected: a YAML
file or from_dict raises a ValueError that lists the valid ones, and the
constructor raises Python's TypeError for an unexpected keyword argument.
config.to_yaml(path) and
AgriboundConfig.from_dict(config.to_dict()) round-trip all fields, and
config.merged(**overrides) returns a validated copy with some fields changed.
from agribound import AgriboundConfig, delineate
config = AgriboundConfig(
study_area="area.geojson",
source="sentinel2",
year=2024,
engine="delineate-anything",
gee_project="my-gee-project",
output_path="fields.gpkg",
)
gdf = delineate(config=config)
delineate(config=config, **kwargs) applies keyword arguments on top of the
configuration (logged); the caller's object is not modified.
Validation¶
At construction the configuration checks, among other things:
- enum fields (
source,engine,output_format,export_method,composite_method,device,s2_cloud_mask,tessera_version, alllulc_*enums,sam_backend,fine_tune_split,google_embedding_backend,aoi_selection); - that the engine supports the source (for
ensemble, every member, including the default members); - that
yearlies in the source's available years (for TESSERA, those oftessera_version), and, forlandsat-pan, that a mission selected bylandsat_pan_missionshas a record overlapping the year ordate_range(for exampleLE07with 2025 orLC09with 2020 is rejected); - that
fine_tune=Truehasreference_boundariesand a fine-tunable engine; export_crs("utm"or a validEPSG:<code>),date_rangeformat and order, numeric ranges, thebandsmapping (1-based indices), thegee_workload_tagformat, andlandsat_pan_missions("auto"or mission IDs fromLE07,LC08andLC09, stored without duplicates in that order);- that a GEE imagery source has a project (from
gee_project, theGEE_PROJECTenvironment variable,gcloud config, or else theproject_idof the service-account key orGOOGLE_APPLICATION_CREDENTIALSfile; see GEE setup).
The output format is inferred from the output_path extension (.gpkg,
.geojson/.json, .parquet/.geoparquet) when output_format is left at
its default; a conflicting explicit format raises.
Fields¶
Defaults are shown in parentheses.
Core¶
| Field | Description |
|---|---|
source ("sentinel2") |
One of the sources. |
engine ("delineate-anything") |
One of the engines. |
year (2024) |
Target year. |
study_area ("") |
Vector file (GeoJSON, GeoPackage, Shapefile, GeoParquet), GEE vector asset ID (projects/... or users/...), "bbox:minx,miny,maxx,maxy" (EPSG:4326) or a WKT geometry (EPSG:4326). Optional only for source="local". |
output_path ("fields.gpkg") |
Output vector file. delineate() without output_path writes fields_<source>_<year>.<ext> in the current directory. |
output_format ("gpkg") |
"gpkg", "geojson" or "parquet" (GeoParquet). |
Earth Engine¶
| Field | Description |
|---|---|
gee_project (None) |
Earth Engine project. Required for GEE imagery sources; when not given, resolved from GEE_PROJECT, then gcloud config, then the project_id of the credentials file (gee_service_account_key, AGRIBOUND_GEE_SERVICE_ACCOUNT_KEY or GOOGLE_APPLICATION_CREDENTIALS; see GEE setup). |
gee_service_account_key (None) |
Service-account JSON key (also read from AGRIBOUND_GEE_SERVICE_ACCOUNT_KEY). See GEE setup. |
gee_high_volume (False) |
Use the high-volume Earth Engine endpoint. |
gee_max_requests (8) |
Maximum concurrent download requests of this process. A WARNING is logged above 40, Earth Engine's default per-project limit. |
gee_workload_tag (None) |
Workload tag for EECU accounting (1-63 characters). |
export_method ("local") |
"local" (direct download), "gdrive" or "gcs" (batch export; see Batch exports). |
gcs_bucket (None) |
Required with export_method="gcs". |
Compositing¶
| Field | Description |
|---|---|
composite_method ("median") |
"median", "greenest" or "max_ndvi" (an alias of greenest). |
date_range (None) |
("YYYY-MM-DD", "YYYY-MM-DD") window (end inclusive) instead of the calendar year. |
cloud_cover_max (20) |
Maximum scene cloud cover in percent (scene filter). |
export_crs ("utm") |
"utm" (UTM zone of the study-area centroid) or "EPSG:<code>". |
s2_cloud_mask ("scl") |
Sentinel-2 pixel mask: "scl" or "cloud_score_plus". |
cloud_score_threshold (0.60) |
Minimum Cloud Score+ cs_cdf kept. |
naip_resolution_m (1.0) |
NAIP export resolution in metres. |
landsat_pan_missions ("auto") |
Landsat missions whose panchromatic band source="landsat-pan" uses (ignored by the other sources). "auto" never mixes the two PAN bandpasses: Landsat 8/9 (LC08, LC09) when the date window overlaps their record (from 2013-03-18), else Landsat 7 (LE07). Or a list of "LE07", "LC08" and "LC09" (a comma-separated string is also accepted): exactly those missions, where their record overlaps the window; Landsat 7 together with Landsat 8 or 9 mixes the two bandpasses in one median, with a WARNING. See Landsat panchromatic. |
tile_size (10000) |
Maximum download tile size in pixels per side (tiles are assembled into one GeoTIFF). |
USGS NAIP Plus¶
| Field | Description |
|---|---|
usgs_service_url |
ImageServer URL (default: the USGS NAIP Plus service). |
usgs_state (None) |
Two-letter state code filter. |
usgs_allow_year_fallback (False) |
Also accept footprints from year ± 1. |
usgs_timeout_s (120), usgs_retries (3) |
Request timeout and retries. |
Embeddings¶
| Field | Description |
|---|---|
tessera_version ("v1") |
"v1", "v1.1" or "v2" (beta). |
tessera_variant (None) |
Dataset variant (geotessera's default for the version when None). |
embedding_cache_dir (None) |
Cache directory for geotessera's Zarr reads and the Source Cooperative tile index (default ~/.cache/agribound for the index; TESSERA reads are not cached on disk). |
google_embedding_backend ("gee") |
"gee" or "source_coop". |
Local input¶
| Field | Description |
|---|---|
local_tif_path (None) |
GeoTIFF for source="local". |
bands (None) |
Canonical band name → 1-based index, e.g. {"R": 1, "G": 2, "B": 3, "NIR": 4}. Honoured for every source. |
Study-area selection¶
| Field | Description |
|---|---|
aoi_selection ("representative_point") |
How predictions are restricted to the study-area geometry (composites cover its bounding box). |
Applied after delineation and SAM refinement, before post-processing, the LULC filter and evaluation:
"representative_point": keep polygons whose representative point (shapely.point_on_surface, always inside the polygon) lies in the study area or on its boundary; fields crossing the outline are kept or dropped whole."intersects": keep polygons that intersect the study area."clip": clip polygons to the study area (slivers are then removed by the area filter)."none": keep every polygon in the composite's bounding box.
The rule and the counts before and after are recorded in the provenance record
(facts.aoi_selection). Reference polygons used for evaluation are selected
with the same rule (with "none": those that intersect the study area).
Post-processing¶
| Field | Description |
|---|---|
min_field_area_m2 (2500.0) |
Minimum polygon area (EPSG:6933); holes smaller than this are filled. Applied before smoothing and again after it (see below), so no output polygon is smaller. |
simplify_tolerance (2.0) |
Douglas-Peucker tolerance in metres (0 disables). |
The pipeline's post-processing is: merge overlapping polygons → area filter
and hole removal → Chaikin smoothing (engine_params["smooth_iterations"],
default 3) → simplification → optional regularisation
(engine_params["regularize"], default "none") → area filter again.
The second area filter runs whenever smoothing, simplification or
regularisation is on (with the defaults it always runs). It removes the
polygons these steps shrank below min_field_area_m2; it does not fill holes
again. The provenance record lists it under facts.postprocess
(min_field_area_applied). Outputs can therefore have fewer polygons than
with agribound 0.1.x, which filtered only before smoothing. Output reuse
does not cover this step (it is not versioned in agribound._results): an
output written by a build without the second filter that has a matching
provenance record is reused as it is. Pass overwrite=True to recompute it.
Smoothing and simplification bias areas low
Measured on four real engine outputs (2026-09-27), the default smoothing
and simplification changed polygon areas by a median of -0.6 % to -4.0 %
(10th percentile -4.7 % to -11.4 %), with median Hausdorff distances of
7-10 m. A rectangle traced with only its four corners loses about 16 %
of its area to the smoothing. Set engine_params["smooth_iterations"]=0
and simplify_tolerance=0 to keep the engine outlines, for example for
area statistics. The full table is in the docstring of
agribound.pipeline._postprocess.
LULC crop filter¶
| Field | Description |
|---|---|
lulc_filter (True) |
Remove polygons whose crop fraction is below the threshold. Needs Earth Engine for every source. |
lulc_crop_threshold (0.3) |
Minimum crop fraction (0-1). |
lulc_dataset ("auto") |
"auto", "nlcd", "cdl", "dynamic_world" or "c3s". |
lulc_tree_crops (False) |
Count tree cover as crop, for tree crops such as orchards and plantations: Dynamic World then uses the crops + trees probability and C3S adds its tree-cover classes, so with these two datasets the filter no longer removes forest. NLCD and CDL are unchanged (NLCD class 82 already includes orchards and vineyards) and still remove forest. See Tree crops. |
lulc_mode ("server") |
"server" (Earth Engine reduceRegions) or "raster" (download a LULC raster in stage A, filter locally). |
lulc_on_error ("raise") |
"raise" aborts on failure; "warn" keeps the unfiltered polygons and records the failure. |
lulc_nodata_policy ("keep") |
Polygons without valid LULC pixels are kept and flagged ("keep") or dropped ("drop"). |
lulc_batch_size (200) |
Polygons per reduceRegions request. |
See LULC crop filter for the datasets and the routing rule.
SAM refinement¶
| Field | Description |
|---|---|
sam_refine (False) |
Refine polygons with box-prompted SAM. |
sam_backend ("sam2") |
"sam2", "sam2.1", "sam3" or "sam3-hf". The two SAM 3 backends are untested in 1.0.1 (not run end to end, because the facebook/sam3 weights are gated); a WARNING is logged when one is loaded. |
sam_model (None) |
Model id (backend default when None). |
sam_min_crop_px (64) |
Minimum padded bounding-box side in pixels. |
sam_crop_padding (0.15) |
Padding per side as a fraction of the box size. |
See SAM refinement and SAM 3 is untested.
Compute¶
| Field | Description |
|---|---|
device ("auto") |
"auto" (CUDA, then MPS, then CPU), "cuda", "mps" or "cpu". |
n_workers (4) |
PyTorch data-loader worker processes (0 = main process). |
seed (42) |
Seed for Python, NumPy, torch and Lightning (see Reproducibility). |
Fine-tuning and evaluation¶
| Field | Description |
|---|---|
reference_boundaries (None) |
Reference polygons for fine-tuning or evaluation. |
fine_tune (False) |
Fine-tune the engine on the reference before inference. |
fine_tune_epochs (20) |
Training epochs. |
fine_tune_val_split (0.2) |
Fraction of chips used for validation. |
fine_tune_split ("block") |
"block" (spatial blocks), "random" or "column". |
fine_tune_block_size_m (5000.0) |
Block edge length for the block split. |
fine_tune_split_column (None) |
Reference column used as group id for the column split. |
When reference_boundaries is set and fine_tune is False, the output is
evaluated against it (see Evaluation).
The training-chip settings are engine parameters, not configuration fields.
For example, engine_params["chip_size"] defaults, for GeoAI, to a size
derived from the reference fields (1.25 × the 90th-percentile field
bounding-box side, rounded up to a multiple of 32 px and clamped to
256-1024 px); see
Fine-tuning.
Caching and provenance¶
| Field | Description |
|---|---|
cache_dir (None) |
Directory for intermediates (default <output dir>/.agribound_cache). |
overwrite (False) |
Re-run and replace an existing output. |
provenance (True) |
Write <output_path>.provenance.json. |
Engine parameters¶
| Field | Description |
|---|---|
engine_params ({}) |
Engine-specific options, documented per engine in Engines and the API reference. Some engines reject options they cannot honour (for example Delineate-Anything options that the selected backend does not support, or unknown ensemble-level keys). |
YAML¶
study_area: my_region.geojson
source: sentinel2
year: 2024
engine: delineate-anything
gee_project: my-gee-project
composite_method: median
cloud_cover_max: 20
export_crs: utm
min_field_area_m2: 2500
simplify_tolerance: 2.0
lulc_filter: true
lulc_crop_threshold: 0.3
engine_params:
da_model: large_v2
conf_threshold: 0.15
output_path: fields.gpkg
seed: 42
agribound delineate --config config.yml
agribound delineate --config config.yml --year 2023 --engine-param conf_threshold=0.2
agribound delineate --config config.yml --dry-run # print the resolved YAML and exit
Options given explicitly on the command line override the YAML; everything
else, including study_area, comes from the file.