Evaluation¶
agribound.evaluate.evaluate(predicted, reference, ...) compares predicted
field polygons with reference field polygons at the object level and returns
a dictionary of metrics. evaluate_frame returns the per-reference-field table
the summary is computed from, and pixels_per_field the number of pixels a
field spans at a given ground sample distance. The same function backs the
agribound evaluate command and the pipeline's automatic evaluation.
from agribound.evaluate import evaluate
from agribound.io import read_vector
metrics = evaluate(
read_vector("fields.gpkg"),
read_vector("reference.gpkg"),
iou_threshold=0.5,
strata="county", # a reference column
size_bins=[0, 0.5, 1, 2, 5, 10, float("inf")], # hectares
bootstrap=1000,
boundary_tolerance_m=10.0,
)
print(metrics["f1"], metrics["area_weighted_recall"], metrics["bootstrap"]["ci"]["recall"])
agribound evaluate -p fields.gpkg -r reference.gpkg --strata-column county \
--size-bins 0,0.5,1,2,5,10,inf --bootstrap 1000 --boundary-tolerance-m 10 -o metrics.json
Example 20
(examples/20_stratified_evaluation.py)
uses these functions against the NMOSE registry in the San Juan Basin: strata,
size classes, bootstrap intervals, the 10 m boundary tolerance, the per-field
table and pixels per field. It also compares Delineate-Anything v2 on Landsat,
Sentinel-2, SPOT 6/7 and NAIP of one year (--resolution-year, default 2018;
--no-resolution skips it) and the pre-trained FTW and Delineate-Anything
engines with and without the LULC crop filter (--no-lulc-comparison skips
it); --predicted PATH evaluates an existing layer and skips both
comparisons.
In the pipeline¶
When reference_boundaries is set and fine_tune is False, delineate()
calls evaluate(output, reference) with the defaults (IoU 0.5, one-to-one
matching, no tolerance metrics) after the metadata step. With a study area,
the reference polygons are first selected with the same aoi_selection rule
as the predictions (with "none": the references that intersect the study
area), so a field crossing the outline is kept or dropped on both sides alike.
The metrics are stored in gdf.attrs["evaluation_metrics"] and in the
provenance record (facts.evaluation, with the reference counts in
facts.evaluation_reference).
Incomplete reference layers
Predictions that overlap no reference field count as false positives. If
the reference covers only some fields in the area (a registry of some
fields), precision is not meaningful; use recall, or the precision among
predictions that overlap a reference field,
count_tp / (count_tp + count_fp - count_fp_unassigned).
Matching¶
A prediction p and a reference field r can match when
IoU(p, r) = |p ∩ r| / |p ∪ r| >= iou_threshold (compared with a relative
tolerance of 1e-9) and the intersection has positive area.
matching="one_to_one"(default since 1.0): all candidate pairs at or above the threshold are sorted by descending IoU (ties by reference, then prediction position), and a pair is accepted when neither polygon has been matched yet. This is a greedy assignment in the spirit of COCO-style detection evaluation (processed in order of IoU, since there are no confidence scores), not an optimal assignment. With non-overlapping polygons andiou_threshold > 0.5the matching is unique.matching="many_to_one": the 0.1.x rule. Each reference field is matched to its highest-IoU prediction at or above the threshold, and one prediction may match several reference fields, so precision can exceed the fraction of predictions that are matched.
TP = matched reference fields, FP = unmatched predictions, FN = unmatched reference fields.
Summary metrics¶
Kept from 0.1.x:
| Key | Definition |
|---|---|
precision |
TP / (TP + FP) |
recall |
TP / (TP + FN) |
f1 |
2 · precision · recall / (precision + recall) |
iou_mean |
mean IoU of matched pairs (0.0 when nothing matched) |
over_segmentation |
fraction of reference fields with more than one prediction at IoU ≥ threshold (0.1.x definition) |
under_segmentation |
fraction of predictions with more than one reference field at IoU ≥ threshold (0.1.x definition) |
area_error_mean_m2 |
mean absolute area difference of matched pairs |
count_predicted, count_reference, count_tp, count_fp, count_fn, iou_threshold |
counts and the threshold |
With non-overlapping polygons and a threshold above 0.5, the two 0.1.x
segmentation keys are always 0; use oversegmentation_mean and
undersegmentation_mean below to measure splits and merges.
Added in 1.0:
| Key | Definition |
|---|---|
best_iou_mean |
mean over all reference fields of the highest IoU with any prediction (0 when none overlaps); not conditional on a match |
area_weighted_recall |
area of matched reference fields / area of all reference fields |
area_weighted_precision |
area of matched predictions / area of all predictions |
area_weighted_f1 |
harmonic mean of the two |
reference_area_ha, predicted_area_ha, median_area_ha |
total areas and the median reference field area |
area_error_signed_mean_m2 |
mean of (predicted − reference) area over matched pairs |
oversegmentation_mean, undersegmentation_mean |
means over reference fields of OS = 1 − |r ∩ p′| / |r| and US = 1 − |r ∩ p′| / |p′|, where p′ is the prediction with the largest intersection with r (the maximum-overlap pairing of Persello & Bruzzone, 2010; the form of Clinton et al.'s, 2010, measures). A reference field that no prediction overlaps has OS = 1 and no US. |
hausdorff_mean_m, hausdorff_median_m |
Hausdorff distance between the boundaries of matched pairs |
boundary_distance_mean_m, boundary_distance_median_m |
symmetric mean boundary distance of matched pairs, (∫∂r d(x, ∂p) dx + ∫∂p d(x, ∂r) dx) / (|∂r| + |∂p|) |
boundary_sample_spacing_m, boundary_sample_spacing_max_m |
the sampling argument, and the largest spacing Δ used for any matched pair |
count_fp_unassigned |
false positives that overlap no reference field |
count_*_dropped, count_*_repaired |
rows excluded (null, empty, non-polygonal, zero area) and invalid geometries repaired |
matching, equal_area_crs, distance_crs, utm_epsg_codes |
how the metrics were computed |
With boundary_tolerance_m:
| Key | Definition |
|---|---|
boundary_recall |
length of reference boundaries within the tolerance of any predicted boundary / total reference boundary length |
boundary_precision |
the same for predicted boundaries against reference boundaries |
boundary_f1 |
harmonic mean of the two |
coverage_within_tolerance |
fraction of reference fields that are matched and have a symmetric mean boundary distance ≤ tolerance |
coverage_within_tolerance_area |
the same, weighted by reference area |
Boundary lengths are summed per polygon (an edge shared by two fields counts once for each). The within-tolerance lengths are computed exactly (the length of one layer's boundary inside the union of circular-arc buffers of the other layer's boundaries), without polygonal buffer approximation.
Empty denominators: rates (keys ending in precision, recall, f1,
iou_mean, the coverage keys) and the 0.1.x keys over_segmentation,
under_segmentation, area_error_mean_m2 are 0.0; other means and distances
are NaN. When both layers are empty the rates are 1.0 (0.1.x behaviour).
Strata, size classes and confidence intervals¶
strata: a reference column name, an aligned Series or an array. Metrics are also reported per stratum (per_stratum); rows with a missing value count only in the overall metrics.size_bins:"auto"(1-2-5 series spanning the reference areas) or strictly increasing edges in hectares of the reference field area; classes are[lower, upper)except the last,[lower, upper]. Metrics are also reported per class (per_size_class).- In a group, a prediction counts with its home reference field: the matched one, else the one it overlaps most. Predictions overlapping no reference field count only in the overall metrics.
bootstrap=N: percentile bootstrap intervals at 95 % (BOOTSTRAP_CONFIDENCE_LEVEL), resampling reference fields with replacement (each stratum separately whenstratais given), from child seeds ofnumpy.random.SeedSequence(bootstrap_seed)(default 42). The intervals treat the reference fields as an independent (stratified) random sample: they do not model spatial autocorrelation, so they are too narrow when errors are spatially clustered, and they are not design-based estimates for a probability sample (Stehman & Foody, 2019).
Geometry and coordinate systems¶
- Invalid geometries are repaired with
shapely.make_valid(method="structure", keep_collapsed=False)and only their polygonal parts are kept. This needs shapely >= 2.1 (GEOS >= 3.10); with an older shapely anImportErroris raised when a repair is needed (there is no fallback to another repair method). - Areas and IoU are computed in
equal_area_crs(default EPSG:6933). Edges are split into pieces of at most 50 m before reprojection, so a long straight edge is not replaced by one chord. A WARNING is logged when a given CRS changes any polygon's area by more than 0.01 % relative to EPSG:6933; for example ESRI:54009 (Mollweide, a spherical projection) is not equal-area on WGS 84 coordinates. - Distances (Hausdorff, boundary distances, tolerance metrics) are computed in the WGS 84 / UTM zone of each reference field (predicted boundaries in their own zone for boundary precision).
- Boundaries are sampled at points at most Δ apart
(
boundary_sample_spacing_m, default 1 m;Noneselectsmax(1 m, perimeter / 1000), which is faster on large fields). The Hausdorff distance is then underestimated by at most Δ/2 and the mean distance is within Δ/4 of its exact value. - Layers that cross the antimeridian are not supported.
Public constants: MATCHING_METHODS, BOUNDARY_SAMPLE_SPACING_M (1.0, the
default spacing), BOUNDARY_MAX_SAMPLES (1000, used only with spacing None)
and BOOTSTRAP_CONFIDENCE_LEVEL (0.95).
Run time: on 50,603 NMOSE reference fields against 35,225 predictions,
evaluate() took about 56 s with the defaults and 33 s with
boundary_sample_spacing_m=None (measured 2026-09-27 on the development
machine).
Per-field table¶
evaluate_frame(...) returns one row per evaluated reference field, indexed
like the reference layer: area_m2, area_ha, perimeter_m, utm_epsg,
matched, pred_index, iou, pred_area_m2, best_iou,
n_overlapping_pred, max_overlap_pred_index, oversegmentation,
undersegmentation, hausdorff_m, boundary_mean_distance_m,
boundary_sample_spacing_m, plus stratum/size_class and the tolerance
columns when requested.
Pixels per field¶
pixels_per_field(reference, gsd_m) returns A / gsd_m² per field, the
number of gsd_m × gsd_m pixels whose total area equals the field area. It is
a quick measure of how well a sensor can resolve the fields (the agent's
estimate_resolvability tool uses it).
References¶
- Clinton, N., Holt, A., Scarborough, J., Yan, L., & Gong, P. (2010). Accuracy assessment measures for object-based image segmentation goodness. Photogrammetric Engineering & Remote Sensing 76(3), 289-299. https://doi.org/10.14358/PERS.76.3.289
- Persello, C., & Bruzzone, L. (2010). A novel protocol for accuracy assessment in classification of very high resolution images. IEEE Transactions on Geoscience and Remote Sensing 48(3), 1232-1244. https://doi.org/10.1109/TGRS.2009.2029570
- Stehman, S. V., & Foody, G. M. (2019). Key issues in rigorous accuracy assessment of land cover products. Remote Sensing of Environment 231,
- https://doi.org/10.1016/j.rse.2019.05.018