Skip to content

Published FTW Polygons

Data-access helpers for the published Fields of The World polygons. See FTW polygon query.

ftw_query

Query the published Fields of The World (FTW) Global prediction polygons.

Data-access helpers for already-published FTW polygons (PRUE model predictions on Source Cooperative, CC-BY-4.0; see :mod:agribound.ftw_arrow for the layouts and the confidence column). The module does not run FTW inference, host FTW data, or treat FTW predictions as ground truth.

Two backends:

  • "pyarrow": the public GeoParquet on Source Cooperative (default layout by-admin-conf), or any GeoParquet file, directory or glob given as source_url.
  • "manifest": local or HTTP(S) GeoParquet tiles listed in a manifest, or a local tile directory.

query_ftw

query_ftw(study_area: Any, year: int | str | None = None, label: str | None = 'field', clip: bool = True, output_path: str | Path | None = None, output_format: str | None = None, source_url: str | None = None, manifest_path: str | Path | None = None, tile_dir: str | Path | None = None, cache_dir: str | Path | None = None, source_backend: str = 'auto', max_features: int | None = None, columns: list[str] | tuple[str, ...] | None = None, deduplicate: bool = True, dst_crs: str | int | None = None, min_confidence: float | None = None, keep_null_confidence: bool = True, layout: str = DEFAULT_FTW_LAYOUT, provenance: bool = True) -> gpd.GeoDataFrame

Query published FTW polygons for a study area.

Parameters:

Name Type Description Default
study_area object

AOI as a GeoJSON/GPKG/Shapefile/GeoParquet path, "bbox:minx,miny, maxx,maxy" string, bbox tuple/list (minx, miny, maxx, maxy) in EPSG:4326, WKT string, Shapely geometry, GeoSeries, or GeoDataFrame-like object.

required
year int, str, or None

Optional prediction-year filter, applied to determination:datetime (by-admin-conf layout), else to a year column, else to a time column. The by-admin-conf layout publishes 2024 and 2025 only; other years log a WARNING (and return no rows) when that layout is queried, by default or through a source_url under its prefix (including the https://data.source.coop/ftw/global-data/ alias).

None
label str or None

Optional label filter. Defaults to "field". If the tile lacks a label column (the by-admin-conf layout holds fields only), no label filtering is applied.

'field'
clip bool

If True, clip returned polygons to the AOI. The published metrics:area and metrics:perimeter describe the whole polygon, so for every polygon that crosses the AOI boundary they are recomputed from the clipped geometry (area in EPSG:6933, m²; geodesic perimeter on WGS 84, m) and the added boolean column agribound:clipped is True. If False, return the full published polygons that intersect the AOI, with their published metrics.

True
output_path str, Path, or None

Optional destination path. Supported formats are those accepted by :func:agribound.io.vector.write_vector, including GeoParquet, GeoJSON, and GeoPackage.

None
output_format str or None

Optional output format override.

None
source_url str or None

PyArrow backend: GeoParquet file, directory, S3 prefix or glob (default: the prefix of layout; https://data.source.coop/ftw/global-data/ URLs are mapped to the raw S3 prefix). Manifest backend: base URL used to resolve relative tile paths in a manifest, or a URL/path to a manifest when manifest_path is not provided. HTTP/HTTPS candidate tiles are downloaded to cache_dir before reading.

None
manifest_path str, Path, or None

Path or HTTP/HTTPS URL to a local tile manifest. The manifest may be a vector file with tile geometries or a tabular file with tile paths and bbox columns.

None
tile_dir str, Path, or None

Directory containing local GeoParquet tiles. Used to resolve relative paths in a manifest. If no manifest is provided, a manifest is built from tile-level GeoParquet metadata in this directory.

None
cache_dir str, Path, or None

Directory for downloaded remote manifests or candidate tiles, and for the cached partition-bounding-box index of the PyArrow backend (default ~/.cache/agribound/ftw).

None
source_backend str

Source backend: "auto", "pyarrow", or "manifest". In auto mode, local manifests/tile directories use the manifest backend; otherwise the public FTW GeoParquet source is queried with PyArrow.

'auto'
max_features int or None

Optional row limit for PyArrow-backed preview or smoke-test queries.

None
columns list[str], tuple[str, ...], or None

Optional tile columns to read and return. Internal filter and deduplication columns are read when needed but dropped from the result unless explicitly requested.

None
deduplicate bool

Drop repeated polygons, such as the same polygon read from two overlapping tiles. Rows are duplicates when their normalized geometry is identical (SHA-1 of the normalized WKB) and, whenever the prediction year can be read, so is the year, so the same polygon predicted in 2024 and 2025 is kept once per year. The first row is kept; rows with a missing or empty geometry are always kept. The published id is not used as the key because it is not unique per polygon in the by-admin-conf layout: in US_NM (read 2026-09-27) 190,940 ids in 2024 and 495,230 in 2025 are each shared by 2-4 different polygons, typically hundreds of kilometres apart, while no geometry repeats within a year.

True
dst_crs str, int, or None

Optional CRS for the returned GeoDataFrame.

None
min_confidence float or None

Keep polygons with confidence >= min_confidence (0-100 scale). The dataset README (predictions/vectors/README.md on Source Cooperative, read 2026-09-26) recommends 69 (:data:agribound.ftw_arrow.RECOMMENDED_MIN_CONFIDENCE, raw 0.4). Values in (0, 1] log a WARNING (they look like raw 0-1 values). Raises :class:ValueError if the source has no confidence column.

None
keep_null_confidence bool

Keep polygons whose confidence is null (default True). A null confidence means the 500 m confidence raster has no data at the field's point-on-surface, not a low score. False drops them, with or without min_confidence. Coverage is uneven: on 2026-09-27 the confidence was null for 99.7 % of the rows of AU_NSW and all rows of US_NM (1.8 % for France), so there min_confidence filters almost nothing unless null values are dropped. A WARNING is logged when min_confidence is set and null-confidence polygons are kept.

True
layout str

Published layout used when source_url is None: "by-admin-conf" (default; alpha/results-by-admin-conf, partitioned by country and subdivision; only the partitions whose bounding box intersects the AOI are read) or "raw" (legacy alpha/results with label/time columns).

DEFAULT_FTW_LAYOUT
provenance bool

With output_path, also write <output_path>.provenance.json (default True): the query parameters, the resolved backend and source, the counts of duplicates dropped, clipped and returned polygons, and the package versions (see :func:query_provenance_record).

True

Returns:

Type Description
GeoDataFrame

Published FTW polygons intersecting the AOI. The bbox struct column is dropped unless requested in columns. For the PyArrow backend attrs["ftw_query"] records the source, the number of files listed and opened, the filters, the number of duplicates dropped, the number of returned polygons with a null confidence (n_null_confidence, when the source has that column) and n_returned.

Notes

This function is a query/download helper for existing FTW prediction polygons. It does not run FTW inference and should not be used to treat FTW predictions as ground truth.

Source code in agribound/ftw_query.py
def query_ftw(
    study_area: Any,
    year: int | str | None = None,
    label: str | None = "field",
    clip: bool = True,
    output_path: str | Path | None = None,
    output_format: str | None = None,
    source_url: str | None = None,
    manifest_path: str | Path | None = None,
    tile_dir: str | Path | None = None,
    cache_dir: str | Path | None = None,
    source_backend: str = "auto",
    max_features: int | None = None,
    columns: list[str] | tuple[str, ...] | None = None,
    deduplicate: bool = True,
    dst_crs: str | int | None = None,
    min_confidence: float | None = None,
    keep_null_confidence: bool = True,
    layout: str = DEFAULT_FTW_LAYOUT,
    provenance: bool = True,
) -> gpd.GeoDataFrame:
    """Query published FTW polygons for a study area.

    Parameters
    ----------
    study_area : object
        AOI as a GeoJSON/GPKG/Shapefile/GeoParquet path, ``"bbox:minx,miny,
        maxx,maxy"`` string, bbox tuple/list ``(minx, miny, maxx, maxy)`` in
        EPSG:4326, WKT string, Shapely geometry, GeoSeries, or
        GeoDataFrame-like object.
    year : int, str, or None
        Optional prediction-year filter, applied to ``determination:datetime``
        (``by-admin-conf`` layout), else to a ``year`` column, else to a
        ``time`` column. The ``by-admin-conf`` layout publishes 2024 and 2025
        only; other years log a WARNING (and return no rows) when that layout
        is queried, by default or through a *source_url* under its prefix
        (including the ``https://data.source.coop/ftw/global-data/`` alias).
    label : str or None
        Optional label filter. Defaults to ``"field"``. If the tile lacks a
        ``label`` column (the ``by-admin-conf`` layout holds fields only), no
        label filtering is applied.
    clip : bool
        If ``True``, clip returned polygons to the AOI. The published
        ``metrics:area`` and ``metrics:perimeter`` describe the whole
        polygon, so for every polygon that crosses the AOI boundary they are
        recomputed from the clipped geometry (area in EPSG:6933, m²;
        geodesic perimeter on WGS 84, m) and the added boolean column
        ``agribound:clipped`` is True. If ``False``, return the full
        published polygons that intersect the AOI, with their published
        metrics.
    output_path : str, Path, or None
        Optional destination path. Supported formats are those accepted by
        :func:`agribound.io.vector.write_vector`, including GeoParquet,
        GeoJSON, and GeoPackage.
    output_format : str or None
        Optional output format override.
    source_url : str or None
        PyArrow backend: GeoParquet file, directory, S3 prefix or glob (default:
        the prefix of *layout*; ``https://data.source.coop/ftw/global-data/``
        URLs are mapped to the raw S3 prefix). Manifest backend: base URL used
        to resolve relative tile paths in a manifest, or a URL/path to a
        manifest when ``manifest_path`` is not provided. HTTP/HTTPS candidate
        tiles are downloaded to ``cache_dir`` before reading.
    manifest_path : str, Path, or None
        Path or HTTP/HTTPS URL to a local tile manifest. The manifest may be a
        vector file with tile geometries or a tabular file with tile paths and
        bbox columns.
    tile_dir : str, Path, or None
        Directory containing local GeoParquet tiles. Used to resolve relative
        paths in a manifest. If no manifest is provided, a manifest is built
        from tile-level GeoParquet metadata in this directory.
    cache_dir : str, Path, or None
        Directory for downloaded remote manifests or candidate tiles, and for
        the cached partition-bounding-box index of the PyArrow backend
        (default ``~/.cache/agribound/ftw``).
    source_backend : str
        Source backend: ``"auto"``, ``"pyarrow"``, or ``"manifest"``.
        In auto mode, local manifests/tile directories use the manifest backend;
        otherwise the public FTW GeoParquet source is queried with PyArrow.
    max_features : int or None
        Optional row limit for PyArrow-backed preview or smoke-test queries.
    columns : list[str], tuple[str, ...], or None
        Optional tile columns to read and return. Internal filter and
        deduplication columns are read when needed but dropped from the result
        unless explicitly requested.
    deduplicate : bool
        Drop repeated polygons, such as the same polygon read from two
        overlapping tiles. Rows are duplicates when their normalized geometry
        is identical (SHA-1 of the normalized WKB) and, whenever the
        prediction year can be read, so is the year, so the same polygon
        predicted in 2024 and 2025 is kept once per year. The first row is
        kept; rows with a missing or empty geometry are always kept. The
        published ``id`` is not used as the key because it is not unique per
        polygon in the ``by-admin-conf`` layout: in ``US_NM`` (read
        2026-09-27) 190,940 ids in 2024 and 495,230 in 2025 are each shared by
        2-4 different polygons, typically hundreds of kilometres apart, while
        no geometry repeats within a year.
    dst_crs : str, int, or None
        Optional CRS for the returned GeoDataFrame.
    min_confidence : float or None
        Keep polygons with ``confidence >= min_confidence`` (0-100 scale). The
        dataset README (``predictions/vectors/README.md`` on Source
        Cooperative, read 2026-09-26) recommends 69
        (:data:`agribound.ftw_arrow.RECOMMENDED_MIN_CONFIDENCE`, raw 0.4).
        Values in (0, 1] log a WARNING
        (they look like raw 0-1 values). Raises :class:`ValueError` if the
        source has no ``confidence`` column.
    keep_null_confidence : bool
        Keep polygons whose confidence is null (default True). A null
        confidence means the 500 m confidence raster has no data at the
        field's point-on-surface, not a low score. False drops them, with or
        without *min_confidence*. Coverage is uneven: on 2026-09-27 the
        confidence was null for 99.7 % of the rows of ``AU_NSW`` and all rows
        of ``US_NM`` (1.8 % for France), so there *min_confidence* filters
        almost nothing unless null values are dropped. A WARNING is logged
        when *min_confidence* is set and null-confidence polygons are kept.
    layout : str
        Published layout used when *source_url* is None: ``"by-admin-conf"``
        (default; ``alpha/results-by-admin-conf``, partitioned by country and
        subdivision; only the partitions whose bounding box intersects the
        AOI are read) or ``"raw"`` (legacy ``alpha/results`` with
        ``label``/``time`` columns).
    provenance : bool
        With *output_path*, also write ``<output_path>.provenance.json``
        (default True): the query parameters, the resolved backend and
        source, the counts of duplicates dropped, clipped and returned
        polygons, and the package versions (see
        :func:`query_provenance_record`).

    Returns
    -------
    geopandas.GeoDataFrame
        Published FTW polygons intersecting the AOI. The ``bbox`` struct
        column is dropped unless requested in *columns*. For the PyArrow
        backend ``attrs["ftw_query"]`` records the source, the number of files
        listed and opened, the filters, the number of duplicates dropped, the
        number of returned polygons with a null confidence
        (``n_null_confidence``, when the source has that column) and
        ``n_returned``.

    Notes
    -----
    This function is a query/download helper for existing FTW prediction
    polygons. It does not run FTW inference and should not be used to treat FTW
    predictions as ground truth.
    """
    if layout not in FTW_VECTOR_LAYOUTS:
        raise ValueError(f"Unknown FTW layout {layout!r}. Choose from {tuple(FTW_VECTOR_LAYOUTS)}")
    parameters = {
        "study_area": _describe_study_area(study_area),
        "year": year,
        "label": label,
        "clip": clip,
        "deduplicate": deduplicate,
        "source_url": source_url,
        "manifest_path": None if manifest_path is None else str(manifest_path),
        "tile_dir": None if tile_dir is None else str(tile_dir),
        "source_backend": source_backend,
        "layout": layout,
        "min_confidence": min_confidence,
        "keep_null_confidence": keep_null_confidence,
        "max_features": max_features,
        "columns": None if columns is None else list(columns),
        "dst_crs": None if dst_crs is None else str(dst_crs),
        "output_format": output_format,
    }

    def finish(result: gpd.GeoDataFrame, info: dict[str, Any]) -> gpd.GeoDataFrame:
        out = _finalize_result(result, requested_columns, output_path, output_format, dst_crs)
        info["n_returned"] = len(out)
        if clip and "agribound:clipped" in result.columns:
            info["n_clipped"] = int(result["agribound:clipped"].sum())
        out.attrs["ftw_query"] = info
        if output_path is not None and provenance:
            from agribound.provenance import write_provenance

            record = query_provenance_record(parameters, info, aoi_4326)
            out.attrs["provenance_path"] = str(write_provenance(output_path, record))
        return out

    requested_columns = _normalize_columns(columns)
    aoi = _coerce_study_area(study_area)
    aoi = _ensure_crs(aoi, "EPSG:4326")
    aoi_4326 = aoi.to_crs("EPSG:4326")
    aoi_geom = _union_geometry(aoi_4326)

    if aoi_geom is None or aoi_geom.is_empty:
        result = _empty_ftw_gdf(requested_columns, crs="EPSG:4326")
        return finish(result, {"backend": None, "note": "empty study area"})

    backend = _resolve_source_backend(
        source_backend=source_backend,
        source_url=source_url,
        manifest_path=manifest_path,
        tile_dir=tile_dir,
    )
    if backend == "pyarrow":
        if year is not None and _targets_by_admin_conf(source_url, layout):
            _warn_unpublished_year(year)
        result = query_ftw_arrow(
            study_area_bounds=tuple(aoi_4326.total_bounds),
            source_url=source_url,
            year=year,
            label=label,
            columns=requested_columns,
            max_features=max_features,
            min_confidence=min_confidence,
            keep_null_confidence=keep_null_confidence,
            layout=layout,
            index_cache_dir=cache_dir,
        )
        query_info = {"backend": "pyarrow", **dict(result.attrs.get("ftw_query") or {})}
        result = _ensure_crs(result, "EPSG:4326")
        if result.empty and len(result.columns) <= 1:
            defaults = _BY_ADMIN_EMPTY_COLUMNS if layout == "by-admin-conf" else None
            result = _empty_ftw_gdf(requested_columns, crs="EPSG:4326", default_columns=defaults)
        if deduplicate and not result.empty:
            n_before = len(result)
            result = _deduplicate_ftw(result)
            query_info["n_duplicates_dropped"] = n_before - len(result)
        if clip and not result.empty:
            result = _clip_to_aoi(result, aoi_4326)
        _report_null_confidence(
            result, query_info.get("min_confidence"), keep_null_confidence, query_info
        )
        return finish(result, query_info)

    min_confidence = validate_min_confidence(min_confidence)

    manifest, tile_base = _load_or_build_manifest(
        manifest_path=manifest_path,
        tile_dir=tile_dir,
        source_url=source_url,
        cache_dir=cache_dir,
    )
    manifest = _ensure_crs(manifest, "EPSG:4326").to_crs("EPSG:4326")
    candidates = _select_candidate_tiles(manifest, aoi_4326.total_bounds)
    manifest_info: dict[str, Any] = {
        "backend": "manifest",
        "manifest": None if manifest_path is None else str(manifest_path),
        "tile_base": None if tile_base is None else str(tile_base),
        "n_candidate_tiles": len(candidates),
    }

    if candidates.empty:
        result = _empty_ftw_gdf(requested_columns, crs="EPSG:4326")
        return finish(result, manifest_info)

    parts: list[gpd.GeoDataFrame] = []
    for row in candidates.itertuples(index=False):
        tile_id = _row_value(row, "tile_id", default=None)
        tile_ref = _row_value(row, "tile_path", default=None)
        if tile_ref is None:
            logger.warning("Skipping FTW tile without tile_path: %s", row)
            continue

        tile_path = _resolve_tile_path(
            tile_ref,
            tile_base=tile_base,
            tile_dir=tile_dir,
            cache_dir=cache_dir,
        )
        try:
            tile = _read_ftw_tile(tile_path, requested_columns)
        except Exception as exc:
            logger.warning("Failed reading FTW tile %s: %s", tile_path, exc)
            continue

        tile = _prepare_tile(tile, aoi_4326, label=label, year=year)
        tile = _filter_confidence(tile, min_confidence, keep_null_confidence)
        if tile.empty:
            continue

        if "source_tile_id" not in tile.columns:
            tile["source_tile_id"] = (
                str(tile_id) if tile_id is not None else Path(str(tile_ref)).stem
            )
        parts.append(tile)

    if parts:
        result = gpd.GeoDataFrame(
            pd.concat(parts, ignore_index=True, sort=False),
            geometry="geometry",
        )
        result = _ensure_crs(result, "EPSG:4326")
    else:
        result = _empty_ftw_gdf(requested_columns, crs="EPSG:4326")

    manifest_info["n_tiles_read"] = len(parts)
    if deduplicate and not result.empty:
        n_before = len(result)
        result = _deduplicate_ftw(result)
        manifest_info["n_duplicates_dropped"] = n_before - len(result)

    if clip and not result.empty:
        result = _clip_to_aoi(result, aoi_4326)

    _report_null_confidence(result, min_confidence, keep_null_confidence)
    return finish(result, manifest_info)

query_provenance_record

query_provenance_record(parameters: dict[str, Any], query_info: dict[str, Any], aoi_4326: GeoDataFrame) -> dict[str, Any]

Provenance record written next to a :func:query_ftw output.

Parameters:

Name Type Description Default
parameters dict

The query arguments (the study area described, not embedded).

required
query_info dict

attrs["ftw_query"] of the result: backend, the source and filters, n_duplicates_dropped, n_clipped and n_returned (for the PyArrow backend also the files listed and opened; for the manifest backend the candidate tiles and the tiles read).

required
aoi_4326 GeoDataFrame

The study area in EPSG:4326 (its bounds are recorded).

required

Returns:

Type Description
dict

JSON-serialisable record with kind="query_ftw".

Source code in agribound/ftw_query.py
def query_provenance_record(
    parameters: dict[str, Any], query_info: dict[str, Any], aoi_4326: gpd.GeoDataFrame
) -> dict[str, Any]:
    """Provenance record written next to a :func:`query_ftw` output.

    Parameters
    ----------
    parameters : dict
        The query arguments (the study area described, not embedded).
    query_info : dict
        ``attrs["ftw_query"]`` of the result: ``backend``, the source and
        filters, ``n_duplicates_dropped``, ``n_clipped`` and ``n_returned``
        (for the PyArrow backend also the files listed and opened; for the
        manifest backend the candidate tiles and the tiles read).
    aoi_4326 : geopandas.GeoDataFrame
        The study area in EPSG:4326 (its bounds are recorded).

    Returns
    -------
    dict
        JSON-serialisable record with ``kind="query_ftw"``.
    """
    import datetime as _dt
    import platform
    import sys

    from agribound._repro import collect_versions
    from agribound._version import __version__
    from agribound.provenance import PROVENANCE_SCHEMA_VERSION, to_jsonable

    bounds = [float(v) for v in aoi_4326.total_bounds] if len(aoi_4326) else None
    return to_jsonable(
        {
            "schema_version": PROVENANCE_SCHEMA_VERSION,
            "kind": "query_ftw",
            "agribound_version": __version__,
            "created_utc": _dt.datetime.now(_dt.UTC).isoformat(timespec="seconds"),
            "data": (
                "published Fields of The World (FTW) Global prediction polygons (model "
                "outputs, CC-BY-4.0), not reference boundaries"
            ),
            "parameters": parameters,
            "aoi_bounds_4326": bounds,
            "query": query_info,
            "versions": collect_versions(("pyarrow", "fsspec", "s3fs")),
            "platform": platform.platform(),
            "python": sys.version,
        }
    )

ftw_arrow

PyArrow backend for the published Fields of The World (FTW) Global polygons.

The polygons are predictions of the PRUE model published on Source Cooperative (source.coop/ftw/global-data, CC-BY-4.0). This module reads them; it does not run FTW inference, and the polygons are model output, not ground truth.

Layouts

"by-admin-conf" (default; :data:FTW_VECTOR_LAYOUTS) predictions/vectors/alpha/results-by-admin-conf/admin:country_code=<CC>/ <name>.parquet: fiboa/vecorel GeoParquet partitioned by country (large countries by subdivision), with columns id, geometry, bbox, metrics:area, metrics:perimeter, determination:datetime (timestamp, UTC; 1 January of the prediction year), determination:method, admin:country_code, admin:subdivision_code and confidence. Years 2024 and 2025 are published (collection temporal extent 2024-01-01 to 2025-12-31, checked 2026-09-26). "raw" The older predictions/vectors/alpha/results Spark output (1000 parts) with geometry, time, label (field, non_field_background, field_boundaries) and bbox columns and no confidence.

Only the files whose data bounding box intersects the study area are opened (partitioning=None). The bounding box of a file comes from the Parquet row-group statistics of its bbox columns, read from the file footer; the GeoParquet geo metadata bbox is used only when those statistics are missing, because the published geo bboxes of subdivided countries are wrong (see :func:footer_bbox). For remote sources the boxes are read once and cached in a small JSON index, keyed by each file's path, size and modification time, under index_cache_dir (default $XDG_CACHE_HOME/agribound/ftw or ~/.cache/agribound/ftw). The per-partition STAC items reachable from predictions/vectors/collection.json carry correct bboxes (checked for Australia on 2026-09-27), but they need one request per partition plus one per country sub-catalog, so they are not faster than the footers and are not used.

Confidence

confidence is on a 0-100 scale: the 500 m PRUE confidence raster sampled at each field's point-on-surface and rescaled raw / 0.578178 * 100 (clamped to 100). The dataset README recommends confidence >= 69 (raw 0.4) as the default reliability filter (:data:RECOMMENDED_MIN_CONFIDENCE). Null means that the confidence raster has no data in the cell of the field's point-on-surface ("Cells with no data become null"), not a low score, so null values are kept unless keep_null_confidence=False. Confidence describes 500 m cell-level model reliability, not the geometric accuracy of individual polygons. Source: predictions/vectors/README.md and collection.json of the dataset (read 2026-09-26).

FTW_VECTOR_LAYOUTS module-attribute

FTW_VECTOR_LAYOUTS: dict[str, str] = {'by-admin-conf': FTW_GLOBAL_DATA_S3 + 'predictions/vectors/alpha/results-by-admin-conf/', 'raw': FTW_GLOBAL_DATA_S3 + 'predictions/vectors/alpha/results/'}

Published polygon layouts (see the module docstring).

DEFAULT_FTW_LAYOUT module-attribute

DEFAULT_FTW_LAYOUT = 'by-admin-conf'

PUBLISHED_YEARS module-attribute

PUBLISHED_YEARS: tuple[int, ...] = (2024, 2025)

Prediction years in the published by-admin-conf layout (as of 2026-09-26).

RECOMMENDED_MIN_CONFIDENCE module-attribute

RECOMMENDED_MIN_CONFIDENCE = 69.0

Reliability filter recommended by the dataset README (confidence >= 69, raw 0.4).

query_ftw_arrow

query_ftw_arrow(study_area_bounds: tuple[float, float, float, float] | list[float], source_url: str | Path | None = None, year: int | str | None = None, label: str | None = 'field', columns: list[str] | tuple[str, ...] | None = None, max_features: int | None = None, *, min_confidence: float | None = None, keep_null_confidence: bool = True, layout: str = DEFAULT_FTW_LAYOUT, index_cache_dir: str | Path | None = None) -> gpd.GeoDataFrame

Query published FTW polygons that intersect a bounding box.

Parameters:

Name Type Description Default
study_area_bounds tuple

(minx, miny, maxx, maxy) in EPSG:4326.

required
source_url (str, Path or None)

GeoParquet file, directory, S3 prefix or glob (e.g. ".../results-by-admin-conf/admin:country_code=*/*.parquet"). https://data.source.coop/ftw/global-data/... is mapped to the raw S3 prefix. None uses the prefix of layout.

None
year (int, str or None)

Prediction year, matched against determination:datetime, else a year or time column ([Jan 1, Jan 1 of the next year) in UTC).

None
label str or None

Keep rows with this label (raw layout); ignored when the source has no label column (by-admin-conf holds fields only).

'field'
columns list or None

Extra columns to read (all default columns present are read anyway).

None
max_features int or None

Stop after this many rows (preview queries).

None
min_confidence float or None

Keep rows with confidence >= min_confidence (0-100 scale).

None
keep_null_confidence bool

Keep rows whose confidence is null (default True). False drops them, also when min_confidence is None.

True
layout str

"by-admin-conf" or "raw"; selects the default source.

DEFAULT_FTW_LAYOUT
index_cache_dir (str, Path or None)

Directory for the cached partition-bounding-box index.

None

Returns:

Type Description
GeoDataFrame

Rows in EPSG:4326. attrs["ftw_query"] records the source, the number of files listed and opened, and the filters applied.

Raises:

Type Description
ValueError

For an unknown layout, or a confidence filter on a source without a confidence column.

Source code in agribound/ftw_arrow.py
def query_ftw_arrow(
    study_area_bounds: tuple[float, float, float, float] | list[float],
    source_url: str | Path | None = None,
    year: int | str | None = None,
    label: str | None = "field",
    columns: list[str] | tuple[str, ...] | None = None,
    max_features: int | None = None,
    *,
    min_confidence: float | None = None,
    keep_null_confidence: bool = True,
    layout: str = DEFAULT_FTW_LAYOUT,
    index_cache_dir: str | Path | None = None,
) -> gpd.GeoDataFrame:
    """Query published FTW polygons that intersect a bounding box.

    Parameters
    ----------
    study_area_bounds : tuple
        ``(minx, miny, maxx, maxy)`` in EPSG:4326.
    source_url : str, Path or None
        GeoParquet file, directory, S3 prefix or glob (e.g.
        ``".../results-by-admin-conf/admin:country_code=*/*.parquet"``).
        ``https://data.source.coop/ftw/global-data/...`` is mapped to the raw
        S3 prefix. *None* uses the prefix of *layout*.
    year : int, str or None
        Prediction year, matched against ``determination:datetime``, else a
        ``year`` or ``time`` column (``[Jan 1, Jan 1 of the next year)`` in UTC).
    label : str or None
        Keep rows with this ``label`` (``raw`` layout); ignored when the
        source has no ``label`` column (``by-admin-conf`` holds fields only).
    columns : list or None
        Extra columns to read (all default columns present are read anyway).
    max_features : int or None
        Stop after this many rows (preview queries).
    min_confidence : float or None
        Keep rows with ``confidence >= min_confidence`` (0-100 scale).
    keep_null_confidence : bool
        Keep rows whose confidence is null (default True). False drops
        them, also when *min_confidence* is None.
    layout : str
        ``"by-admin-conf"`` or ``"raw"``; selects the default source.
    index_cache_dir : str, Path or None
        Directory for the cached partition-bounding-box index.

    Returns
    -------
    geopandas.GeoDataFrame
        Rows in EPSG:4326. ``attrs["ftw_query"]`` records the source, the
        number of files listed and opened, and the filters applied.

    Raises
    ------
    ValueError
        For an unknown layout, or a confidence filter on a source without a
        ``confidence`` column.
    """
    if layout not in FTW_VECTOR_LAYOUTS:
        raise ValueError(f"Unknown FTW layout {layout!r}. Choose from {tuple(FTW_VECTOR_LAYOUTS)}")
    min_confidence = validate_min_confidence(min_confidence)
    source = _normalize_source_coop_path(str(source_url or FTW_VECTOR_LAYOUTS[layout]))
    minx, miny, maxx, maxy = (float(v) for v in study_area_bounds)
    filesystem, files = _list_parquet_files(source)
    bounds = _file_bounds(filesystem, files, source, index_cache_dir)
    selected = [
        f.path
        for f in files
        if bounds.get(f.path) is None or _bbox_intersects(bounds[f.path], (minx, miny, maxx, maxy))
    ]
    info: dict[str, Any] = {
        "source": source,
        "layout": layout if source_url is None else None,
        "n_files_listed": len(files),
        "n_files_opened": len(selected),
        "year": None if year is None else int(year),
        "label": label,
        "min_confidence": min_confidence,
        "keep_null_confidence": keep_null_confidence,
    }
    if not selected:
        gdf = _empty_result(columns)
        gdf.attrs["ftw_query"] = info
        return gdf

    ds = _ds()
    dataset = ds.dataset(selected, filesystem=filesystem, format="parquet", partitioning=None)
    schema_names = set(dataset.schema.names)
    if "geometry" not in schema_names:
        raise ValueError("FTW GeoParquet source must contain a geometry column.")

    expr = _bbox_filter(dataset, schema_names, minx, miny, maxx, maxy)
    if label is not None and "label" in schema_names:
        expr = expr & (ds.field("label") == str(label))
    if year is not None:
        year_expr = _year_filter_expression(dataset, schema_names, int(year))
        if year_expr is not None:
            expr = expr & year_expr
    conf_expr = _confidence_expression(schema_names, min_confidence, keep_null_confidence)
    if conf_expr is not None:
        expr = expr & conf_expr

    read_columns = _select_columns(dataset.schema.names, columns)
    scanner = dataset.scanner(columns=read_columns, filter=expr)
    table = scanner.head(int(max_features)) if max_features is not None else scanner.to_table()
    gdf = _table_to_geodataframe(table)
    if not gdf.empty and year is not None:
        gdf = _filter_year(gdf, year)
    info["n_rows"] = len(gdf)
    gdf.attrs["ftw_query"] = info
    return gdf