Skip to content

Module vignette 6: process drone imagery

Notebook: View the drone notebook in the repository. GitHub displays the cells; clone or download the file to run them.

Use this module for local drone TIFF packages or HDF5 inputs. Discovery is recursive, the original package remains traceable, and TIFF inputs enter the same established HDF5, ENVI, topographic, and BRDF correction machinery.

Scientific branches

DRONE
TIFF + ancillary + manifest
  -> working H5 -> ENVI -> topo/BRDF -> corrected MicaSense
  -> affine cross-sensor translation -> Landsat-like raster
  -> full or polygon spectral library -> QA

NORMAL NEON
NEON H5 -> ENVI -> topo/BRDF -> corrected hyperspectral
  -> spectral convolution -> Landsat-like raster
  -> full or polygon spectral library -> QA

OPTIONAL VALIDATION
actual Landsat
            \
drone-like --- actual-Landsat grid -> pairwise QA
            /
NEON-like

The paths are shared through corrected ENVI and then diverge. Drone data use reviewed affine coefficients derived by the independent bulk workflow. NEON hyperspectral data use spectral convolution. The corrected native MicaSense raster remains a first-class output and is never overwritten. A translated product is Landsat-like; it is not an actual Landsat observation.

Prepare inputs and coefficients

Place valid drone HDF5 files or reflectance TIFF packages under one input directory. TIFF packages may include aligned terrain/view sidecars and are matched to the bundled field manifest unless drone_manifest_path overrides it. See the detailed tutorial for the complete TIFF contract.

Production translation consumes a static, versioned registry generated from the completed compact bulk output. It does not fit coefficients. The production weighting is fixed to site_balanced, reflecting intended transfer to new flightlines and sites rather than optimization of the pooled pixel fit. SpectralBridge validates equation direction, sensor identities, the exact 18 physical pair-band mappings, wavelengths, finite coefficients, confidence, and bulk-run provenance before writing a translated product.

Run it

from spectralbridge import run_drone_pipeline

results = run_drone_pipeline(
    input_h5_dir="drone_inputs",
    output_dir="drone_outputs",
    polygon_path="plots.geojson",
    extraction_mode="polygon",
    apply_topo=True,
    apply_brdf=True,
    apply_translation=True,
    translation_strict=False,
    drone_manifest_timezone="UTC",
)

print(results["processed"])
print(results["translation_outputs"])
print(results["matched_source_outputs"])

run_drone_pipeline() now treats missing requested products as an incomplete run, rather than returning a successful-looking result. It validates the working H5, raw and corrected ENVI pairs, requested full or polygon Parquet, every requested translated target and library, stage/provenance JSON, flight QA, and the final report. Requested topo and BRDF corrections must actually be applied. A failure raises DronePipelineIncompleteError; its results attribute and drone_qa_summary.json retain per-flight reasons. Set raise_on_incomplete=False only when intentionally collecting a structured partial batch result. A polygon run with no intersecting pixels is incomplete, not a successful extraction. Optional actual-Landsat comparison remains non-blocking when no acceptable observation is available. Flights whose required scientific geometry cannot be validated or safely repaired are reported separately in results["blocked"]; they are not mislabeled as software failures.

Use extraction_mode="full" for all corrected pixels. Omitting extraction_mode preserves the earlier behavior: polygon extraction when a polygon is supplied, otherwise raster and QA outputs only. Translation remains opt-in and uses the packaged site-balanced registry by default. It checks corrected input values against the bulk fit's numeric domain before writing; apparently fractional input is refused rather than silently rescaled. Normal mode writes caution bands with warnings, while translation_strict=True refuses them. A reject coefficient is never applied.

Standalone and comparison QA

Every translated run produces per-target translation JSON and PNG diagnostics without NEON or network access. They show band mapping, coefficient evidence, corrected-versus-translated values, reflectance shifts, valid fractions, nodata preservation, and unusual translated values.

Set landsat_qa=True to request an overlapping Landsat Collection 2 Level 2 scene from Microsoft Planetary Computer. Install this optional support with:

python -m pip install "earthlab-spectralbridge[landsat]"

Alternatively, use landsat_product="landsat_stack.tif" to supply an analysis-ready multiband raster or a previously cached observation manifest. The comparison selects the closest acceptable scene within landsat_search_days, applies the Landsat QA pixel mask, and aggregates the translated drone product onto the actual Landsat grid using valid-aware area averaging. Failed searches, cloud rejection, or insufficient overlap are recorded as optional-QA limitations and do not fail the core drone run.

Add comparison_neon_product="neon_landsat_like.img" to compare an existing normal-pipeline NEON product on the same Landsat support. The three reported pairs are drone-like versus actual, NEON-like versus actual, and drone-like versus NEON-like. NEON is optional and is never processed by drone code.

Restart and outputs

Solar geometry deserves a separate review before accepting corrected products. The original HDF5 or TIFF package is authoritative and read-only. Discovery ignores generated __working.h5 files. Every run starts from the original source, then reuses or rebuilds the derived working copy according to its stage signature. Passing a working H5 as a source is rejected because it loses the provenance needed to decide whether a repair remains valid.

For HDF5 input, the working-copy stage compares embedded solar arrays with an independent scene-center position computed from the acquisition datetime and georeference. A valid array is preserved. A missing or materially inconsistent array is replaced only in the working copy when the datetime has an explicit timezone, scene coordinates are valid, the expected sun is above the horizon, and manifest provenance authorizes the repair. The canonical Solar_Zenith_Angle and Solar_Azimuth_Angle datasets are then used by topographic/BRDF correction. The source H5 fingerprint is verified before and after repair. If those conditions are not met, the flight is blocked_scientific instead of being silently corrected with questionable geometry.

For TIFF input, aligned solar rasters take precedence over explicit scalar angles, which take precedence over manifest-derived geometry. Naive manifest times are localized with drone_manifest_timezone (default "UTC"); pass a verified IANA zone when the campaign used local civil time. Ambiguous DST times, malformed dates, missing coordinates, conflicting duplicate manifest rows, and below-horizon solutions block the affected flight while the campaign continues.

Each flight audit records source and UTC-normalized acquisition time, timezone, coordinate source and scene center, embedded and expected angle summaries, zenith and circular-azimuth residuals, validation status, geometry actually used, repair decision/reason, and source fingerprint. The calculation reuses the package's approximate NOAA-style solar-position equations, not a precise ephemeris. The 5° validation tolerance is inclusive and applies to zenith plus circular azimuth residuals. These are workflow validation bounds, not universal calibration limits, and the code never substitutes 90 - angle as an undocumented repair.

On a production VM, inspect existing working H5 files without rerunning corrections or rewriting products:

PYTHONPATH=src python scripts/diagnose_drone_solar_geometry.py \
  /home/jovyan/data-store/SpectralBridge_Drone_2023_2024_Production/flight_outputs \
  --output /home/jovyan/data-store/SpectralBridge_Drone_2023_2024_Production/solar_validation/solar_geometry.csv \
  --candidate-timezone UTC --candidate-timezone America/Denver

The two timezone candidates in the CSV are hypothetical interpretations until the field manifest's convention is independently verified. Review source dataset units, scale/fill attributes, collection date, scene CRS, and circular residuals together. A filename/manifest date mismatch requires provenance review because a package filename need not be the acquisition date.

Valid working H5, corrected ENVI, translated ENVI, and translated Parquet products are reused when their inputs and translation signatures match. Legacy H5 sun-angle arrays under Reflectance/Metadata, including nested to-sun_* datasets, are exposed through lightweight links in the working copy; source H5 files are not changed. Continue a campaign by rerunning the same command against the original input root and same output root. Missing or signature-invalid stages resume from the first incomplete stage. Do not set overwrite=True for an ordinary continuation.

Each completed flight directory contains spectralbridge_flightline.json and matched native-MicaSense/Landsat-like ENVI pairs. Consequently run_bulk_pipeline(drone_outputs, ..., input_mode="auto") discovers completed drone flights directly; no manual rename or campaign-wide pixel merge is needed. Per-flight Parquets remain authoritative. Set merge_extractions=True only when a legacy consumer explicitly requires the optional run-level merged tables.

Start bulk work with preflight_only=True. Bulk inventories the per-flight Parquet footers, schemas, row counts, sizes, QA state, sensors, and translation availability without scanning pixels or writing into the drone tree. The translated Landsat-like products are applications of an existing coefficient registry, not independent Landsat observations, so the ordinary drone campaign is analyzed as derived_application_verification. Bulk writes compact fits and LOSO outputs to verify application consistency, with coefficient metadata marked diagnostic_application_verification_only. Do not feed those circular diagnostics back into the registry or interpret them as empirical calibration. The canonical shared-band contracts are 4 bands for TM/ETM+ and 5 for OLI/OLI-2; the 6/7-band contracts remain specific to NEON convolution products.

With full extraction, the six canonical per-flight analysis tables are the native corrected __full.parquet, four sensor-specific __landsat_like_<target>_translated_envi.parquet tables, and the optional combined Landsat-like table used by validated production packaging. Polygon extraction uses the corresponding __polygons.parquet forms. Every translated table carries the translation pair, source/target sensor, coefficient hash, and evidence-boundary columns; bulk validates these fields from the Parquet schema.

Artifact Role Regeneration rule
Original H5/TIFF package Authoritative immutable input Never generated or modified
__working.h5 Derived, restartable bridge and repair target Reuse only when source/config/solar signature matches
ENVI, translation, extraction, and QA outputs Derived scientific products Reuse only when their stage contracts validate
spectralbridge_flightline.json and stage/provenance JSON Identity and audit contract Required and deterministic

Optional Landsat failure does not invalidate or recompute core products. See outputs and naming for the complete contract, and review every QA JSON before treating a coefficient application as trustworthy.

Continue