Pipeline
Outputs & File Structure¶
Outputs on disk are the primary interface of SpectralBridge. Downstream analysis should rely on these files rather than return values from the Python API.
Local drone processing
Drone translation and QA contract¶
Drone TIFF packages enter the validated working-H5 bridge and share the normal export and correction machinery through corrected ENVI. Translation is an opt-in affine, wavelength-aware stage using an explicitly selected bulk coefficient family. It never overwrites corrected native MicaSense and does not use convolution.
| Output type | Drone filename pattern | Description |
|---|---|---|
| Flight identity | <flight>/spectralbridge_flightline.json |
Generic flight ID, site, acquisition date, platform, and authoritative source identity used by direct bulk discovery. |
| Working H5 | <flight>__working.h5 |
NeonCube-compatible bridge retaining the original package, manifest, spatial, wavelength, nodata, ancillary, and solar-geometry provenance. |
| Native ENVI | <flight>__envi.(img\|hdr) |
Native drone reflectance exported from the working H5. |
| Corrected MicaSense | <flight>__corrected.(img\|hdr) |
Native MicaSense after requested topographic/BRDF correction; retained independently of translation. |
| Landsat-like raster | <flight>__landsat_like_<target>_translated_envi.(img\|hdr) |
Affine translated product with target-band wavelengths. It is not an actual Landsat observation. |
| Matched native source raster | <flight>__micasense_to_match_tm_etm+_envi.(img\|hdr) or <flight>__micasense_to_match_oli_oli2_envi.(img\|hdr) |
Exact wavelength-selected native MicaSense bands paired with translated targets; 4 bands for TM/ETM+ and 5 for OLI/OLI-2. This is the source side of direct bulk analysis. |
| Translation provenance | <translated_stem>__translation.json |
Coefficient fingerprint/run/weighting, exact equation, sensor pair, per-band wavelength mapping and coefficients, source fingerprint, range checks, output summaries, and warnings. |
| Native library | <flight>__polygons.parquet or <flight>__full.parquet |
Requested polygon or full-pixel extraction from corrected native MicaSense. |
| Landsat-like library | <translated_stem>__polygons.parquet or <translated_stem>.parquet |
Four canonical sensor-specific tables per extraction mode. Each preserves 4 shared TM/ETM+ bands or 5 shared OLI/OLI-2 bands plus source-package, working-H5, corrected/translated raster, coefficient, acquisition-time, and band-mapping provenance. |
| Combined Landsat-like library | <flight>__landsat_like_combined[__polygons].parquet |
Optional sixth per-flight analysis product used by validated production packaging: a combined view of the four sensor-specific translated tables. It is cataloged, not used as the raster regression source. |
| Optional merged libraries | drone_merged.parquet, drone_landsat_like_merged.parquet |
Legacy run-level pixel-table merges written only when merge_extractions=True; not required by bulk analysis. |
| Standalone translation QA | <translated_stem>__translation_qa.(png\|json) |
Corrected-versus-translated distributions, slopes/intercepts, shifts, valid fractions, nodata, range/evidence warnings, and provenance. |
| Optional common-support QA | qa_common_support/*__on_actual_landsat_grid.tif, *__comparison.json, *__landsat_common_support_qa.png |
Drone-like, optional NEON-like, and actual Landsat comparisons after valid-aware aggregation to the actual Landsat grid. |
| Run audit | drone_qa_summary.json |
Core status, optional-QA availability/reasons, source and product paths, correction and translation provenance, software version, timestamps, and warnings. |
| Stage records | <flight>__working.stage.json, <flight>__envi__export.stage.json, <flight>__corrected__correction.stage.json |
Source/config fingerprints, deterministic outputs, status, and timestamp used to decide reuse. |
| QA dashboard | qa/summary/drone_qa_summary.(png|json) |
First-page run status, ingest/correction/translation state, valid support, coefficient provenance, external-validation metrics, and warnings. |
| Publication translation figure | <flight>/qa_publication/*__translation_quality.(png|pdf) |
Compact wavelength-aware MicaSense-to-Landsat-like translation summary. |
| Final QA report | qa_summary.pdf, with qa/report.stage.json |
Dashboard first, followed by available diagnostic/publication pages; regenerable from compact QA products. |
A drone flight is counted as complete only after its requested H5, ENVI, correction, translation, extraction, provenance, and QA artifacts validate. Incomplete runs raise by default while preserving a structured audit; unsafe solar inputs are classified separately as blocked_scientific. Valid derived outputs are reused when their signatures remain current. Continuation always starts from the authoritative original input root—generated __working.h5 files are ignored by discovery and rejected as direct sources. Optional actual-Landsat or NEON comparison unavailability does not invalidate core products.
Cross-run analysis
Independent bulk-pipeline contract¶
The optional spectralbridge-bulk workflow consumes completed or minimally staged scientific flightline directories beneath arbitrary storage folders. Identity comes from a generic manifest or another configured parser, never the outer folder. Completed drone output roots are directly discoverable because each flight contains the generic identity manifest plus matched native-MicaSense, translated Landsat-like ENVI, and per-flight tabular products. Bulk inventories Parquet footers and streams scientifically eligible raster relationships in place, writing only compact products to a separate output; it does not modify normal or drone runs. Canonical NEON names and prebuilt merged Parquets remain compatible inputs.
| Output type | Canonical path | Description |
|---|---|---|
| Flightline catalog | catalog/flightlines.parquet |
Scientific identity, site/date, processing completeness, product availability, profile eligibility, and duplicate/rejection status. |
| Source catalog | catalog/source_files.parquet |
Every upstream product and derived observation source with original path, role/sensor, dimensions, dtype, wavelengths, metadata fingerprint, and selection status. |
| Source-product catalog | catalog/source_products.parquet |
Read-only raw/corrected/target ENVI plus canonical per-flight Parquet inventory, including product key, storage format, semantics, rows, schema, and sizes; derived caches are excluded. |
| Duplicate/rejection catalogs | catalog/duplicates.parquet, catalog/rejected_sources.parquet |
Explicit exclusions; duplicate canonical IDs are never silently double-counted. |
| Structured exclusions | catalog/exclusions.(parquet|json|csv) |
Deterministic reason codes, affected scientific units/products, offending paths, details, and processing stage. |
| Per-flightline statistics | statistics/flightlines/<flight_id>/ |
Mergeable sufficient statistics, signatures, optional bounded sample, and restart/failure status. |
| Collection statistics | statistics/translation_sufficient_statistics.parquet |
Compact flightline/band moments used for hierarchical model fitting. |
| Bulk database | database/spectralbridge_bulk.duckdb |
Catalogs, compact statistics, exclusions, provenance, and modular analysis tables. |
| Bulk observations | database/bulk_observations.parquet |
Explicit legacy dataset-build output; absent from normal analysis. |
| Dataset census | analyses/dataset_census/ |
Metadata-only preflight JSON, report, and Parquet breakdowns. |
| Campaign summary | reports/campaign_summary.md |
Flights, sites/dates, products, schemas, rows, bytes, QA, exclusions, translation availability, and analyses run or intentionally omitted. |
| Analysis decisions | reports/analysis_decisions.json |
Machine-readable scientific eligibility decision separating available translated products from independent regression evidence. |
| Translation analyses | analyses/sensor_translation/ |
Pixel-pooled, per-flightline, per-site, flightline-balanced, and site-balanced regressions. |
| Leave-one-site-out | analyses/leave_one_site_out/ |
Held-out-site generalization metrics. |
| Candidate coefficients | coefficients/candidate_translation_coefficients.(parquet|json) |
Pooled and balanced source-to-target summaries with selected-pair provenance. |
| Translation interpretation | analyses/bulk_results/ |
Pair-band summaries with spectral identity and source/target wavelengths, weighting comparisons, flightline/site stability, LOSO transferability, configurable attention flags, and restart metadata derived only from compact result tables. |
| Translation diagnostics | figures/bulk_results/diagnostics/*.png |
Detailed wavelength-ordered weighting, correction, heterogeneity, and held-out-site comparisons. |
| QA dashboard | figures/bulk_results/summary/bulk_qa_summary.(png|pdf) |
Accepted/excluded counts, rows, fit metrics, correction magnitude, weakest cases, and warning count. |
| Publication panels | figures/bulk_results/publication/*.(png|pdf) |
Translation performance, stability, and unseen-site generalization/failure cases using the shared accessible palette. |
| Translation report | reports/bulk_results/bulk_translation_results.(md|pdf) |
Portable narrative plus a dashboard-first multipage report assembled only from compact outputs. |
| Spectral-library summaries | analyses/spectral_library/ |
Compact species/band summaries, approximate quantiles, medians, group counts, robust/full plot ranges, bounded extreme-spectrum diagnostics, and provenance from an optional existing merged polygon Parquet. |
| Spectral-library reports | figures/spectral_library/ |
Explicit summary or full multipage low-alpha variability PDFs, including primary robust and separate full-range audit views; source observations are read in place and not copied. |
| Bulk manifest | catalog/bulk_manifest.json |
Restart signature, execution settings, counts, and artifact names. |
Normal completed-flightline analysis has no observation population copy. The optional results report reads only compact model outputs and does not require the source archive. NEON convolution products keep their full 6/7-band schemas and synthetic_convolution semantics. Canonical drone targets keep only the 4/5 shared translated bands and use landsat_like_translated plus affine_cross_sensor_translation. Their compact regressions are application-verification diagnostics marked diagnostic_application_verification_only, not new registry candidates or empirical calibration. Mixed evidence classes are cataloged together but are not silently pooled.
Canonical outputs
Per-flightline contract¶
Naming stems come from spectralbridge.paths.FlightlinePaths and spectralbridge.utils.naming.get_flightline_products. Sensor-specific stems come from SensorProductPaths.
| Output type | Canonical filename pattern | Description | Notes / guarantees |
|---|---|---|---|
| Raw ENVI (when available) | <flight_id>_envi.(img|hdr|parquet) |
Direct export of the NEON HDF5 reflectance cube. | The ENVI pair is the upstream raster contract; its Parquet sidecar depends on the extraction path. |
| BRDF model JSON | <flight_id>_brdf_model.json |
Scene-level BRDF coefficient tables written before BRDF application. | Includes iso/vol/geo, kernel settings, ndvi_binning_enabled, and ndvi_edges. |
| BRDF + topographic corrected ENVI | <flight_id>_brdfandtopo_corrected_envi.(img|hdr|json|parquet) |
Physics-informed normalization and correction JSON produced before sensor resampling. | The ENVI pair and correction JSON persist before extraction; the Parquet sidecar is conditional. |
| Sensor-resampled ENVI + Parquet | <flight_id>_<sensor>_envi.(img|hdr|parquet) |
Reflectance cubes resampled into the configured target sensor frame. | ENVI products persist after convolution; Parquet sidecars depend on the extraction path. |
| Merged Parquet | <flight_id>_merged_pixel_extraction.parquet |
Master table that merges Parquet sidecars into one analysis-ready spectral library. | Produced by full-pixel extraction; polygon-only runs may omit it. It is not required for bulk archive discovery. |
| QA artefacts | <flight_id>_qa.png, <flight_id>_qa.json, optional <flight_id>_qa.pdf |
Visual and numeric QA summaries aligned to the merged outputs. | PNG and JSON are expected for completed runs; PDF is optional. |
| QA metrics parquet | <flight_id>_qa_metrics.parquet |
Structured QA metrics by band and sensor. | Emitted alongside QA outputs when QA calculation runs. |
| Synthetic sensor regression diagnostic | qa_plots/<merged_stem>__MS_vs_Landsat_FIXED.(png|json) |
Scatter panels compare spectral-identity/wavelength-matched synthetic MicaSense and Landsat products; the JSON records separate source/target band indices and wavelengths plus the displayed slope, intercept, correlation, R², and sample count. | Both axes come from the same corrected NEON source. This is a descriptive convolution diagnostic, not empirical sensor calibration. |
Drone translation writes a distinct
<flight>__landsat_like_<target>_translated_envi.(img|hdr) pair plus
*__translation.json; it never overwrites <flight>__corrected.*. The JSON
records the fixed affine equation, coefficient-set and bulk-run identities,
site-balanced policy, per-band spectral names/wavelengths, slope/intercept,
confidence status, R²/RMSE, correction magnitude, flightline/site/LOSO evidence,
attention flags, artifact hashes, and evidence-boundary warning. Translated
full/polygon Parquets keep an explicit Landsat-like product identifier and the
same band mapping/provenance as constant columns rather than mixing their values
with native MicaSense columns.
| Stage QA | qa/stages/<order>_<stage>/stage_qa.(json|html) plus optional overview.png | Focused report for one canonical stage with explicit checks and provenance. | Deterministic and restart-safe; missing diagnostics are recorded as NOT EVALUATED. |
| Combined stage QA | qa/combined/combined_qa.(json|html|pdf) plus pipeline_evolution.png | Cross-stage status, pipeline evolution, evidence-backed synthesis, and a printable multi-page summary. | Does more than concatenate stage reports; unsupported translation/Landsat diagnostics remain explicit. The PDF is intended for download and flightline-to-flightline comparison. |
Success criteria
What a successful run looks like¶
- The products required by the requested stages exist and are readable. A completed correction/convolution run has corrected and configured target-sensor ENVI pairs even if full-pixel extraction was not requested.
- The QA PNG renders with its matching JSON:
<flight_id>_qa.pngand<flight_id>_qa.json. - Sensor-specific ENVI and parquet products exist as configured; absence can reflect configuration rather than failure.
- If full-pixel extraction was requested, its merged Parquet should also exist and validate. Polygon-only processing has its own polygon output contract.
Restart safety
Idempotence and validation¶
process_one_flightline and go_forth_and_multiply skip stages whose outputs already exist and validate, so re-running the pipeline does not recompute completed products unless outputs are missing or invalid.
Stage QA uses the same principle: reports are reused when the schema, parameters, thresholds, software version, and input/output artifact fingerprints match.
Drone polygon workflows also attempt to write CSV sidecars next to parquet outputs for portability. The parquet files remain the authoritative outputs.
How to use these files
Recommended downstream usage¶
Load parquet first
Use merged parquet outputs directly for most analysis tasks.
Inspect QA before modeling
Review QA PNG and JSON outputs to confirm spectral health and calibration quality.
Treat ENVI as diagnostic
Intermediate ENVI products remain useful for inspection, but many workflows only need parquet and QA outputs.