Skip to content

Pipeline

Outputs & File Structure

Outputs on disk are the primary interface of SpectralBridge. Downstream analysis should rely on these files rather than return values from the Python API.

Local drone processing

Drone translation and QA contract

Drone TIFF packages enter the validated working-H5 bridge and share the normal export and correction machinery through corrected ENVI. Translation is an opt-in affine, wavelength-aware stage using an explicitly selected bulk coefficient family. It never overwrites corrected native MicaSense and does not use convolution.

Output type Drone filename pattern Description
Flight identity <flight>/spectralbridge_flightline.json Generic flight ID, site, acquisition date, platform, and authoritative source identity used by direct bulk discovery.
Working H5 <flight>__working.h5 NeonCube-compatible bridge retaining the original package, manifest, spatial, wavelength, nodata, ancillary, and solar-geometry provenance.
Native ENVI <flight>__envi.(img\|hdr) Native drone reflectance exported from the working H5.
Corrected MicaSense <flight>__corrected.(img\|hdr) Native MicaSense after requested topographic/BRDF correction; retained independently of translation.
Landsat-like raster <flight>__landsat_like_<target>_translated_envi.(img\|hdr) Affine translated product with target-band wavelengths. It is not an actual Landsat observation.
Matched native source raster <flight>__micasense_to_match_tm_etm+_envi.(img\|hdr) or <flight>__micasense_to_match_oli_oli2_envi.(img\|hdr) Exact wavelength-selected native MicaSense bands paired with translated targets; 4 bands for TM/ETM+ and 5 for OLI/OLI-2. This is the source side of direct bulk analysis.
Translation provenance <translated_stem>__translation.json Coefficient fingerprint/run/weighting, exact equation, sensor pair, per-band wavelength mapping and coefficients, source fingerprint, range checks, output summaries, and warnings.
Native library <flight>__polygons.parquet or <flight>__full.parquet Requested polygon or full-pixel extraction from corrected native MicaSense.
Landsat-like library <translated_stem>__polygons.parquet or <translated_stem>.parquet Four canonical sensor-specific tables per extraction mode. Each preserves 4 shared TM/ETM+ bands or 5 shared OLI/OLI-2 bands plus source-package, working-H5, corrected/translated raster, coefficient, acquisition-time, and band-mapping provenance.
Combined Landsat-like library <flight>__landsat_like_combined[__polygons].parquet Optional sixth per-flight analysis product used by validated production packaging: a combined view of the four sensor-specific translated tables. It is cataloged, not used as the raster regression source.
Optional merged libraries drone_merged.parquet, drone_landsat_like_merged.parquet Legacy run-level pixel-table merges written only when merge_extractions=True; not required by bulk analysis.
Standalone translation QA <translated_stem>__translation_qa.(png\|json) Corrected-versus-translated distributions, slopes/intercepts, shifts, valid fractions, nodata, range/evidence warnings, and provenance.
Optional common-support QA qa_common_support/*__on_actual_landsat_grid.tif, *__comparison.json, *__landsat_common_support_qa.png Drone-like, optional NEON-like, and actual Landsat comparisons after valid-aware aggregation to the actual Landsat grid.
Run audit drone_qa_summary.json Core status, optional-QA availability/reasons, source and product paths, correction and translation provenance, software version, timestamps, and warnings.
Stage records <flight>__working.stage.json, <flight>__envi__export.stage.json, <flight>__corrected__correction.stage.json Source/config fingerprints, deterministic outputs, status, and timestamp used to decide reuse.
QA dashboard qa/summary/drone_qa_summary.(png|json) First-page run status, ingest/correction/translation state, valid support, coefficient provenance, external-validation metrics, and warnings.
Publication translation figure <flight>/qa_publication/*__translation_quality.(png|pdf) Compact wavelength-aware MicaSense-to-Landsat-like translation summary.
Final QA report qa_summary.pdf, with qa/report.stage.json Dashboard first, followed by available diagnostic/publication pages; regenerable from compact QA products.

A drone flight is counted as complete only after its requested H5, ENVI, correction, translation, extraction, provenance, and QA artifacts validate. Incomplete runs raise by default while preserving a structured audit; unsafe solar inputs are classified separately as blocked_scientific. Valid derived outputs are reused when their signatures remain current. Continuation always starts from the authoritative original input root—generated __working.h5 files are ignored by discovery and rejected as direct sources. Optional actual-Landsat or NEON comparison unavailability does not invalidate core products.

Cross-run analysis

Independent bulk-pipeline contract

The optional spectralbridge-bulk workflow consumes completed or minimally staged scientific flightline directories beneath arbitrary storage folders. Identity comes from a generic manifest or another configured parser, never the outer folder. Completed drone output roots are directly discoverable because each flight contains the generic identity manifest plus matched native-MicaSense, translated Landsat-like ENVI, and per-flight tabular products. Bulk inventories Parquet footers and streams scientifically eligible raster relationships in place, writing only compact products to a separate output; it does not modify normal or drone runs. Canonical NEON names and prebuilt merged Parquets remain compatible inputs.

Output type Canonical path Description
Flightline catalog catalog/flightlines.parquet Scientific identity, site/date, processing completeness, product availability, profile eligibility, and duplicate/rejection status.
Source catalog catalog/source_files.parquet Every upstream product and derived observation source with original path, role/sensor, dimensions, dtype, wavelengths, metadata fingerprint, and selection status.
Source-product catalog catalog/source_products.parquet Read-only raw/corrected/target ENVI plus canonical per-flight Parquet inventory, including product key, storage format, semantics, rows, schema, and sizes; derived caches are excluded.
Duplicate/rejection catalogs catalog/duplicates.parquet, catalog/rejected_sources.parquet Explicit exclusions; duplicate canonical IDs are never silently double-counted.
Structured exclusions catalog/exclusions.(parquet|json|csv) Deterministic reason codes, affected scientific units/products, offending paths, details, and processing stage.
Per-flightline statistics statistics/flightlines/<flight_id>/ Mergeable sufficient statistics, signatures, optional bounded sample, and restart/failure status.
Collection statistics statistics/translation_sufficient_statistics.parquet Compact flightline/band moments used for hierarchical model fitting.
Bulk database database/spectralbridge_bulk.duckdb Catalogs, compact statistics, exclusions, provenance, and modular analysis tables.
Bulk observations database/bulk_observations.parquet Explicit legacy dataset-build output; absent from normal analysis.
Dataset census analyses/dataset_census/ Metadata-only preflight JSON, report, and Parquet breakdowns.
Campaign summary reports/campaign_summary.md Flights, sites/dates, products, schemas, rows, bytes, QA, exclusions, translation availability, and analyses run or intentionally omitted.
Analysis decisions reports/analysis_decisions.json Machine-readable scientific eligibility decision separating available translated products from independent regression evidence.
Translation analyses analyses/sensor_translation/ Pixel-pooled, per-flightline, per-site, flightline-balanced, and site-balanced regressions.
Leave-one-site-out analyses/leave_one_site_out/ Held-out-site generalization metrics.
Candidate coefficients coefficients/candidate_translation_coefficients.(parquet|json) Pooled and balanced source-to-target summaries with selected-pair provenance.
Translation interpretation analyses/bulk_results/ Pair-band summaries with spectral identity and source/target wavelengths, weighting comparisons, flightline/site stability, LOSO transferability, configurable attention flags, and restart metadata derived only from compact result tables.
Translation diagnostics figures/bulk_results/diagnostics/*.png Detailed wavelength-ordered weighting, correction, heterogeneity, and held-out-site comparisons.
QA dashboard figures/bulk_results/summary/bulk_qa_summary.(png|pdf) Accepted/excluded counts, rows, fit metrics, correction magnitude, weakest cases, and warning count.
Publication panels figures/bulk_results/publication/*.(png|pdf) Translation performance, stability, and unseen-site generalization/failure cases using the shared accessible palette.
Translation report reports/bulk_results/bulk_translation_results.(md|pdf) Portable narrative plus a dashboard-first multipage report assembled only from compact outputs.
Spectral-library summaries analyses/spectral_library/ Compact species/band summaries, approximate quantiles, medians, group counts, robust/full plot ranges, bounded extreme-spectrum diagnostics, and provenance from an optional existing merged polygon Parquet.
Spectral-library reports figures/spectral_library/ Explicit summary or full multipage low-alpha variability PDFs, including primary robust and separate full-range audit views; source observations are read in place and not copied.
Bulk manifest catalog/bulk_manifest.json Restart signature, execution settings, counts, and artifact names.

Normal completed-flightline analysis has no observation population copy. The optional results report reads only compact model outputs and does not require the source archive. NEON convolution products keep their full 6/7-band schemas and synthetic_convolution semantics. Canonical drone targets keep only the 4/5 shared translated bands and use landsat_like_translated plus affine_cross_sensor_translation. Their compact regressions are application-verification diagnostics marked diagnostic_application_verification_only, not new registry candidates or empirical calibration. Mixed evidence classes are cataloged together but are not silently pooled.

Canonical outputs

Per-flightline contract

Naming stems come from spectralbridge.paths.FlightlinePaths and spectralbridge.utils.naming.get_flightline_products. Sensor-specific stems come from SensorProductPaths.

Output type Canonical filename pattern Description Notes / guarantees
Raw ENVI (when available) <flight_id>_envi.(img|hdr|parquet) Direct export of the NEON HDF5 reflectance cube. The ENVI pair is the upstream raster contract; its Parquet sidecar depends on the extraction path.
BRDF model JSON <flight_id>_brdf_model.json Scene-level BRDF coefficient tables written before BRDF application. Includes iso/vol/geo, kernel settings, ndvi_binning_enabled, and ndvi_edges.
BRDF + topographic corrected ENVI <flight_id>_brdfandtopo_corrected_envi.(img|hdr|json|parquet) Physics-informed normalization and correction JSON produced before sensor resampling. The ENVI pair and correction JSON persist before extraction; the Parquet sidecar is conditional.
Sensor-resampled ENVI + Parquet <flight_id>_<sensor>_envi.(img|hdr|parquet) Reflectance cubes resampled into the configured target sensor frame. ENVI products persist after convolution; Parquet sidecars depend on the extraction path.
Merged Parquet <flight_id>_merged_pixel_extraction.parquet Master table that merges Parquet sidecars into one analysis-ready spectral library. Produced by full-pixel extraction; polygon-only runs may omit it. It is not required for bulk archive discovery.
QA artefacts <flight_id>_qa.png, <flight_id>_qa.json, optional <flight_id>_qa.pdf Visual and numeric QA summaries aligned to the merged outputs. PNG and JSON are expected for completed runs; PDF is optional.
QA metrics parquet <flight_id>_qa_metrics.parquet Structured QA metrics by band and sensor. Emitted alongside QA outputs when QA calculation runs.
Synthetic sensor regression diagnostic qa_plots/<merged_stem>__MS_vs_Landsat_FIXED.(png|json) Scatter panels compare spectral-identity/wavelength-matched synthetic MicaSense and Landsat products; the JSON records separate source/target band indices and wavelengths plus the displayed slope, intercept, correlation, R², and sample count. Both axes come from the same corrected NEON source. This is a descriptive convolution diagnostic, not empirical sensor calibration.

Drone translation writes a distinct <flight>__landsat_like_<target>_translated_envi.(img|hdr) pair plus *__translation.json; it never overwrites <flight>__corrected.*. The JSON records the fixed affine equation, coefficient-set and bulk-run identities, site-balanced policy, per-band spectral names/wavelengths, slope/intercept, confidence status, R²/RMSE, correction magnitude, flightline/site/LOSO evidence, attention flags, artifact hashes, and evidence-boundary warning. Translated full/polygon Parquets keep an explicit Landsat-like product identifier and the same band mapping/provenance as constant columns rather than mixing their values with native MicaSense columns. | Stage QA | qa/stages/<order>_<stage>/stage_qa.(json|html) plus optional overview.png | Focused report for one canonical stage with explicit checks and provenance. | Deterministic and restart-safe; missing diagnostics are recorded as NOT EVALUATED. | | Combined stage QA | qa/combined/combined_qa.(json|html|pdf) plus pipeline_evolution.png | Cross-stage status, pipeline evolution, evidence-backed synthesis, and a printable multi-page summary. | Does more than concatenate stage reports; unsupported translation/Landsat diagnostics remain explicit. The PDF is intended for download and flightline-to-flightline comparison. |

Success criteria

What a successful run looks like

  • The products required by the requested stages exist and are readable. A completed correction/convolution run has corrected and configured target-sensor ENVI pairs even if full-pixel extraction was not requested.
  • The QA PNG renders with its matching JSON: <flight_id>_qa.png and <flight_id>_qa.json.
  • Sensor-specific ENVI and parquet products exist as configured; absence can reflect configuration rather than failure.
  • If full-pixel extraction was requested, its merged Parquet should also exist and validate. Polygon-only processing has its own polygon output contract.

Restart safety

Idempotence and validation

process_one_flightline and go_forth_and_multiply skip stages whose outputs already exist and validate, so re-running the pipeline does not recompute completed products unless outputs are missing or invalid.

Stage QA uses the same principle: reports are reused when the schema, parameters, thresholds, software version, and input/output artifact fingerprints match.

Drone polygon workflows also attempt to write CSV sidecars next to parquet outputs for portability. The parquet files remain the authoritative outputs.

How to use these files

Load parquet first

Use merged parquet outputs directly for most analysis tasks.

Inspect QA before modeling

Review QA PNG and JSON outputs to confirm spectral health and calibration quality.

Treat ENVI as diagnostic

Intermediate ENVI products remain useful for inspection, but many workflows only need parquet and QA outputs.