Reference
JSON schemas and metadata sidecars¶
SpectralBridge writes structured JSON files next to major processing artefacts so correction choices, wavelengths, QA summaries, and provenance remain inspectable after the raster work is done.
Reproducibility
Metadata sidecars record how a product was generated instead of forcing users to infer processing history from filenames alone.
Validation
QA JSON files carry machine-readable checks that complement the PNG and PDF artefacts.
Stage transparency
Each major stage can leave a structured record of what was assumed, fitted, or applied.
Export metadata
ENVI export sidecars¶
The export stage records metadata needed to interpret the ENVI cube and the downstream parquet products derived from it.
- wavelength arrays and band names
- reflectance scaling details
- mask and CRS context
- basic provenance for the exported raster
Correction metadata
Topographic and BRDF sidecars¶
Topographic metadata
Topographic sidecars summarize the slope, aspect, solar geometry, and correction parameters used for the terrain-sensitive part of the workflow.
BRDF metadata
BRDF sidecars preserve coefficient summaries, fit diagnostics, kernel settings, and any NDVI-binning context used while building the correction model.
The streamlined BRDF model JSON remains the most direct structured record of the fitted coefficient surfaces and kernel configuration for a flight line.
Harmonization metadata
Sensor translation sidecars¶
Convolution and harmonization sidecars document how corrected hyperspectral reflectance was translated into target sensor products.
- sensor response assumptions
- wavelength alignment checks
- brightness adjustment coefficients when used
- bandpass integration context and summary metrics
Bulk analysis metadata
Cross-run catalog, analysis, and manifest schemas¶
catalog/flightlines.parquet identifies scientific units and records identity source, site/date, processing completeness, product availability, profile eligibility, QA, scientific evidence status, and duplicate/rejection state. catalog/source_products.parquet records each recognized product's role, semantics, sensor, matching group, processing stage, extraction mode, paths, dimensions, band or column count, wavelengths, dtype, no-data metadata, schema, and source fingerprint. Canonical drone translated tables require 4 TM/ETM+ or 5 OLI/OLI-2 spectral columns plus pair, source/target sensor, coefficient-hash, and evidence-boundary provenance. catalog/exclusions.(parquet|json|csv) stores stable reason codes and offending paths. coefficients/candidate_translation_coefficients.json records equation direction, selected generic pair definitions, evidence class and boundary, candidate status, minimum-reflectance filter, and pixel-pooled plus balanced slope/intercept records. catalog/bulk_manifest.json ties every result to the source inventory, profile, registry, pair selection, package version, and restart signature.
analyses/bulk_results/bulk_results_summary.json records the compact input fingerprints, reporting configuration, derived overview, pair-band summaries, attention flags, artifact paths, restart signature, and wavelength-based matching contract. Its companion Parquets retain typed rows for weighting comparison, flightline stability, site stability, LOSO transferability, pair-band screening, and attention flags. Pair-band screening rows include spectral identity, separate source/target center wavelengths, their difference, and matching basis; band_index remains only the within-pair match ordinal and is not a cross-sensor band identity. Fitted correction is evaluated at one common source value per pair-band as 100 × ((slope × x_reference + intercept) - x_reference) / x_reference; the default reference is that pair-band's pixel-pooled source mean.
Attention thresholds are configurable review rules. A no_configured_warning_triggered status does not establish sensor interchangeability, universal validity, or safe extrapolation beyond the observed sites and reflectance domain.
Bulk regression JSON is an analysis output, not a packaged brightness table. The source catalog and manifest must remain with a coefficient set so its population and upstream processing state stay traceable.
Drone coefficient consumer
Translated-product provenance schema¶
The drone pipeline reads the bulk candidate coefficient Parquet or its JSON companion. It requires one selected weighting family and the existing fields for analysis run, analysis level, weighting, translation pair, source/target sensor, source/target band, shared band index, equation, fit status, slope, and intercept. Rows must encode target = slope * source + intercept, agree with the product registry, contain exactly one finite coefficient for every expected target band, and support an unambiguous wavelength mapping to the corrected MicaSense cube.
The resulting *__translation.json records the coefficient path and SHA-256 fingerprint, bulk run and weighting provenance, source and translated raster fingerprints, per-band source/target wavelengths and affine values, candidate evidence fields, observed training-range checks, output summaries, warning flags, package version, processing time, and restart signature. Translated Parquet tables repeat the stable provenance identifiers as columns and keep the complete mapping in a JSON-valued column.
Candidate coefficients are cross-sensor translation evidence, not packaged brightness adjustments or a universal empirical calibration. A diagnostic_application_verification_only artifact was fitted against outputs produced by an existing registry; it must not be fed back into that registry. The package does not silently choose among weighting families.
QA metadata
QA JSON sidecars¶
Every QA JSON file acts as the machine-readable counterpart to the PNG and optional PDF reports.
Scene-level summaries
Reflectance, masking, correction, geometry, and harmonization summaries live here for dashboards and automated review.
Provenance and workflow context
QA payloads also carry package version, created time, and other details that make a run auditable later.
How to use them
Why these files matter downstream¶
These JSON files are useful when you need to:
- reproduce a published processing run
- compare runs across software versions
- debug a suspicious corrected or harmonized output
- drive dashboards or automated quality gates from structured metadata
The raster and parquet outputs are still the main analysis products, but the JSON sidecars are the best explanation of how those products were created.
Where to go next