Build a bulk cross-run analysis¶
SpectralBridge has three related, but separate, workflows:
- Single flightline: process one hyperspectral input through correction, target-sensor products, extraction, and QA.
- Multiple flightlines: apply that workflow independently to a collection; each flightline keeps its own restart-safe outputs.
- Bulk analysis: discover completed or minimally staged flightlines, validate them, stream immutable products in place, and fit population-level comparisons from compact sufficient statistics.
The bulk workflow is downstream analysis. It does not download inputs, rerun correction or convolution, invoke the drone pipeline, or modify source folders.
Runnable notebooks: Use the local bulk analysis notebook
when the completed-flightline tree is already on disk. Use the
CyVerse drone production notebook when remote
ExportPackages must first be processed. That notebook is a thin call to
run_drone_bulk_production(); remote listing, source-H5 transfer, producer
iteration, canonical validation, completeness gating, bulk execution, result
packaging, and upload live in tested package code. Neither interface changes
the bulk pipeline's scientific definitions.
Generic input model¶
A flightline is the atomic scientific unit. Storage folders around it have no scientific meaning and may be nested arbitrarily:
completed_products/
├── transfer_group_01/
│ ├── flight_alpha/
│ │ ├── spectralbridge_flightline.json
│ │ ├── sensor_a.img
│ │ ├── sensor_a.hdr
│ │ ├── sensor_b.img
│ │ └── sensor_b.hdr
│ └── flight_beta/
│ └── ...
└── any/other/layout/
└── flight_gamma/
└── ...
For a generic flightline, add
spectralbridge_flightline.json to the scientific-unit directory:
{
"flightline_id": "flight-alpha",
"site": "example-site",
"acquisition_date": "2024-06-15"
}
Only flightline_id is required. site and an ISO acquisition_date are
optional. Canonical NEON directory and merged-product names remain supported
by a built-in identity parser. Custom Python callers can provide other identity
parsers without changing discovery.
Multiple flightlines under one parent remain independent. If one scientific ID
appears in different directories, all copies are reported in
catalog/duplicates.parquet and excluded; SpectralBridge does not guess which
copy is authoritative.
Prebuilt *_merged_pixel_extraction.parquet and
*_polygons_merged_pixel_extraction.parquet inputs remain supported through
input_mode="merged_parquet". Automatic mode prefers identifiable flightline
directories and otherwise uses this compatibility path.
Completed drone output roots are native input_mode="auto" inputs. Each
flight's spectralbridge_flightline.json supplies identity, while the matched
native-MicaSense and affine translated Landsat-like ENVI pairs describe the
available translation relationship. Canonical per-flight Parquets are cataloged
from their footers for row counts, schemas, sizes, roles, and provenance. Bulk
reads these products in place; do not rename them and do not create a
campaign-wide pixel Parquet first.
Availability is not the same as independent calibration evidence. Drone
Landsat-like products were created from the same flight's native MicaSense
observation with an existing coefficient registry. Bulk accepts the canonical
4-band TM/ETM+ and 5-band OLI/OLI-2 source/target pairs and fits them as
application-verification diagnostics. The resulting coefficient and LOSO
artifacts are marked diagnostic_application_verification_only: they should
reproduce the supplied registry and must not be treated as a new empirical
calibration. NEON convolution uses the full 6-band TM/ETM+ and 7-band OLI/OLI-2
target contracts and remains a separate synthetic_convolution evidence class.
Analysis profiles and minimal archives¶
Processing completeness, product availability, and analysis eligibility are
separate catalog fields. The built-in translation profile requires a complete
requested sensor relationship, not a raw or corrected hyperspectral cube. A
minimal archive containing only the two target ENVI products and their headers
is therefore valid. Raw/corrected cubes, QA, plots, HTML, and unrelated sensors
are optional unless a custom profile requires them.
Product recognition and translation relationships are centralized in a
ProductRegistry. The installed defaults describe SpectralBridge's current
matched MicaSense/Landsat products, while ProductDescriptor and
TabularProductDescriptor record product roles, semantics, processing stage,
extraction mode, required columns, and expected band count.
TranslationPair supports other sensor names, matching groups, compatible
source/target schemas, and explicit source-to-target band mappings. The
double-underscore drone names are intentionally distinct from the single-
underscore NEON convolution names.
Band numbers are sensor-local and are never treated as a cross-sensor identity. The installed relationships resolve named spectral identities against the packaged center wavelengths and passbands. Thus Landsat 5 TM blue band 1 corresponds to Landsat 8 OLI blue band 2; OLI band 1 is coastal aerosol. The matched MicaSense products already have sensor-specific subsets: their TM/ETM+ blue band is band 1, while their OLI/OLI-2 blue band is band 2 because that product also contains the coastal-aerosol match.
To select one installed relationship, use its pair key:
spectralbridge-bulk /data/completed_products \
--output-dir /data/bulk_analysis \
--analysis translation \
--translation-pair MicaSense_to-match_OLI_and_OLI-2__to__Landsat_8_OLI \
--preflight-only
--sensor is repeatable and retains only relationships whose source and target
are both selected. A requested relationship never requires unrelated products
that merely happen to be supported by the package.
Run preflight first¶
Use a fresh output directory outside the read-only source tree:
from spectralbridge import run_bulk_pipeline
result = run_bulk_pipeline(
"/data/completed_products",
"/data/bulk_analysis",
analysis="translation",
preflight_only=True,
)
print(result["preflight"])
The structured preflight result reports discovered, accepted, duplicate, and
excluded flightlines; the selected profile and relationship keys; required and
optional product roles; available and regression-eligible relationships;
per-flight tabular product counts, rows, bytes, and schema consistency; QA
states; selected source bytes; estimated compact output and bounded diagnostic
sample; temporary-disk estimate; whether a pixel dataset was requested;
exclusion counts; package version; and all output locations. The linked census
adds site/date/product inventories.
Preflight reads paths, sizes, modification times, ENVI headers, small QA JSON, and Parquet footers where applicable. It does not scan or copy raster pixels.
Review these outputs before the full run:
catalog/flightlines.parquetcatalog/source_products.parquetcatalog/exclusions.parquetcatalog/exclusions.jsoncatalog/exclusions.csvcatalog/duplicates.parquetanalyses/dataset_census/dataset_census.mdanalyses/dataset_census/by_product.parquetanalyses/dataset_census/missing_products.parquetreports/campaign_summary.mdreports/analysis_decisions.json
Missing sidecars, zero-byte files, unreadable metadata, invalid dimensions,
incompatible bands, incomplete requested pairs, duplicate products, duplicate
identities, transient source disappearance, and streaming-analysis failures have stable
reason codes. on_invalid="exclude" is the population-safe default. Valid
flightlines continue when another one fails. Use on_invalid="error" when any
ineligible flightline should make the call raise after diagnostic catalogs are
written. Invalid optional products remain visible but do not invalidate an
otherwise eligible flightline.
Run the eligible population analysis¶
After reviewing preflight, rerun without preflight_only:
result = run_bulk_pipeline(
"/data/completed_products",
"/data/bulk_analysis",
analysis="translation",
)
For completed-flightline inputs, eligible source and target ENVI products are
opened together and read in deterministic aligned windows. Spatial dimensions,
transform, and CRS must match. A pixel is eligible only when the existing
sensor-wide ENVI no-data/finite rules and the requested reflectance threshold
all pass. Each chunk is immediately reduced into numerically stable bivariate
moments. No sensor Parquet or observations.parquet is created.
For a canonical drone campaign produced by run_drone_pipeline(), this full
call streams the matched native and translated rasters, writes bounded
sufficient statistics, and emits application-verification regression, LOSO,
and candidate-coefficient artifacts. The analysis decision and coefficient JSON
record derived_application_verification and
diagnostic_application_verification_only; these outputs verify coefficient
application and are not independent sensor calibration. A mixed catalog may
contain both drone and NEON flights, but a full run intentionally does not pool
drone application-verification with NEON synthetic-convolution evidence. Run
those evidence classes separately when coefficients are required.
The first pass writes compact per-flightline statistics. Site, global,
flightline-balanced, and site-balanced fits combine those checkpoints. LOSO
training uses global statistics - held-out site statistics; one bounded second
pass computes exact absolute-error and held-out metrics. extraction_workers
remains the backward-compatible name for concurrent flightline readers and
defaults to one so disk bandwidth is not oversubscribed.
Completed flightline products (read only)
|
v
Catalog + validation
|
v
Aligned chunked direct reads
|
v
Per-flightline sufficient statistics
|
v
Population aggregation
+-- translation models
+-- site and balanced models
+-- LOSO
+-- bounded diagnostics
For merged-Parquet inputs, input_kind="full" selects full-pixel tables,
"polygon" selects polygon tables, and "both" deliberately combines them.
Combining both can count polygon pixels a second time and should be a conscious
scientific choice.
The translation equation is always recorded explicitly:
target reflectance = slope × source reflectance + intercept
Outputs include pixel-pooled, per-flightline, per-site, flightline-balanced, and site-balanced fits, plus leave-one-site-out validation. Pixels are nested within flightlines and sites; the balanced results prevent the largest rasters from silently becoming the only effective replicates. Built-in matched MicaSense/Landsat products are synthetic products from one corrected source, so their coefficients are diagnostic candidates—not empirical field calibration. Translation coefficients are also distinct from the upstream percentage brightness adjustment.
Output contract¶
bulk_analysis/
├── catalog/
│ ├── flightlines.parquet
│ ├── source_files.parquet
│ ├── source_products.parquet
│ ├── exclusions.parquet
│ ├── exclusions.json
│ ├── exclusions.csv
│ ├── duplicates.parquet
│ ├── rejected_sources.parquet
│ └── bulk_manifest.json
├── statistics/
│ ├── translation_sufficient_statistics.parquet
│ ├── diagnostic_sample.parquet # optional and globally bounded
│ └── flightlines/<flightline-id>/
│ ├── sufficient_statistics.parquet
│ ├── diagnostic_sample.parquet
│ ├── statistics_metadata.json
│ └── status.json
├── database/
│ ├── spectralbridge_bulk.duckdb
│ └── bulk_observations.parquet # explicit dataset build only
├── analyses/
│ ├── dataset_census/
│ ├── sensor_translation/
│ ├── leave_one_site_out/
│ ├── bulk_results/ # optional compact interpretation
│ └── spectral_library/ # optional compact summaries
├── coefficients/
│ ├── candidate_translation_coefficients.parquet
│ └── candidate_translation_coefficients.json
├── tables/
├── figures/
│ ├── bulk_results/
│ │ ├── summary/ # one-page PNG/PDF dashboard
│ │ ├── diagnostics/ # detailed diagnostic PNGs
│ │ └── publication/ # three PNG/PDF manuscript panels
│ └── spectral_library/ # optional summary/full PDFs
├── reports/
│ └── bulk_results/ # optional Markdown and PDF reports
└── logs/
DuckDB stores catalogs, exclusions, sufficient statistics, models, QA summaries,
and provenance. In completed-flightline mode the compatibility
bulk_observations view is empty because analysis does not require an observation
layer. Original merged Parquets remain virtual read-in-place inputs.
Interpret a completed translation run¶
The reusable results layer starts after run_bulk_pipeline has completed. It
requires only these compact artifacts:
catalog/bulk_manifest.jsoncoefficients/candidate_translation_coefficients.parquetanalyses/sensor_translation/per_flightline.parquetanalyses/sensor_translation/per_site.parquetanalyses/leave_one_site_out/leave_one_site_out.parquet
The source raster archive, per-flightline sufficient-statistics checkpoints, diagnostic sample, observation cache, and DuckDB database are not opened or required. This makes the completed bulk output independently portable for results review.
from spectralbridge import summarize_bulk_results
from spectralbridge.bulk import BulkResultsConfig
results = summarize_bulk_results(
"/data/bulk_analysis",
config=BulkResultsConfig(
r2_review_threshold=0.90,
loso_r2_review_threshold=0.80,
),
make_figures=True,
make_report=True,
)
print(results["overview"])
print(results["attention_flags"])
Each pair-band uses one common representative source reflectance for comparing
the pixel-pooled, flightline-balanced, and site-balanced fitted corrections.
By default that value is the pixel-pooled source mean; set
representative_source_value to use one explicit reflectance instead. The
calculation is recorded as:
100 × ((slope × source value + intercept) - source value) / source value
The generated analyses/bulk_results/ tables separate six questions:
weighting_comparison.parquet: do pooled and balanced coefficients agree, and what correction does each imply at the common value?flightline_stability.parquet: how broad is the per-flightline slope and fit distribution, and which flightline is weakest?site_stability.parquet: how much do site coefficients differ, and which site is weakest?loso_transferability.parquet: how well does a model trained without each site transfer to that held-out site?attention_flags.parquet: which weak fits, large corrections, weighting differences, heterogeneity, site dependence, incomplete results, or weak LOSO evaluations cross configured review thresholds?pair_band_summary.parquet: what evidence should be reviewed together for each translation relationship and band?
bulk_results_summary.json records all thresholds and compact input hashes.
An unchanged rerun reuses valid outputs; changing inputs or configuration
rebuilds them. A corrupt or missing downstream figure/report is regenerated
without rerunning coefficients or source-raster analysis. The optional figures
and Markdown/PDF reports contain no hard-coded
production statistics: every count and metric is calculated from the supplied
completed result tables. Plot axes are ordered by spectral identity/wavelength,
and labels show both the physical source wavelength and target sensor-local band
and wavelength. The first page answers whether the run completed and ranks the
most weighting-sensitive, site-dependent, and weakest LOSO cases. The
pair-band summary also records spectral identity, separate source and target
wavelengths, wavelength difference, and the matching basis. Known built-in
sensor rows with incompatible identities fail reporting instead of being
silently plotted together.
High R² indicates a strong relationship, but it does not establish sensor
interchangeability. A pair-band with
no_configured_warning_triggered has passed only the configured screen; it has
not received universal scientific approval. In particular, inspect weighting
sensitivity, flightline and site heterogeneity, the observed reflectance domain,
and LOSO performance before promoting any candidate coefficient. These bulk
translation candidates also remain distinct from the fixed upstream brightness
adjustment.
Analysis and pixel-level dataset construction are separate operations:
Completed flightline products
|
v
EXPLICIT DATASET BUILD
|
v
Harmonized pixel-level dataset
Use the explicit compatibility builder only when the combined pixel dataset is itself the requested scientific product:
from spectralbridge import build_harmonized_dataset
dataset = build_harmonized_dataset(
"/data/completed_products",
"/data/harmonized_pixel_dataset",
threads=8,
memory_limit="16GB",
temp_directory="/scratch/spectralbridge_bulk",
)
materialize_observations=True remains temporarily available for backward
compatibility. New code should use build_harmonized_dataset so the disk cost
and scientific intent are explicit. A production analysis invocation is:
spectralbridge-bulk /data/completed_products \
--output-dir /data/bulk_analysis \
--extraction-workers 1 \
--extraction-chunk-size 1024 \
--diagnostic-sample-size 100000 \
--diagnostic-seed 42
Visualize an existing polygon spectral library¶
The intentional merged polygon spectral library is different from the removed
temporary observation cache. Pass its existing Parquet path separately: it is
opened read-only and is never copied into the bulk output. The adapter inspects
the footer before analysis, detects the species and available hierarchy fields,
and selects one compatible wavelength-bearing spectral stage. It prefers
corr_*_wl*nm columns when present; multiple other stages require an explicit
choice. Wavelengths are parsed from column names, numerically sorted, and kept
paired to their original values. Duplicate wavelengths, duplicate band indices,
non-numeric spectral columns, and schemas without physical wavelengths fail
explicitly.
from spectralbridge import inspect_spectral_library_preflight, run_bulk_pipeline
from spectralbridge.bulk import SpectralLibraryPlotConfig
plots = SpectralLibraryPlotConfig(
spectral_stage="corr",
species_sort="count_desc", # or "alphabetical"
panels_per_page=4,
trace_batch_size=2_000,
summary_band_batch_size=32,
max_traces_per_group=None, # default: show every valid trace
raster_dpi=150,
plot_y_quantiles=(0.005, 0.995),
species_y_scale="global_robust",
spectral_plot_minimum_reflectance=None,
)
library = "/data/library/polygons_merged_pixel_extraction.parquet"
preflight = inspect_spectral_library_preflight(library, config=plots)
print(preflight["schema"])
print(preflight["counts"])
print(preflight["largest_species"])
print(preflight["expected_pages"])
result = run_bulk_pipeline(
"/data/completed_products",
"/data/bulk_analysis",
spectral_library=library,
make_summary_plots=True,
make_full_spectral_reports=True,
spectral_library_config=plots,
)
print(result["preflight"]["spectral_library_schema"])
print(result["spectral_library"]["reports"])
Use the public preflight call before launching a bulk run. It reads the Parquet in place and returns the source size and signature, detected fields and spectral stage, row and hierarchy counts, largest and median species sizes, groups over 10,000 and 100,000 valid spectra, trace-cap state, expected pages, and estimated logical scans. It writes no summaries or PDFs. The recommended production sequence is: (1) preflight, (2) run with both plot flags false to create compact summaries and inspect robust bounds/extremes, (3) enable summary plots, then (4) enable the full suite.
make_summary_plots=True writes the between-species median overview and species
observation-count report. make_full_spectral_reports=True also writes the
species low-alpha ensemble, nested quantile envelopes, and available site,
flightline, and within-species hierarchy reports. It also implies the two
summary PDFs. No report is generated by default because large libraries can
require substantial sequential I/O and rendering time.
The trace alpha is deterministic:
alpha = min(0.03, max(0.003, 0.2 / sqrt(rendered trace count)))
This makes tens of spectra individually legible and progressively lowers alpha
for dense groups. It is a qualitative spectral trace density, not a normalized
probability density. The raw trace layer is rendered batch by batch into a
fixed-size raster; titles, axes, annotations, and the median remain vector PDF
elements. By default every complete finite spectrum contributes. Sampling is
used only when max_traces_per_group (or
--spectral-max-traces-per-group) is explicitly set, in which case a seeded
DuckDB hash order makes the cap reproducible.
The compact products are species_summary.parquet,
species_band_summary.parquet, species_quantiles.parquet,
species_median_spectra.parquet, group_counts.parquet,
species_plot_ranges.parquet, and extreme_spectra.parquet. Quantiles default
to 2.5, 10, 25, 50, 75, 90, and 97.5 percent and use DuckDB
approx_quantile in bounded band batches.
Visualization validity is intentionally separate from regression validity.
Every selected-stage value must be present and finite; explicit nodata sentinels
from recognized Parquet metadata and nodata_values (default -9999) are
excluded within nodata_tolerance. Finite negative corrected reflectance
remains visible by default. Set spectral_plot_minimum_reflectance only when a
visualization-specific lower threshold is scientifically warranted. This does
not change translation/regression validity.
The primary species report defaults to one common robust range, estimated from
the 0.5th and 99.5th percentiles without loading the library into memory. Values
outside it are clipped graphically only: they remain in analytical medians,
quantiles, min/max, and counts. Plot-range provenance records the bounds,
method, value and spectrum exceedance counts, and per-species indicators. The
bounded extreme_spectra.parquet ranks up to
max_extreme_spectra_per_species observations per species with source context.
species_y_scale="global_full" uses one observed full range, while
"per_group_robust" maximizes within-species detail at the cost of direct
cross-species scale comparison. The separate full-range audit PDF is always
part of the full suite; other scale combinations are not generated
automatically.
Memory is bounded by one Arrow trace batch plus at most the configured page's fixed-size panel rasters during plotting, and by one configured band batch during summary calculation. The companion JSON records source signature, package version, run ID, detected schema, counts, wavelength range, exact plot configuration, alpha rule, rasterization, sampling, visualization validity, plot ranges, extreme diagnostics, report-cost estimates, report paths, page counts, and sizes.
The full suite uses these canonical names when the corresponding grouping field exists:
spectral_library_species_variability.pdfspectral_library_species_variability_full_range.pdfspectral_library_species_quantiles.pdfspectral_library_species_medians.pdfspectral_library_observation_counts.pdfspectral_library_hierarchical_variability.pdfspectral_library_site_variability.pdfspectral_library_flightline_variability.pdf
These reports are additional products. Dataset census, translation, cross- sensor, LOSO, and other QA/statistical outputs remain unchanged.
Restart and provenance¶
The source tree remains read only. The run manifest records package version, available git commit, profile, complete registry/pair configuration, source and output roots, accepted/excluded units, execution settings, timestamps, and artifact names. Raster fingerprints use path, size, modification time, header hash, and readable ENVI metadata. Every flightline checkpoint records its source signatures, algorithm/schema version, band and validity configuration, sampling configuration, and package version.
An unchanged complete invocation returns status="reused". Changing a selected
source invalidates only that flightline checkpoint; changing a profile,
relationship, registry, integrity rule, or streaming setting invalidates the
affected configuration. A failed streaming analysis is cataloged while other
flightlines continue. Pass force=True or
--force for an intentional rebuild.
Advanced extension¶
The immutable AnalysisProfile, ProductDescriptor, ProductRegistry, and
TranslationPair types are importable from spectralbridge.bulk. Custom
registries belong in calling code or a future package extension, not in the
source archive. Discovery/statistics construction is separate from the independently
callable run_dataset_census, run_sensor_translation, and
run_leave_one_site_out analysis modules, so new analyses do not need to
reimplement source validation.
The large production campaign that motivated these checks is a validation case, not a package data model: no site code, storage hierarchy, flightline count, or campaign-specific folder name is embedded in the bulk architecture.