Scorecard
Field, Coverage, and Geometry are three separate fidelity questions (TL field accuracy, shadow-mask accuracy, out-of-plane accuracy vs the BELLHOP3D reference solver) that are never collapsed into one number. Composite — the first column, sorted highest-first by default — is a secondary 60/20/20 blend of the three, shown for convenience; it is not the scientific result. Core err uses a winsorized (outlier-clipped) RMSE over BELLHOP-insonified cells so isolated caustic-adjacent spikes don't dominate. Tie-breaking is deterministic: lower core error, then lower coverage error, then lower receiver error, then panel id — never a "within N points" window. Each panel's numerical self-check (reciprocity, convergence, canonical-grid integrity — Qualified / Provisional / Invalid) is informational and does not gate ranking eligibility.
Visual comparison
Per-metric bar comparison across all models vs BELLHOP3D reference — best on the left, worst on the right
Metric glossary
How each scorecard column is defined and scored vs BELLHOP3D
= model TL · = BELLHOP3D TL, per grid cell ( cells) · = receiver · shadow cutoff 120 dB.
100×(1 − clamp(Efield, 0, 1)). Higher wins. How close the predicted TL field is to BELLHOP3D — the primary "did the physics track the reference" question, kept separate from mask and geometry accuracy.
RMSE over BELLHOP-insonified cells dB, but each per-cell residual is winsorized (clipped) to ±20 dB before squaring — a robust statistic so isolated caustic-adjacent spikes (a known soft spot of geometric-spreading TL) don't dominate. The dominant 50% term of Field Fidelity.
Core-region RMSE after a 3×3×3 box blur is applied to both the model and BELLHOP3D fields first. Isolates whether large-scale energy distribution tracks the reference, separately from fine interference-pattern (constructive/destructive) mismatch. 25% of Field Fidelity.
The transmission loss in dB at the single receiver point each panel reports (canonical 41×31). Labeled a single-receiver diagnostic — the harness code accepts an array of receiver errors so more receiver points can be added later without a scoring rewrite.
Absolute deviation of the model's receiver TL from BELLHOP3D's, median-clipped at 25 dB across the (currently single-element) receiver-error array. 25% of Field Fidelity.
In a reciprocal medium, source→receiver and receiver→source loss must match. Feeds the Validation gate (must be ≤ 3 dB, or absent for Provisional) rather than a score weight.
How much the receiver TL shifts when the ray fan is refined (count ). Feeds the Validation gate (must be ≤ 5 dB, or absent for Provisional) rather than a score weight.
100×(1 − clamp(Ecoverage, 0, 1)). Higher wins. false-shadow = fraction of BELLHOP-lit cells the model marks shadow; false-light = fraction of BELLHOP-shadow cells the model marks lit — a balanced average, so an unbalanced lit/shadow split can't dilute either failure mode.
Share of grid cells the model marks as insonified (TL below the 120 dB shadow cutoff). A raw diagnostic, not a scored quantity — Coverage Fidelity uses the directional false-shadow/false-light rates above.
Average symmetric distance between the model's and BELLHOP3D's shadow/lit boundary voxels, from a 6-connected grid-distance search (not exact Euclidean — an approximate, grid-quantized diagnostic). Reported when the full canonical grid is available; prefer this over a raw Hausdorff distance, which one outlier voxel can dominate.
100×(1 − clamp(Etotal, 0, 1)), computed for every canonical model. If Geometry is unavailable for a model, its weight is re-normalized across the remaining dimensions and the result is marked provisional (shown as "· *"). This is a convenience summary, not the scientific result — Field, Coverage, and Geometry each have their own leader.
Root-mean-square gap between the model's transmission-loss field and the BELLHOP3D reference solver over every cell of the 101×49×31 grid, shadow zones included. A raw diagnostic — Field Fidelity uses the robust, insonified-only Core err instead.
Diff overlay
TL(model) − TL(BELLHOP3D) · centerline y = 12 km range–depth slice