UWA Ray Bench
Hover a panel for c(z) · D(x,y)
Simulation Control
Fan Beams
Rays per fan, spread across a fixed ±20° vertical / ±15° horizontal sweep
Elevation
41
Azimuth
31
Playback
Scrub position
0%
Speed
Volume
Opacity
50%
Camera
★ composite awaiting canonical metrics
Benchmark
Model vs BELLHOP3D reference
BELLHOP3Dreference · BELLHOP3D solver
Loading reference…
reference previewBELLHOP3DBELLHOP3D reference · 3D ray trace▶ View standalone
Fable 5max · model
Awaiting Fable panel…
model previewFable 5awaiting score▶ Open comparison
Gemini 3.1 Prohigh · model
Awaiting Gemini panel…
model previewGemini 3.1 Proawaiting score▶ Open comparison
GPT 5.6sol · ultra model
Awaiting GPT panel…
model previewGPT 5.6 Solawaiting score▶ Open comparison
Opus 4.8max · model
Awaiting Opus panel…
model previewOpus 4.8awaiting score▶ Open comparison
Sakana Fuguultra · model
Awaiting Fugu panel…
model previewSakana Fuguawaiting score▶ Open comparison

Scorecard

Field, Coverage, and Geometry are three separate fidelity questions (TL field accuracy, shadow-mask accuracy, out-of-plane accuracy vs the BELLHOP3D reference solver) that are never collapsed into one number. Composite — the first column, sorted highest-first by default — is a secondary 60/20/20 blend of the three, shown for convenience; it is not the scientific result. Core err uses a winsorized (outlier-clipped) RMSE over BELLHOP-insonified cells so isolated caustic-adjacent spikes don't dominate. Tie-breaking is deterministic: lower core error, then lower coverage error, then lower receiver error, then panel id — never a "within N points" window. Each panel's numerical self-check (reciprocity, convergence, canonical-grid integrity — Qualified / Provisional / Invalid) is informational and does not gate ranking eligibility.

Visual comparison

Per-metric bar comparison across all models vs BELLHOP3D reference — best on the left, worst on the right

Metric glossary

How each scorecard column is defined and scored vs BELLHOP3D

T^i = model TL · Ti = BELLHOP3D TL, per grid cell i (N=101·49·31 cells) · R = receiver · shadow cutoff 120 dB.

Field FidelityTL field accuracy · 0–100 headline
Efield=0.50Ecore+ 0.25Esmooth+ 0.25Erecv

100×(1 − clamp(Efield, 0, 1)). Higher wins. How close the predicted TL field is to BELLHOP3D — the primary "did the physics track the reference" question, kept separate from mask and geometry accuracy.

Core errrobust insonified error · dB
1|C| iC (clip(T^iTi,20))2

RMSE over BELLHOP-insonified cells C={i:Ti<120} dB, but each per-cell residual is winsorized (clipped) to ±20 dB before squaring — a robust statistic so isolated caustic-adjacent spikes (a known soft spot of geometric-spreading TL) don't dominate. The dominant 50% term of Field Fidelity.

Smoothed errlarge-scale TL error · dB

Core-region RMSE after a 3×3×3 box blur is applied to both the model and BELLHOP3D fields first. Isolates whether large-scale energy distribution tracks the reference, separately from fine interference-pattern (constructive/destructive) mismatch. 25% of Field Fidelity.

TL(R)receiver loss · dB
TL(R)=20 log10 |p(R)|pref

The transmission loss in dB at the single receiver point R each panel reports (canonical 41×31). Labeled a single-receiver diagnostic — the harness code accepts an array of receiver errors so more receiver points can be added later without a scoring rewrite.

TL(R) errreceiver error · dB
ΔTL(R)=| TLmdl(R) TLref(R)|

Absolute deviation of the model's receiver TL from BELLHOP3D's, median-clipped at 25 dB across the (currently single-element) receiver-error array. 25% of Field Fidelity.

Recipself-consistency · dB
εrecip=| TL(SR) TL(RS)|

In a reciprocal medium, source→receiver and receiver→source loss must match. Feeds the Validation gate (must be ≤ 3 dB, or absent for Provisional) rather than a score weight.

Conv ΔTL(R)convergence · dB
Δconv=| TL2n(R) TLn(R)|

How much the receiver TL shifts when the ray fan is refined (count n2n). Feeds the Validation gate (must be ≤ 5 dB, or absent for Provisional) rather than a score weight.

Coverage Fidelityshadow-mask accuracy · 0–100 headline
Ecoverage=0.5( false-shadow+false-light)

100×(1 − clamp(Ecoverage, 0, 1)). Higher wins. false-shadow = fraction of BELLHOP-lit cells the model marks shadow; false-light = fraction of BELLHOP-shadow cells the model marks lit — a balanced average, so an unbalanced lit/shadow split can't dilute either failure mode.

Insonif%coverage · %
|{i:TLi<120}| N×100

Share of grid cells the model marks as insonified (TL below the 120 dB shadow cutoff). A raw diagnostic, not a scored quantity — Coverage Fidelity uses the directional false-shadow/false-light rates above.

Boundary Δshadow-boundary distance · km

Average symmetric distance between the model's and BELLHOP3D's shadow/lit boundary voxels, from a 6-connected grid-distance search (not exact Euclidean — an approximate, grid-quantized diagnostic). Reported when the full canonical grid is available; prefer this over a raw Hausdorff distance, which one outlier voxel can dominate.

Compositesecondary blend · points secondary
Etotal=0.60Efield+ 0.20Ecoverage+ 0.20Egeom

100×(1 − clamp(Etotal, 0, 1)), computed for every canonical model. If Geometry is unavailable for a model, its weight is re-normalized across the remaining dimensions and the result is marked provisional (shown as "· *"). This is a convenience summary, not the scientific result — Field, Coverage, and Geometry each have their own leader.

TL RMSEfull-field error · dB
1N i=1N (T^iTi)2

Root-mean-square gap between the model's transmission-loss field and the BELLHOP3D reference solver over every cell of the 101×49×31 grid, shadow zones included. A raw diagnostic — Field Fidelity uses the robust, insonified-only Core err instead.

Diff overlay

TL(model) − TL(BELLHOP3D) · centerline y = 12 km range–depth slice