The Atlas · v2.0 · 2026-05-03 · horizontal coverage

The Atlas of Public Astronomical Archives

An atlas of public astronomical archives — laid out by wavelength. A single electromagnetic spectrum runs across the top. Below it, every mission is a horizontal bar that spans only the frequencies it actually covers. Bar saturation tracks how much of the data has been individually studied.

"Most of astronomy's data is processed but unread. This map shows where the unread bands live on the spectrum."

Pan the spectrum
Click any bar · hover for layers

Confidence

High · proxy or published
Med · peer-reading estimate
Low · provisional

EM Bands

Click any archive bar above to see full layer breakdown and a link to the data archive.

How to read this atlas

Public astronomy archives are huge and growing. But "stored" is not the same as "studied." The Atlas separates the two — laid out by where, on the electromagnetic spectrum, each archive actually lives.

The top bar is the electromagnetic spectrum, in true log10(frequency / Hz). It runs from radio at the left to gamma-ray at the right, with the eight band labels embedded inside. Frequency ticks above and wavelength callouts below give two ways to anchor.

Beneath the spectrum, every archive is a horizontal bar drawn over the exact frequency range it observes. Wide bars are wide-band facilities (Gaia, ALMA, Spitzer, Suzaku); narrow bars are line-tracking or single-window instruments (NVSS, EHT, Chang'e-3 LUT). Bars are grouped into rows by band and packed greedily so they don't overlap.

Within each bar, the saturated top fraction is how much of the archive's data has been individually mined in the literature; the lighter bottom fraction is the still-stored reservoir below. Hover any bar for the layer summary; click any bar to expand the full breakdown and open the archive's portal.

Methodology — the vector mining score

Each archive is broken into its meaningful data layers (Gaia: astrometry / RVS / BP-RP photometry / BP-RP spectra / epoch variability / non-single stars). Each layer carries two numbers in [0,1]:

  • processing — how mature the data product is. Raw single-frame = low; calibrated catalog = high.
  • mined — how much of the layer has been individually studied in the literature. Bulk catalog citations alone do not lift this number; we want per-source or per-window papers.

The bar's top is the mined value of the most-mined layer. The bar's bottom is the mined value of the least-mined layer with processing ≥ 0.5 — i.e. the data is ready to use; readers haven't gotten there. The gap and the reservoir below are the editorial point.

Source proxies, in priority order: SIMBAD/NED references-per-source for catalog layers; ADS dedicated-paper counts for instrument-specific and time-domain layers; and a curator override where the proxies lie. Each mined value carries a confidence flag.

High — defensible proxy or published count Med — survives a peer reading, not a deep audit Low — provisional, expert override needed

References & data sources