streamobs.background.storage module#

Persistence layer for precomputed background CMD histogram grids.

class streamobs.background.storage.BackgroundStorage(base_path=None, survey_name='lsst', **kwargs)[source]#

Bases: object

Save and load precomputed color–magnitude diagram (CMD) histogram grids.

One compressed parquet file per (source_type, bands) combination, e.g. stars_gr.parquet. Inside that file there is one row per (maglim_b2, maglim_b1) grid point, where b1 = bands[0] (color band) and b2 = bands[1] (reference/magnitude band).

CMD counts are stored as density (counts per deg²). Multiply by the target pixel area to obtain expected object counts.

File format — one row per (maglim_b2, maglim_b1) pair:

maglim_b2 | maglim_b1
| color_edge_min | color_edge_max | n_color
| mag_edge_min   | mag_edge_max   | n_mag
| counts  (list of n_color × n_mag floats per deg², row-major)

Bin edges are derived from (edge_min, edge_max, n_bins) on load. Bin centers are not stored; compute them from edges when needed.

When reading for a specific (maglim_b2, maglim_b1) pair, pyarrow predicate pushdown is used so that only the relevant row groups are read from disk — the full file is never loaded into memory.

Parameters:
  • base_path (str, optional) – Root directory for resource files. Defaults to {package_root}/data/background/.

  • survey_name (str, optional) – Survey identifier (e.g. 'lsst'). Used in the file path.

Examples

>>> storage = BackgroundStorage(survey_name='lsst')
>>> path = storage.get_path('stars', ('g', 'r'))
>>> # -> .../data/background/lsst/stars_gr.parquet
exists(source_type: str, bands: tuple) → bool[source]#

Check whether the resource file for this (source_type, bands) combination exists on disk.

Parameters:
  • source_type (str) – 'stars' or 'galaxies'.

  • bands (tuple of str) – Band names, e.g. ('g', 'r').

Return type:

bool

get_path(source_type: str, bands: tuple) → str[source]#

Build the file path for a (source_type, bands) combination.

Parameters:
  • source_type (str) – 'stars' or 'galaxies'.

  • bands (tuple of str) – Band names in order, e.g. ('g', 'r').

Returns:

Absolute path, e.g. {base_path}/lsst/stars_gr.parquet.

Return type:

str

load_all(source_type: str, bands: tuple) → dict[source]#

Load the full CMD histogram grid from the parquet file.

Returns:

{(maglim_b2, maglim_b1): {'cmd_hist', 'color_edges', 'mag_edges'}} where b1 = bands[0], b2 = bands[1].

Return type:

dict

load_data(source_type: str, bands: tuple, maglim_b2: float, maglim_b1: float) → dict[source]#

Load the CMD histogram for a specific (maglim_b2, maglim_b1) pair.

Uses pyarrow predicate pushdown — only the relevant row groups are read from disk.

Parameters:
  • source_type (str) – 'stars' or 'galaxies'.

  • bands (tuple of str) – Band names, e.g. ('g', 'r').

  • maglim_b2 (float) – Magnitude limit for bands[1] (reference band).

  • maglim_b1 (float) – Magnitude limit for bands[0] (color band).

Returns:

{'cmd_hist', 'color_edges', 'mag_edges'}.

Return type:

dict

save_data(data: dict, source_type: str, bands: tuple, **kwargs)[source]#

Persist the full CMD histogram grid to a single parquet file.

Parameters:
  • data (dict) – Full grid keyed by (maglim_b2, maglim_b1), each value being a dict with keys cmd_hist (counts per deg²), color_edges, mag_edges. b1 = bands[0], b2 = bands[1].

  • source_type (str) – 'stars' or 'galaxies'.

  • bands (tuple of str) – Band names, e.g. ('g', 'r').

  • **kwargs –

    compressionstr, optional

    Parquet compression codec. Default 'zstd'.