Selection functions from Balrog (DECam surveys)#
This technote covers how the streamobs selection-function products are derived for the DECam surveys — DES Y6 Gold and DELVE DR3 Gold — from Balrog synthetic-source-injection (SSI) catalogues, and how to re-run the derivation on a machine that holds the injections.
It is the DECam counterpart to Selection-Function Derivation Methodology, which
defines the products, the delta_mag convention and the truth-anchoring recipe
survey-agnostically. Read that page first; this one only covers what is
specific to Balrog, and the one problem that is unique to DES/DELVE: the star
classifier cannot be evaluated on injected sources.
Everything here is driven by a single script,
scripts/des/balrog_selection_function.py, which carries a schema adapter per
survey (--survey delve, --survey des_y6). It streams the HDF5 in chunks,
imports nothing from streamobs, and depends only on numpy / h5py / pandas /
healpy (plus healsparse for .hsp depth maps and xgboost for the DES
surrogate) — so it runs unmodified on the machine that holds the catalogues and
emits only ~10 KB of CSVs plus optional depth maps.
Why Balrog needs special handling#
A Balrog catalogue is an injection catalogue: known sources are painted into real survey images, the images are re-reduced, and the recovered measurements are matched back to what went in. That gives exactly what a selection function needs — a truth table and a measured table for the same objects, in the real survey’s own noise, masking and crowding.
Two things make it awkward in practice:
The star/galaxy classifier the survey actually publishes cannot be recomputed on the injections. See The EXT_XGB problem below.
The truth star/galaxy label is not in the injection files. For DES it has to be joined in from the parent deep-field catalogue.
Both are solved, and both cost accuracy in ways that are quantified rather than hidden.
The EXT_XGB problem#
DES Y6 Gold’s best star/galaxy classifier is EXT_XGB, an XGBoost model over
six features (Bechtol et al. 2025, arXiv:2501.05739,
Table A.2): CONC, BDF_T, BDF_T_ERR, log10(BDF_S2N),
WAVG_SPREAD_MODEL_I, WAVG_SPREADERR_MODEL_I. Its integer classes come from
thresholding the continuous output (their Eq. A5):
EXT_XGB = (XGB_PRED < 0.865) + (XGB_PRED < 0.110)
+ (XGB_PRED < 0.045) + (XGB_PRED < 0.015)
so EXT_XGB = 0 is the purest stellar class and 0 <= EXT_XGB <= 1 is the
“complete” stellar selection. EXT_XGB = -9 means no classification.
It cannot be evaluated on injections. Three of the six features are never measured for injected sources. The Gold paper says so directly (App. A.2):
the XGBoost classification has not been implemented on simulation–injection–recovery tests due to the fact that the CONC parameter was not measured for the simulated object samples
and the Balrog paper (Anbajagane et al. 2025,
arXiv:2501.05683, Sec. 4.4) reaffirms it and
falls back to EXT_MASH — which that same paper shows carries >10% galaxy
contamination when used to select stars. Verified directly against both
released Balrog files: no EXT_XGB, no XGB_PRED, no SPREAD_MODEL, no
CONC.
Neither paper offers a workaround. streamobs needs one, because the quantity the injector consumes is
classification_eff(delta_mag) = P(EXT_XGB <= 1 | true star, detected)
for the cut that will actually be applied to the real catalogue.
What we do instead: surrogate + deconvolution#
The key observation is that Balrog has the truth label but not EXT_XGB, while
the real Y6 Gold catalogue has EXT_XGB but no truth label — and the two
overlap in feature space.
Step 1 — train a surrogate on the real catalogue.
scripts/des/build_des_xgb_surrogate.py trains a classifier S on real
des_y6_gold rows against the label EXT_XGB <= 1, using only features that
can be constructed identically on both sides. All are fitvd/SOF quantities,
and the Balrog files are the _sof products, so there is no cross-pipeline
mismatch:
feature |
real Y6 Gold |
Balrog matched file |
|---|---|---|
|
as named |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
as named |
|
Balrog’s (N,4) arrays are ordered g,r,i,z.
The GAp (Gaussian-aperture) concentration matters and is not
interchangeable with the aperture one: the paper’s CONC is a Gaussian-weighted
flux dilation ratio (their Sec. III.4 — flux through a PSF-sized Gaussian
versus one 5% wider, calibrated on the PSF model), not an aperture-magnitude
difference, and the paper explicitly flags GAp fluxes as the Balrog-appropriate
tool. Measured feature importance bears this out: conc_gap is the single most
important feature.
Step 2 — deconvolve. S is not EXT_XGB, so applying it to Balrog measures
eff_S = P(S=1 | star), not eff_X. The confusion between them is measurable
on the real catalogue, per magnitude bin:
a(m) = P(S=1 | EXT_XGB <= 1) b(m) = P(S=1 | EXT_XGB > 1)
eff_S = a*eff_X + b*(1 - eff_X) -> eff_X = (eff_S - b) / (a - b)
The reducer applies this inversion with --confusion, only where the classes
are well separated (a - b > 0.2; as a -> b the inversion amplifies noise
without bound). This removes the first-order bias rather than quoting the
surrogate’s disagreement rate as an irreducible systematic.
Without the deconvolution
classification_effdescribes the surrogate, notEXT_XGB. The reducer warns if--confusionis omitted.
How good is it, and where does it break#
Two checks establish this, both run before anything downstream was built.
CONC reconstruction (--precheck). The real catalogue carries CONC
itself, so the headroom is directly measurable — regress CONC on the
Balrog-available features and see what is recoverable:
i-band |
R² |
residual scatter |
|
|---|---|---|---|
18–19 |
0.851 |
0.0003 |
0.0060 |
20–21 |
0.912 |
0.0006 |
0.0063 |
22–23 |
0.878 |
0.0011 |
0.0045 |
23–24 |
0.689 |
0.0020 |
0.0043 |
24–25 |
0.393 |
0.0034 |
0.0049 |
CONC is essentially recoverable brightward of i ≈ 23 (residual ~20× below its
own spread) and degrades faintward, which is where the deconvolution does its
work.
Domain shift. A surrogate trained on one catalogue and applied to another is
only valid if the feature distributions agree. Comparing real versus Balrog at
matched magnitude flagged one feature hard: PSF_T is offset by ~1.4σ at
every magnitude (real median 1.035–1.045, Balrog 0.882–0.886). A constant
offset independent of magnitude is a pipeline difference, not a population one
(cf. the PSFEx/Piff mismatch noted in the Balrog paper, Sec. 3.4). PSF_T and
the derived psf_bdf_T are therefore excluded from the feature set, along
with conc_gap_r, which is redundant — fitvd fits one shape jointly across
bands with per-band fluxes, so GAP - BDF is band-independent by construction.
Dropping all three changes AUC, agreement and recall by less than 1e-4, so they carried no unique information. Final performance on 16.8M real rows:
AUC 0.9946, selection agreement 0.9802, recall 0.8813, precision 0.9491
a(m)falls from 1.000 at g ≈ 17 to 0.642 at g ≈ 27, whileb(m)stays at 0.002–0.011 — the surrogate is very pure but loses recall faintward, which is exactly what the deconvolution corrects. Left uncorrected,classification_effwould be ~17% low by g ≈ 25.
The residual systematic is that a and b are measured on the real
catalogue’s mixed star+galaxy population, while what is needed is the same
quantity conditioned on true stars. That dependence is second-order, and it is
checked externally rather than assumed away (see Validation below).
The truth star/galaxy label#
DES Y6 Balrog has no truth label column. It does inject real stars — the
Balrog paper is explicit that 1/3 of injections (weight_scheme == Y3_weight,
confirmed at 33.3% in the file) apply no star/galaxy selection at all, and that
Y6 uses real deep-field stars rather than Y3’s simulated ones. But the paper’s
own star definition is the colour-based classifier of the parent Y3
deep-field catalogue (Hartley & Choi et al. 2021), explicitly “classified as
stars using photometric colors (and not morphology)”.
That label is recoverable because the injected id is the deep-field
COADD_OBJECT_ID — verified to match 100% of a 5M-injection sample.
scripts/des/build_des_truth_labels.py downloads the four deep-field parquets
and builds the lookup. The KNN_CLASS encoding is not documented in the
released JSON sidecars (Oracle type info only), so it was established
empirically on bright deep-field objects where morphology is unambiguous:
|
meaning |
median |
frac point-like |
median |
|---|---|---|---|---|
2 |
star |
−0.004 |
0.924 |
−0.21 |
1 |
galaxy |
0.891 |
0.015 |
+0.50 |
0 |
unclassified |
1.324 |
0.149 |
+0.15 |
KNN_CLASS = 0 is not “galaxy” — only 21% of those objects have any NIR
photometry, i.e. the colour classifier had nothing to run on. They are 22.5% of
injections (every injection has in_VHS_footprint == 1, so this means “too
faint for usable NIR”, not “outside NIR coverage”), and the reducer drops them
from the numerator and the denominator of every curve: they are neither
known stars nor known galaxies, so they cannot inform either.
Do not substitute a morphological proxy. A
|bdf_T| < 0.02cut on the injected deep-field morphology selects 7.64% of injections but is only 36.8% pure — per 5M injections it picks up ~141k true stars against ~140k galaxies and ~102k unclassified. That is the difference between a stellar selection function and a mostly-galaxy one.
DELVE’s Balrog-of-the-Stars, by contrast, ships an explicit truth_STAR boolean
at roughly 50/50, so none of this applies there.
Which population defines the depth#
The truth anchor asks “at what magnitude does the truth-based scatter of
(obs − true) reach σ = 0.2171?” — but whose scatter? The reducer’s
--anchor-sample selects between the population that has passed the
reference-band S/N cut and the one that has not, and for DES the two answers are
a magnitude apart. Measured on the full catalogue:
band |
|
|
|---|---|---|
g |
−0.297 |
+0.395 |
r |
−0.280 |
−0.637 |
i |
−0.180 |
−0.207 |
z |
−0.124 |
−0.116 |
detected moves g deeper by 0.40 while moving r shallower by 0.64 — a 1.03 mag
swing between adjacent bands, which cannot be real depth. The cause is
structural: the S/N cut is applied in the reference band only, so the g
anchor is truncated by its own cut (self-referential — asking where the scatter
of objects selected for having small scatter reaches a threshold that selection
enforces), while the r/i/z anchor samples are selected on g S/N, distorting
each differently. nosnr gives smooth, same-sign shifts ordered correctly with
wavelength, and is the default.
The trade-off is that the sample photo-error curve then reaches σ = 0.2171
somewhat faintward of delta_mag = 0 rather than exactly at it. That is already
true of every truth-anchored streamobs release for a related reason — the maglim
is anchored to the sample curve while get_photo_error returns the catalog
curve — which is why they all carry skip_snr_maglim_check in the test registry.
Roman’s anchor sample does formally include the detection cut, but its own
methodology footnote records the difference there as ≲ 0.01 mag, i.e. Roman’s
anchor is numerically indistinguishable from the no-cut one. DES diverges only
because its S/N cut truncates much harder relative to its error inflation
(1.46 versus Roman’s ~2), leaving far less headroom between “reported error
passes S/N > 5” and “truth scatter reaches 0.217”. So nosnr is the choice that
matches Roman’s effective convention, even though detected matches its code.
Sanity check to keep: the per-band anchor shifts must be smooth and same-sign. A band whose shift departs sharply from its neighbours is signalling a photometry, selection or sentinel-masking problem in that band, not a real depth feature. That check is what caught this.
Balrog gotchas that cost real time#
DES ships two files, DELVE one.
fiducial_matched_measured_sof.hdf5(119,667,013 rows) is exactlyfiducial_injected_sof.hdf5(145,724,947) masked bydetected_0.5arcsec_match_radius, in the same row order — a lockstep walk, no hash join. The reducer precomputes the cumulative detected count so its chunk reader is idempotent, because the catalogue is traversed twice (zero-point audit, then products).Use
wide_tilename, nottilename, for anything per-tile in DES.tilenameis the deep-field source tile — only 181 values (COSMOS_C*,SN-*_C*). DELVE has onlytilenameand it is the wide tile.Undetected rows are 0.0, not NaN and not the sentinel, so they pass every
== 0flag test. The detection gate is what keeps them out.Failed measurements use
-9.999e9, coincident withmeas_bdf_flags != 0. Real-catalogue sentinels are heterogeneous and are not nulls:-9999(CONC),-99(WAVG_SPREAD_MODEL),+99/+37.5(MAG_AUTO/capped mags),EXT_XGB = -9. Value-based masking is mandatory;isnasilently misses them.meas_bdf_deblend_flagsis a broken constant in both surveys; AND-ing it zeroes the catalogue.meas_mask_flagsis identically 0. Use neither.Foreground masking is a property of the sky, not the object. In DES the flag exists only in the measured file, so 7.5% of detected rows carry it while undetected rows have no value at all. Cutting on it directly would drop masked objects from the numerator but keep them in the denominator. The reducer instead rebuilds it as a positional HEALPix mask (nside 4096; a pixel is masked if >50% of its detections are flagged, masking 8.9% of populated pixels) and applies it to detected and undetected rows alike. In DELVE
FLAGS_FOREGROUNDships for every row — and putting it in the numerator alone there drags the bright-end plateau from ~0.95 down to ~0.46.Blends: the Balrog paper flags injections within 1.5″ of a real object (<0.2% of the catalogue) and removes them from every analysis. The reducer does the same via
meas_BALROG_FLAG_BLEND.flags_bad_zpdoes not mean what its name suggests in DES. It is set on 11.1% of rows but shows no correlation with per-tile photometric offset (Pearson r = −0.002 on an 8% sample), and neither paper defines it. Do not use it as a QA cut.Extinction. DELVE’s
truth_FLUX_*is dereddened and needsA_bandadded back. DES needsbdf_mag_deredden + ext_mags— de-redden at the deep field, re-redden at the injection position. The deciding diagnostic is the per-tile spread of (obs − true), not its median: rawbdf_maghas the smallest median offset yet a 12× wider spread (0.0300 vs 0.0025), because extinction varies tile to tile. With the correct choice, 0 of 538 DES tiles deviate by more than 0.05 mag.Per-tile zero points differ wildly between the two surveys. DES is clean; DELVE V4 has ~22% of tiles off by >0.05 mag, some by −0.72 (close to 2.5·log₁₀2, suggesting a factor-of-two flux convention issue on the truth side for those tiles). Left uncorrected the DELVE truth anchor returns g ≈ 19.8 instead of ≈24.5.
--tile-zp rejectis the default and is the treatment DES itself applies to this class of artifact.Rejected tiles are a sample cut, not a footprint cut. The curves are
delta_mag-keyed and the imaging is homogeneous within a survey, so a curve measured on good-ZP tiles applies to the whole footprint. Ship the maglim map unmasked. Do not extrapolate this across surveys.
Running it#
DES Y6#
Three steps; the first two are one-offs.
PY=/astro/store/shiren/conda-envs/stream_team/envs/streamobs/bin/python
# 1. truth labels (downloads ~5.3 GB of deep-field parquet, emits ~36 MB)
$PY scripts/des/build_des_truth_labels.py --data-dir /path/to/y3_deep_fields
# 2. EXT_XGB surrogate + confusion table (reads the real des_y6_gold catalogue)
$PY scripts/des/build_des_xgb_surrogate.py --out artifacts/des_y6 --precheck
# 3. the reduction itself
$PY scripts/des/balrog_selection_function.py \
--survey des_y6 --band g \
--catalog /path/to/fiducial_injected_sof.hdf5 \
--measured /path/to/fiducial_matched_measured_sof.hdf5 \
--truth-labels /path/to/des_y6_deepfield_truth.parquet \
--surrogate artifacts/des_y6/des_y6_xgb_surrogate.json \
--confusion artifacts/des_y6/des_y6_xgb_confusion.csv \
--maglim-map g=data/surveys/des_yr6/des_y6_5_sig_maglim_band_g_nside_512.hsp \
r=data/surveys/des_yr6/des_y6_5_sig_maglim_band_r_nside_512.hsp \
i=data/surveys/des_yr6/des_y6_5_sig_maglim_band_i_nside_512.hsp \
z=data/surveys/des_yr6/des_y6_5_sig_maglim_band_z_nside_512.hsp \
--write-maglim --maglim-nside 128 \
--corrections scripts/des/des_photoerror_corrections.yaml \
--out ./des_y6_products --tag des_y6
The des_y6_5_sig_* maps supply the spatial structure only; their absolute
scale is truth-anchored from the injections, so a map delivered on any S/N
convention is corrected automatically rather than by assuming a
+2.5·log₁₀2 = 0.753 mag shift. (Those maps are genuinely 5σ — medians g 25.32 /
r 25.13 / i 24.58 / z 23.88 — despite a long-stale comment in the config that
quoted the 10σ number.)
DELVE DR3 Gold — runbook for the machine holding the injections#
DELVE’s products are not shipped with streamobs; derive them where the
catalogue lives. The DELVE schema needs no surrogate and no truth-label join —
truth_STAR and FLAGS_FOREGROUND ship for every row, and the classifier is
bdf_extended_class_dr3gold, vendored into the reducer and verified
bit-identical to desqr on 1.5M rows. It needs only BDF_T and BDF_S2N, both
of which Balrog measures, and it reuses the DES Y6 Gold interpolation nodes, so
DES and DELVE come out classified identically — which is what makes the two
releases comparable.
PY=/path/to/streamobs-env/bin/python # numpy, h5py, pandas, healpy, healsparse
D=/path/to/dr3_gold_maglim_maps # the eight DR3 HealSparse depth maps
$PY scripts/des/balrog_selection_function.py \
--survey delve --band g \
--catalog /path/to/BalrogOfTheStars_Catalog_V4.hdf5 \
--maglim-map g=$D/delve_dr32_g_maglim_wmean.hsp,$D/delve_dr311+dr312_g_maglim_Nov28th.hsp \
r=$D/delve_dr32_r_maglim_wmean.hsp,$D/delve_dr311+dr312_r_maglim_Nov28th.hsp \
i=$D/delve_dr32_i_maglim_wmean.hsp,$D/delve_dr311+dr312_i_maglim_Nov28th.hsp \
z=$D/delve_dr32_z_maglim_wmean.hsp,$D/delve_dr311+dr312_z_maglim_Nov28th.hsp \
--tile-zp reject --write-maglim --maglim-nside 512 \
--corrections scripts/des/delve_photoerror_corrections.yaml \
--out ./delve_dr3_gold_products --tag delve_dr3_gold --chunk 4000000
Run it from the repository root, or give an absolute path to the script — it
imports nothing from streamobs, so it does not need to be on sys.path.
Roughly 35 minutes and ~48 GB of RAM for all four bands on the full V4 catalogue, most of it spent reading the eight nside-16384 depth maps. Two bands is about 20 minutes and ~38 GB.
--maglim-nside 512 matches the des/yr6 depth grid, so the two DECam releases
are directly comparable pixel for pixel. Note that this argument only ever
degrades: asking for a resolution finer than the input map’s own nside is
silently a no-op, while the output filename still takes the value you asked for.
Both depth maps are required. The DR3.2 and DR3.1.1+3.1.2 map sets are
exactly disjoint — measured on a 1M-injection sample spread over the whole
catalogue, DR3.2 covers 42.2% and DR3.1.1+3.1.2 covers 57.6%, with 0.00%
overlap and a 99.76% union. They are complementary halves of the footprint, not
competing vintages. Passing only one silently drops ~half the injections, which
in testing produced an all-zero efficiency table. Comma-separate them; first
listed wins on overlap. They are HealSparse at nside 16384 and are queried
sparsely, never densified — --maglim-nside only sets the resolution of maps
written out.
Reference band is g for both releases; colour is carried by the per-band depth maps, not by separate per-band efficiency tables.
Check the audit JSON before shipping: the per-tile ZP spread, the truth-anchored m5 per band (within ~0.1 mag of the map median after the shift), the error-inflation factor (order 1–2, not 47) and the applied map shift.
Each run emits four photo-error curves, not two: the sample/catalog pair
measured on the detected population, and a _nocut pair measured without the
reference-band S/N cut. streamobs applies the first to the reference band and
the second to every other band, whose photometry is forced at the reference
band’s position and so is not conditioned on its own detection. The pairs are
identical brightward of the depth and diverge faintward, where the S/N cut
truncates the detected sample and its measured scatter turns over rather than
continuing to rise.
Copy back only the CSVs, the depth maps and the audit JSON — they total a few MB.
Validation#
Against the published integrated numbers. Bechtol et al. Table A.3 gives stellar efficiency/contamination over two magnitude ranges, which the derived curves must reproduce:
0 <= EXT_XGB <= 1→ 98.0% / 4.0% for 17.5 ≤ i ≤ 22.5 and 94.3% / 12.5% for 16.5 ≤ i ≤ 23.5;EXT_XGB = 0→ 92.1% / 1.0% and 79.3% / 1.5%. These constrain the integral of the curve; the paper publishes per-magnitude performance only graphically (their Fig. 3).Against an external deep truth sample — done, and it passes.
scripts/des/validate_des_xgb_with_splash.pymatches the real Y6 Gold catalogue at 0.5″ against SPLASH-SXDF (Mehta et al. 2018) in the same SXDF field DES used for their own faint-domain validation. On 151,580 SPLASH-classified matches the stellar completeness reproduces Table A.3:selection
range
this work
paper
EXT_XGB ≤ 117.5–22.5
0.988
0.980
EXT_XGB ≤ 116.5–23.5
0.964
0.943
EXT_XGB = 017.5–22.5
0.947
0.921
EXT_XGB = 016.5–23.5
0.811
0.793
This is the external check on the deconvolution’s residual assumption, and it is the quantity streamobs actually ships.
Contamination is not measurable from this comparison (0.223 vs the paper’s 0.040) and the discrepancy is a property of the truth sample. SPLASH’s
STAR_FLAGis pure but incomplete: flag = 1 sits in a tight point-source locus (median HSCFLUX_RADIUS3.07 vs 6.76 for flag = 0), yet the DES-selected objects it calls “galaxy” have medianFLUX_RADIUS3.351 with p16 = 3.082 — over half of them lie on the stellar locus. They are stars that failed SPLASH’s IRAC-dependent star criterion, not DES misclassifications. A pure-but-incomplete star label yields an unbiasedP(selected | star)and an inflatedP(not star | selected). The same limitation means this field cannot validate the galaxy-misclassification product; that needs a complete galaxy label, e.g. the HSC PDR3 concentrationi_psfflux_mag − i_cmodel_magDES used for their Fig. 3.DES against DELVE in
delta_magspace. The two footprints overlap over only 369 deg² (7.1% of DES) and that overlap is edge-dominated, so a sky-based cross-check is weak. Comparing the curves indelta_magneeds no sky overlap at all, and is precisely what thedelta_magkeying is for.
Bright-end plateau. In the Roman and LSST products the per-object flag cut sits in the numerator and pulls the bright-end plateau down to ~0.91 — a property of the adopted cut, not the instrument. DES Y6 Balrog does not behave that way: its flagged fraction is tiny (0.056% for
meas_flags/meas_bdf_flags, plus 0.2% blend-flagged), because injections are placed on a sparse 20″ hexagonal grid specifically to avoid injection-injection blending. Sodetection_effplateaus at ~1.0, and that is expected rather than a missing cut. Do not “fix” it by reaching for a harsher flag selection, and do not read the agreement with the legacy DES product’s 1.00 plateau as validation — that product reached it for a different reason.