Working with Files

Open in JupyterLite

DASCore reads and writes local pathlib.Path objects and remote UPath resources supported by each format. See Working with remote patches for cache policy.

Writing and reading

Write through Patch.io, then open the result as a spool:

from pathlib import Path

import dascore as dc

patch = dc.get_example_patch()
path = Path("output_file.h5")
patch.io.write(path, "DASDAE")
round_trip = dc.spool(path)[0]

Remote destinations use the same interface through universal-pathlib, whose UPath objects support URLs implemented by fsspec:

from upath import UPath

remote_path = UPath("memory://dascore/tutorial/output.pkl")
patch.io.write(remote_path, "pickle")
remote_patch = dc.spool(remote_path)[0]

Format support varies by backend: a formatter that requires a local seekable file may use the remote cache, while formats built on fsspec-compatible stores can operate directly.

Compression and chunking

DASDAE, PRODML, and NETCDF_CF writers take an encoding dict as xarray’s to_netcdf does: keys are variable names ("data" or a coordinate name; PRODML takes only "data" and "time"), values are options such as zlib, complevel, shuffle, chunksizes, and fletcher32. NETCDF_CF passes it to xarray; DASDAE and PRODML translate it for h5py as xarray’s h5netcdf engine does, clamping chunksizes to each array’s length. A write option the writer does not accept raises a ParameterError.

encoding = {"data": {"zlib": True, "complevel": 4, "chunksizes": (50, 1000)}}
patch.io.write(UPath("memory://dascore/tutorial/small.h5"), "DASDAE", encoding=encoding)
MemoryPath('/dascore/tutorial/small.h5', protocol='memory')

Zarr

The ZARR format writes a patch as a zarr store (a directory) laid out as xarray’s to_zarr writes it, so any xarray user can open it. DASCore reads stores other tools write this way when they hold one payload variable (data, or the only variable): a payload at the root is the one patch, otherwise each child group holding one is a patch. It needs zarr>=3, installed with pip install "dascore[extras]". file_version picks the zarr format ("3", the default, or "2"), and encoding is passed to xarray as is: chunks sets the chunk shape, shards groups chunks into fewer files (zarr format 3 only), and compressors sets the codecs; xarray raises a ValueError for other keys, such as chunksizes.

One patch is written at the store root, which xr.open_zarr opens. Several (including the pieces of a patch with holes) are written one group each, named as Spool.get_patch_names() names them, with “/” replaced by “_” and repeats suffixed __1, __2, … (as DASDAE does), and xr.open_datatree(path, engine="zarr") opens one node per patch; encoding applies to every group. Writing never merges patches; merge first (e.g. spool.chunk(time=None)) to store one array per contiguous block. Writing replaces a store already at the path, but not a directory or file which is not a store. A directory of stores works as a directory spool.

zarr_path = UPath("memory://dascore/tutorial/patch.zarr")
encoding = {"data": {"chunks": (50, 500), "shards": (100, 1000)}}
patch.io.write(zarr_path, "ZARR", encoding=encoding)
zarr_patch = dc.read(zarr_path)[0]

A store on object storage is a UPath carrying its storage options, and dc.read, dc.scan, and dc.spool take it as they take a local store. A scan reads only metadata and coordinates, and loading a patch reads only the chunks its selection touches. A remote directory of stores cannot be indexed.

store = UPath("s3://bucket/das.zarr", anon=True)
spool = dc.spool(store)
start = spool.get_contents()["time_min"].min()
first_second = spool.select(time=(start, start + dc.to_timedelta64(1)))[0]

Writing builds the new store beside the path and then moves it into place. On object storage that move is a copy, so a reader can see a partly replaced store while it runs.

Scanning metadata

dc.scan returns PatchSummary objects without loading patch arrays.

summary_index = 0
summary = dc.scan("examples://terra15_das_1_trimmed.hdf5")[summary_index]

print(summary.source_path)
print(summary.get_coord_summary("time").min)
/home/runner/work/dascore/dascore/.test_data_cache/0.0.0/terra15_das_1_trimmed.hdf5
2021-10-11T22:43:37.380047104

Scanning also reports source_format, source_version, and source_patch_key, which together identify the formatter and patch inside a multi-patch file.

Use the source path and the summary’s position to load the same patch through its spool. Scan results and file-spool rows have the same order, and the spool uses each row’s source_patch_key when materializing it. For compact directory listings, use scan or scan_to_df.

source_spool = dc.spool(
    summary.source_path,
    file_format=summary.source_format,
    file_version=summary.source_version,
)
loaded = source_spool[summary_index]

dc.scan_payloads(...) returns PatchMeta objects – patch metadata without the data – with each formatter’s full CoordManager, attrs, and dtype. With snap=False, formats that store per-sample coordinates expose their exact values:

payload = dc.scan_payloads(summary.source_path, snap=False)[0]
time = payload.coords.get_coord("time")

Payload scans still may read and retain large coordinate arrays, so reserve them for targeted files.

Patch.attrs contains non-coordinate metadata. Access coordinate bounds and steps from Patch.summary, PatchSummary.get_coord_summary(...), or the flattened summary returned by flat_dump().

Directory spools

dc.spool(directory) creates a lazy, indexed view of every supported file below that directory:

from dascore import examples

directory = examples.spool_to_directory(dc.get_example_spool("diverse_das"))
spool = (
    dc.spool(directory)
    .select(acquisition_key="DAS2.*")
    .select(time=(..., "2022-01-01"))
    .chunk(time=2, overlap=0.5)
)

Iteration is where patch arrays are loaded:

for patch in spool:
    processed = patch.detrend("time")

The hidden .dascore_index.sqlite3 stores file metadata and enables selections before patch data is loaded. Opening a new directory creates the index; call spool.update() after files change. See the spool tutorial and index note for path attributes and index lifecycle.

update() scans new or modified files, removes rows for deleted files, and leaves unchanged rows alone. File size and modification time identify candidates for rescanning.

An optional .inventory companion describes the observing system and is not scanned as data.

The index and inventory are companions to the archive rather than data sources themselves, so both use reserved hidden names and are excluded from formatter discovery.

External formats

Patch.io also converts patches to structures used by Pandas, Xarray, and ObsPy. See the external conversion recipe.