Extending DASCore

Namespaces let external packages expose methods such as patch.my_plugin.method() without subclassing DASCore objects or adding optional dependencies to core. File formats use FiberIO, not namespaces.

Host Base class Entry-point group Registry
Patch PatchNameSpace dascore.patch_namespace patch.csv
Spool SpoolNameSpace dascore.spool_namespace spool.csv
Inventory InventoryNameSpace dascore.inventory_namespace inventory.csv
AnnotationSet AnnotationNameSpace dascore.annotation_namespace annotation.csv

The base classes live in dascore.utils.namespace. A method receives the host object, so it behaves like a normal host method.

Local namespaces

For notebooks, tests, or private code, define a named subclass and import it before access:

import dascore as dc
from dascore.constants import PatchType
from dascore.utils.namespace import PatchNameSpace


class MyPatchNamespace(PatchNameSpace):
    name = "my_ext"

    @dc.patch_function()
    def peak_to_peak(patch: PatchType) -> float:
        return patch.data.max() - patch.data.min()


patch = dc.get_example_patch()
value = patch.my_ext.peak_to_peak()

The name must be a public Python identifier. Defining a duplicate host/name pair warns and replaces the earlier class.

Plugin packages

Reusable extensions should register an entry point. Given the class above in dascore_extra.patch_namespace:

[project.entry-points."dascore.patch_namespace"]
my_ext = "dascore_extra.patch_namespace:MyPatchNamespace"

DASCore loads it on first access. Spool, Inventory, and AnnotationSet use the same pattern with the base class and group from the table.

After publishing the package, add it to the matching file under dascore/plugin_registry/ so missing-plugin errors can tell users what to install:

package_name,package_url,namespace
dascore-extra,https://github.com/yourorg/dascore-extra,my_ext

DASCore’s own patch.io, inventory.io, and annotation_set.io follow this model.

Patch processors

A PatchProcessor subclass is a whole operation, and its fields are the parameters:

import dascore as dc
from dascore.constants import PatchType


class Scale(dc.PatchProcessor):
    """Multiply the data by a factor."""

    factor: float = 2.0

    @staticmethod
    def scale(patch: PatchType, /, factor: float = 2.0) -> PatchType:
        """Multiply the data by a factor."""
        return Scale(factor=factor).run(patch)

    def kernel(self, data):
        return data * self.factor


patch = dc.get_example_patch()
assert Scale.scale(patch, factor=3).equals(Scale(3)(patch))

The staticmethod is the patch function: what the registry holds and what a stored history resolves to. A class outside DASCore declares it on the class itself, since nothing outside DASCore can write into Patch. One of DASCore’s own instead writes a real method in the body of Patch – so that a type checker, an IDE and help read the whole class, which a synthesized signature does not allow:

class Patch(NamespaceOwner, PatchMeta):
    def scale(self, factor: float = 2.0) -> Self:
        """Multiply the data by a factor."""
        return Scale(factor=factor).run(self)

Either way, every field must appear among the method’s parameters, with the same default, and a field the class refuses by position must be keyword-only, or the class is refused: the two are one declaration of the same parameters. The method may take more than the class stores, since a body may resolve something before building the instance. -> Self is what run promises for one of DASCore’s own: metadata comes back metadata, and a Patch subclass comes back that subclass. The operation is documented once, with the class, and the method carries that docstring.

Calling a processor checks the patch, runs it, and records history and lineage ids, exactly as a decorated patch function does. Its operation id digests the name, __version__ and validated fields, so Scale(3) and Scale(3.0) are one operation.

An operation whose metadata changes, or whose kernel needs numbers read from the metadata, splits the work so that no array is touched until the kernel runs:

  • get_metadata(meta) gets a PatchMeta – the coords, attrs, dtype and backend of the data, without the data – and returns two things: the result’s metadata, built through the coord manager and meta.new, and what the kernel needs as ints, floats, bools, slices, None, Ellipsis, sequences of those, or numeric arrays. Deriving the one usually computes the other, so they come back together.
  • kernel(data, **plan) computes the array; it never sees a patch.
  • numpy_kernel(data, **plan) may be supplied alongside or instead of kernel for NumPy, SciPy or Numba implementations. NumPy inputs prefer it. Other backends use kernel if available, otherwise run on NumPy copies with a NumpyFallbackWarning and convert the result back to the input backend.
  • reconcile(result, out) can use the kernel result to finalize metadata. Return a PatchMeta to attach the result array, or call out.to_patch(data) to return a patch with value-dependent data and coordinates.
class RemoveMean(dc.PatchProcessor):
    """Remove the mean along a dimension."""

    dim: str = "time"

    @staticmethod
    def remove_mean(patch: PatchType, /, dim: str = "time") -> PatchType:
        """Remove the mean along a dimension."""
        return RemoveMean(dim=dim).run(patch)

    def get_metadata(self, meta):
        return meta, {"axis": meta.get_axis(self.dim)}

    def kernel(self, data, *, axis):
        return data - data.mean(axis=axis, keepdims=True)


out = RemoveMean(dim="distance")(patch)

A class which writes no kernel changes metadata alone: applied to a patch it hands the data straight through, and it runs on a patch’s metadata as readily as on the patch, so its patch function is annotated PatchMetaType rather than PatchType. One of DASCore’s own has its method written in PatchMeta so that metadata carries it, and Patch inherits it; the data-less contract is one DASCore holds only itself to, so a class defined outside DASCore is a method of neither.

Class variables declare the rest: required_dims, required_coords, and required_attrs for what a patch must hold, data_type for the result’s, history for how a call is recorded, and __version__, bumped when the same fields mean a different result. name is the method’s name and the registry’s, and name = None registers nothing and expects no method. Transpose and TileApply show get_metadata changing a patch’s shape.

Array backends

Patch functions receive the original data backend. Use its array API namespace instead of NumPy when the function should support multiple backends:

import dascore as dc
from dascore.constants import PatchType
from dascore.utils.array_api import array_namespace


@dc.patch_function()
def double(patch: PatchType) -> PatchType:
    xp = array_namespace(patch.data)
    return patch.new(data=xp.multiply(patch.data, 2))

Plain NumPy-based functions may raise or return a NumPy-backed patch for other arrays. Operators and ufuncs stay in the array namespace when possible; otherwise they convert through NumPy and emit NumpyFallbackWarning.

NumPy functions such as np.mean(patch) and reductions over excluded dtypes use this fallback. NumPy functions which map to a patch method, such as np.transpose, run that method instead and stay in the array namespace. __array__-only objects report numpy as their backend.

A backend package can give a processor its own kernel without changing the operation. The kernel takes the processor, the data, and whatever the class’s get_metadata plans (abs plans nothing; normalize plans axis, and window when one is given):

import numpy as np

from dascore.proc.basic import Abs
from dascore.core.processor import register_kernel


@register_kernel(Abs, "mybackend")
def _abs_mybackend(processor, data):
    return np.abs(data)

The data backend selects the kernel. A registered kernel takes precedence over the same class’s numpy_kernel and kernel; a subclass’s own implementation takes precedence over its parent’s registrations.

The operation may also be named by its registry tag – bare for one of DASCore’s own, package:name otherwise – and one kernel may serve several backends:

@register_kernel("abs", ("mybackend", "otherbackend"))
def _abs_shared(processor, data):
    return np.abs(data)

Kernels are registered when their module is imported. To have DASCore import it, rather than asking users to, declare the module as a dascore.kernels entry point for each backend it serves, named exactly as backend_name spells that backend. DASCore imports it the first time data of that backend meet a processor:

[project.entry-points."dascore.kernels"]
mybackend = "mypackage.kernels"
otherbackend = "mypackage.kernels"

See Contributing, Testing, and Documentation for package conventions.