cascade.data#
The home for Cascade pipeline building tools
- class cascade.data.ApplyModifier(dataset: Union[Dataset[T], IteratorDataset[T]], func: Callable[[T], Any], p: Optional[float] = None, seed: Optional[int] = None, *args: Any, **kwargs: Any)[source]#
-
Modifier that applies a function to given dataset’s items in each __getitem__ call.
Can be applied to Iterators too.
- __init__(dataset: Union[Dataset[T], IteratorDataset[T]], func: Callable[[T], Any], p: Optional[float] = None, seed: Optional[int] = None, *args: Any, **kwargs: Any) None[source]#
-
- Parameters:
-
-
dataset (Dataset) – A dataset to modify
-
func (Callable) – A function to be applied to every item of a dataset - each
__getitem__callsfuncon an item obtained from a previous dataset -
p (Optional[float], by default None) – The probability [0, 1] with which to apply
func -
seed (Optional[int], by default None) – Random seed is used when p is not None
-
Examples
>>> from cascade import data as cdd >>> ds = cdd.Wrapper([0, 1, 2, 3, 4]) >>> ds = cdd.ApplyModifier(ds, lambda x: x ** 2)
Now function will only be applied when items are retrieved
>>> assert [item for item in ds] == [0, 1, 4, 9, 16]
- class cascade.data.Assessor(id: Optional[str] = None, position: Optional[str] = None)[source]#
-
The container for the info on the people who were in charge of labeling process, their experience and other properties.
This is a dataclass, so any additional fields will not be recorded if added. If it needs to be extended, please create a new class instead.
- class cascade.data.BaseDataset(*args: Any, data_card: Optional[DataCard] = None, **kwargs: Any)[source]#
-
Base class of any object that constitutes a step in a data-pipeline
See also
- class cascade.data.BaseModifier(dataset: BaseDataset[T], *args: Any, **kwargs: Any)[source]#
-
Base class for Modifiers, mostly unifies metadata management
- __init__(dataset: BaseDataset[T], *args: Any, **kwargs: Any) None[source]#
-
Constructs a Modifier. Modifier represents a step in a pipeline - some data transformation
- Parameters:
-
dataset (BaseDataset[T]) – A dataset to modify
- from_meta(meta: List[Dict[Any, Any]]) None[source]#
-
Calls the same method as base class but does it cascade-like which allows to roll a list of meta on a pipeline
- Parameters:
-
meta (Meta) – Meta of a single object or a pipeline
- get_meta() List[Dict[Any, Any]][source]#
-
Overrides base method enabling cascade-like calls to previous datasets. The metadata of a pipeline that consist of several modifiers can be easily obtained with
get_metaof the last block.
- update_meta(meta)[source]#
-
Updates
_meta_prefix, which then updates dataset’s meta whenget_meta()is called- Parameters:
-
meta (Union[Meta, MetaBlock, Config]) – The object to update with
- Raises:
-
ValueError – If the list passed and it is not of the unit length
- class cascade.data.BruteforceCacher(dataset: BaseDataset[T], *args: Any, **kwargs: Any)[source]#
-
Special modifier that calls all previous pipeline in __init__ loading everything in memory.
Examples
>>> from cascade import data as cdd >>> ds = cdd.Wrapper([0 for _ in range(1000000)]) >>> ds = cdd.ApplyModifier(ds, lambda x: x + 1) >>> ds = cdd.ApplyModifier(ds, lambda x: x + 1) >>> ds = cdd.ApplyModifier(ds, lambda x: x + 1)
Cache heavy upstream part
>>> ds = cdd.BruteforceCacher(ds)
- class cascade.data.Composer(datasets: List[Dataset[Any]], *args: Any, **kwargs: Any)[source]#
-
Unifies two or more datasets element-wise.
Example
>>> from cascade import data as cdd >>> items = cdd.Wrapper([0, 1, 2, 3, 4]) >>> labels = cdd.Wrapper([1, 0, 0, 1, 1]) >>> ds = cdd.Composer((items, labels)) >>> assert ds[0] == (0, 1)
- __init__(datasets: List[Dataset[Any]], *args: Any, **kwargs: Any) None[source]#
-
- Parameters:
-
datasets (Iterable[Dataset]) – Datasets of the same length to be unified
- class cascade.data.Concatenator(datasets: List[Dataset[T]], *args: Any, **kwargs: Any)[source]#
-
Unifies several Datasets under one, calling them sequentially in the provided order.
Examples
>>> from cascade.data import Wrapper, Concatenator >>> ds_1 = Wrapper([0, 1, 2]) >>> ds_2 = Wrapper([2, 1, 0]) >>> ds = Concatenator((ds_1, ds_2)) >>> assert [item for item in ds] == [0, 1, 2, 2, 1, 0]
- __init__(datasets: List[Dataset[T]], *args: Any, **kwargs: Any) None[source]#
-
Creates concatenated dataset from the list of datasets provided
- class cascade.data.CyclicSampler(dataset: Dataset[T], num_samples: int, *args: Any, **kwargs: Any)[source]#
-
A Sampler that iterates
num_samplestimes through an input Dataset in cyclic mannerExample
>>> from cascade.data import CyclicSampler, Wrapper >>> ds = Wrapper([1,2,3]) >>> ds = CyclicSampler(ds, 7) >>> assert [item for item in ds] == [1, 2, 3, 1, 2, 3, 1]
- class cascade.data.DataCard(name: Optional[str] = None, desc: Optional[str] = None, source: Optional[str] = None, goal: Optional[str] = None, labeling_info: Optional[LabelingInfo] = None, size: Optional[Union[int, Tuple[int]]] = None, metrics: Optional[Dict[str, Any]] = None, schema: Optional[Dict[Any, Any]] = None, **kwargs: Any)[source]#
-
The container for the information on dataset. The set of fields here is general and can be extended by providing new keywords into __init__.
Example
>>> from cascade.data import DataCard, Assessor, LabelingInfo >>> person = Assessor(id=0, position="Assessor") >>> info = LabelingInfo(who=[person], process_desc="Labeling description") >>> dc = DataCard( ... name="Dataset", ... desc="Example dataset", ... source="Database", ... goal="Every dataset should have a goal", ... labeling_info=info, ... size=100, ... metrics={"quality": 100}, ... schema={"label": "value"}, ... custom_field="hello")
- __init__(name: Optional[str] = None, desc: Optional[str] = None, source: Optional[str] = None, goal: Optional[str] = None, labeling_info: Optional[LabelingInfo] = None, size: Optional[Union[int, Tuple[int]]] = None, metrics: Optional[Dict[str, Any]] = None, schema: Optional[Dict[Any, Any]] = None, **kwargs: Any) None[source]#
-
- Parameters:
-
-
name (Optional[str]) – The name of dataset
-
desc (Optional[str]) – Short description
-
source (Optional[str]) – The source of data. Can be URL or textual description of source
-
goal (Optional[str]) – The datasets have a goal - what should be achieved using this data?
-
labeling_info (Optional[LabelingInfo]) – The instance of dataclass describing labeling process placed here
-
size (Union[int, Tuple[int], None]) – This can usually be done automatically - number of items or shape of the table.
-
metrics (Optional[Dict[str, Any]]) – Dictionary with names and values of metrics. Any quality metrics can be included
-
schema (Optional[Dict[Any, Any]]) – Schema dictionary describing table datasets, their columns, data types, possible values, etc.
-
- class cascade.data.Dataset(*args: Any, data_card: Optional[DataCard] = None, **kwargs: Any)[source]#
-
An abstract class to represent a dataset with __len__ method present. Inheritance of this class should mean the presence of length.
If your dataset does not have length defined you can use IteratorModifier
See also
- class cascade.data.Filter(dataset: Dataset, filter_fn: Callable, *args: Any, **kwargs: Any)[source]#
-
Filter for Datasets with length. Uses a function to create a mask of items that will be stored once and applied for each access.
Example
Here we select only even numbers from a dataset
>>> from cascade.data import Filter, Wrapper >>> ds = Wrapper([0, 1, 2, 3]) >>> ds = Filter(ds, lambda x: x % 2 == 0) >>> list(ds) [0, 2]
- __init__(dataset: Dataset, filter_fn: Callable, *args: Any, **kwargs: Any) None[source]#
-
Filter a dataset using a filter function. Does not accumulate items in memory, will store only an index mask.
- Parameters:
-
-
dataset (Dataset) – A dataset to filter
-
filter_fn (Callable) – A function to be applied to every item of a dataset - should return bool. Will be called on every item on
__init__.
-
- Raises:
-
RuntimeError – If
filter_fnraises an exception
- class cascade.data.FolderDataset(root: str, *args: Any, **kwargs: Any)[source]#
-
Basic “folder of files” dataset. Accepts root folder in which considers all files. Is abstract - getitem is not defined, since it is specific for each file type.
See also
cascade.utils.FolderImageDataset
- class cascade.data.IteratorDataset(*args: Any, data_card: Optional[DataCard] = None, **kwargs: Any)[source]#
-
An abstract class to represent a dataset as an iterable object
- class cascade.data.IteratorFilter(dataset: IteratorDataset, filter_fn: Callable, *args: Any, **kwargs: Any)[source]#
-
Filter for datasets without length
Does not filter on init, returns only items that pass the filter
Example
Here we select only even numbers from a dataset
>>> from cascade.data import IteratorFilter, IteratorWrapper >>> ds = IteratorWrapper([0, 1, 2, 3]) >>> ds = IteratorFilter(ds, lambda x: x % 2 == 0) >>> list(ds) [0, 2]
- __init__(dataset: IteratorDataset, filter_fn: Callable, *args: Any, **kwargs: Any) None[source]#
-
Constructs a Modifier. Modifier represents a step in a pipeline - some data transformation
- Parameters:
-
dataset (BaseDataset[T]) – A dataset to modify
- class cascade.data.IteratorModifier(dataset: IteratorDataset[T], *args: Any, **kwargs: Any)[source]#
-
The Modifier for Iterator datasets
- __init__(dataset: IteratorDataset[T], *args: Any, **kwargs: Any) None[source]#
-
Constructs a Modifier. Modifier represents a step in a pipeline - some data transformation
- Parameters:
-
dataset (BaseDataset[T]) – A dataset to modify
- class cascade.data.IteratorWrapper(data: Iterable[T], *args: Any, **kwargs: Any)[source]#
-
Wraps IteratorDataset around any Iterable. Does not have map-like interface.
- class cascade.data.LabelingInfo(who: Optional[List[Assessor]] = None, process_desc: Optional[str] = None, docs: Optional[str] = None)[source]#
-
The container for the information on labeling process, people involved, description of the process, documentation links.
This is a dataclass, so any additional fields will not be recorded if added. If it needs to be extended, please create a new class instead.
- class cascade.data.Modifier(dataset: BaseDataset[T], *args: Any, **kwargs: Any)[source]#
-
Basic pipeline building block in Cascade. Every block which is not a data source should be a successor of Sampler or Modifier.
This structure enables having a data pipeline which consists of uniform blocks each of them has a reference to the previous one in its
_datasetfieldBasically Modifier defines an arbitrary transformation on every dataset’s item that is applied in a lazy manner on each
__getitem__call.Applies no transformation if
__getitem__is not overriddenDoes not change the length of a dataset. See Sampler for this functionality
Example
from cascade.data import Modifier, Wrapper class PowModifier(Modifier): def __init__(self, ds, p, *args, **kwargs): super().__init__(ds, *args, **kwargs) self.p = p def get(self, index): number = self._dataset[index] return number ** self.p ds = Wrapper([0, 1, 2, 3, 4]) ds = PowModifier(ds, 2) assert list(ds) == [0, 1, 4, 9, 16]
- class cascade.data.RandomSampler(dataset: Dataset[T], num_samples: Optional[int] = None, *args: Any, **kwargs: Any)[source]#
-
Shuffles a dataset
- class cascade.data.RangeSampler(dataset: Dataset[T], start: Optional[int] = None, stop: Optional[int] = None, step: int = 1, *args: Any, **kwargs: Any)[source]#
-
Implements Python range as a Dataset
Example
>>> from cascade.data import RangeSampler, Wrapper >>> ds = Wrapper([1, 2, 3, 4, 5]) >>> # Define start, stop and step exactly as in range() >>> sampler = RangeSampler(ds, 1, 5, 2) >>> for item in sampler: ... print(item) ... 2 4 >>> ds = Wrapper([1, 2, 3, 4, 5]) >>> sampler = RangeSampler(ds, 3) >>> for item in sampler: ... print(item) ... 1 2 3
- class cascade.data.Sampler(dataset: Dataset[T], num_samples: int, *args: Any, **kwargs: Any)[source]#
-
Defines certain sampling over a Dataset.
Its distinctive feature is that it changes the number of items in dataset.
Can be used to build a batch sampler, random sampler, etc.
- class cascade.data.SchemaModifier(dataset: Dataset, *args: Any, **kwargs: Any)[source]#
-
Data validation modifier
When
self._datasetis called and has self.in_schema defined, wrapsself._datasetinto validator, which is anotherModifierthat checks the output of__getitem__of the dataset that was wrapped.- In the end it will look like this:
-
- If
in_schemais not None: -
dataset = SchemaModifier(ValidationWrapper(dataset)) - If
in_schemais None: -
dataset = SchemaModifier(dataset)
- If
How to use it: 1. Define pydantic schema of input
from typing import List, Tuple import pydantic class AnnotImage(pydantic.BaseModel): image: List[List[List[float]]] segments: List[List[int]] bboxes: List[Tuple[int, int, int, int]]
-
Use schema as
in_schema
from cascade.data import SchemaModifier class ImageModifier(SchemaModifier): in_schema = AnnotImage
3. Create a regular
Modifierby subclassing ImageModifier.class IDoNothing(ImageModifier): def get(self, idx): item = self._dataset[idx] return item
4. That’s all. Schema check will be held automatically every time
self._dataset[idx]is accessed. If it is not likeAnnotImage, cascade.data.ValidationError will be raised.- __init__(dataset: Dataset, *args: Any, **kwargs: Any) None[source]#
-
Constructs a Modifier. Modifier represents a step in a pipeline - some data transformation
- Parameters:
-
dataset (BaseDataset[T]) – A dataset to modify
- class cascade.data.SimpleDataloader(data: Sequence[T], batch_size: int = 1)[source]#
-
Simple batch builder - given a sequence and a size of batch breaks it in the subsequences
>>> from cascade.data import SimpleDataloader >>> dl = SimpleDataloader([0, 1, 2], 2) >>> [item for item in dl] [[0, 1], [2]]
- exception cascade.data.ValidationError(message: Optional[str] = None, error_index: Optional[int] = None)[source]#
-
Base class to raise if data validation failed
Can provide additional information about the fail
- class cascade.data.Wrapper(obj: Sequence[T], *args: Any, **kwargs: Any)[source]#
-
Wraps Dataset around any list-like object
- get(index: Any) T[source]#
-
Return an item from wrapped container
- Parameters:
-
index (Any) –
- Returns:
-
Item from a container
- Return type:
-
T
Example
>>> from cascade.data import Wrapper >>> ds = Wrapper([1, 2, 3]) >>> print(ds.get(0)) 1
- get_meta() List[Dict[Any, Any]][source]#
-
Wrapper’s meta adds obj_type field to the default meta of a Modifier
- Return type:
-
Meta
Example
>>> from cascade.data import Wrapper >>> ds = Wrapper([1, 2, 3]) >>> print(ds.get_meta()) [{'name': 'cascade.data.dataset.Wrapper', ..., 'len': 3, 'obj_type': "<class 'list'>"}]
- cascade.data.dataset(f: Callable[[...], Any], do_validate_in: bool = True) Callable[[...], FunctionDataset][source]#
-
Thin wrapper to turn any function into a Cascade’s Dataset. Use this if the function is the data source
Will return FunctionDataset object. To get results of the execution use
dataset.resultfield- Parameters:
-
f (Callable[..., Any]) – Function that produces data
- Returns:
-
Call this to get a dataset
- Return type:
-
Callable[…, FunctionDataset]
Example
from cascade.data import dataset from cascade.data.functions import FunctionDataset @dataset def read_data(): return [0, 1, 2] x = read_data() assert isinstance(x, FunctionDataset) assert x.result == [0, 1, 2]
- cascade.data.modifier(f: Callable[[...], Any], do_validate_in: bool = True) Callable[[...], FunctionModifier][source]#
-
Thin wrapper to turn any function into Cascade’s Modifier Pass the returning value of a function that was previosly wrapped dataset or modifier. Will replace any dataset with
dataset.resultautomatically if the function argument isFunctionDataset.- Parameters:
-
f (Callable[..., Any]) – Function that modifies data
- Returns:
-
Call this to get a modifier
- Return type:
-
Callable[…, FunctionModifier]
Example
from cascade.data import dataset, modifier from cascade.data.functions import FunctionModifier @dataset def read_data(): return [0, 1, 2] @modifier def mul(inp_data, n): return list(x * n for x in inp_data) x = read_data() x = mul(x, 3) assert isinstance(x, FunctionModifier) assert x.result == [0, 3, 6]
- cascade.data.split(ds: Dataset[T], frac: Optional[float] = 0.5, num: Optional[int] = None) Tuple[RangeSampler[T], RangeSampler[T]][source]#
-
Splits dataset into two cascade.data.RangeSampler
- Parameters:
Example
>>> from cascade import data as cdd >>> ds = cdd.Wrapper([0, 1, 2, 3, 4])
>>> ds1, ds2 = cdd.split(ds) >>> print([item for item in ds1]) [0, 1]
>>> print([item for item in ds2]) [2, 3, 4]
>>> ds1, ds2 = cdd.split(ds, 0.6)
>>> print([item for item in ds1]) [0, 1, 2] >>> print([item for item in ds2]) [3, 4]
>>> ds1, ds2 = cdd.split(ds, num=4)
>>> print([item for item in ds1]) [0, 1, 2, 3] >>> print([item for item in ds2]) [4]
- cascade.data.validate_in(f: Callable[[...], Any]) Callable[[...], Any][source]#
-
Data validation decorator for callables. In each call validates only the input schema using type annotations if present. Does not check return value.
- Parameters:
-
f (Callable[[Any], Any]) – Function to wrap
- Returns:
-
Decorated function
- Return type:
-
Callable[[Any], Any]
Example
from cascade.data import validate_in @validate_in def repeat(a: str, b: int): return a * b repeat("a", 2)
Will raise ValidationError:
repeat(2, 2)