lightning/pytorch_lightning/utilities/data.py

# Copyright The PyTorch Lightning team.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
#     http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

from typing import Any, Iterable, Mapping, Union

import torch
from torch.utils.data import DataLoader, IterableDataset

from pytorch_lightning.utilities import rank_zero_warn

BType = Union[torch.Tensor, str, Mapping[Any, "BType"], Iterable["BType"]]


def extract_batch_size(batch: BType) -> int:
    """
    Recursively unpack a batch to find a torch.Tensor.

    Returns:
        ``len(tensor)`` when found, or ``1`` when it hits an empty or non iterable.
    """
    if isinstance(batch, torch.Tensor):
        return batch.size(0)
    if isinstance(batch, str):
        return len(batch)
    if isinstance(batch, dict):
        sample = next(iter(batch.values()), 1)
        return extract_batch_size(sample)
    if isinstance(batch, Iterable):
        sample = next(iter(batch), 1)
        return extract_batch_size(sample)

    return 1


def has_iterable_dataset(dataloader: DataLoader) -> bool:
    return hasattr(dataloader, "dataset") and isinstance(dataloader.dataset, IterableDataset)


def has_len(dataloader: DataLoader) -> bool:
    """
    Checks if a given Dataloader has ``__len__`` method implemented i.e. if
    it is a finite dataloader or infinite dataloader.

    Raises:
        ValueError:
            If the length of Dataloader is 0, as it requires at least one batch
    """

    try:
        # try getting the length
        if len(dataloader) == 0:
            raise ValueError("`Dataloader` returned 0 length. Please make sure that it returns at least 1 batch")
        has_len = True
    except TypeError:
        has_len = False
    except NotImplementedError:  # e.g. raised by torchtext if a batch_size_fn is used
        has_len = False

    if has_len and has_iterable_dataset(dataloader):
        rank_zero_warn(
            "Your `IterableDataset` has `__len__` defined."
            " In combination with multi-process data loading (when num_workers > 1),"
            " `__len__` could be inaccurate if each worker is not configured independently"
            " to avoid having duplicate data."
        )
    return has_len


def get_len(dataloader: DataLoader) -> Union[int, float]:
    """Return the length of the given DataLoader. If ``__len__`` method is not implemented, return float('inf')."""

    if has_len(dataloader):
        return len(dataloader)

    return float("inf")
Follow up of #2892 (#3202) * Follow up of #2892 * typo * iterabledataset 2020-08-27 19:28:29 +00:00			`# Copyright The PyTorch Lightning team.`
			`#`
			`# Licensed under the Apache License, Version 2.0 (the "License");`
			`# you may not use this file except in compliance with the License.`
			`# You may obtain a copy of the License at`
			`#`
			`# http://www.apache.org/licenses/LICENSE-2.0`
			`#`
			`# Unless required by applicable law or agreed to in writing, software`
			`# distributed under the License is distributed on an "AS IS" BASIS,`
			`# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.`
			`# See the License for the specific language governing permissions and`
			`# limitations under the License.`
Revert "Remove limitation of batch scaler (#4006)" (#4040) This reverts commit 7e756ca11f2d3938ef73010fe63aa45d699bf0bc. 2020-10-10 01:03:23 +00:00
Expose `extract_batch_size` method and add corresponding tests. (#8357) * expose extract_batch and make public * first pass * early return * add changelog * move to utilities/data.py * add test_data.py * tests are passing * precommit hook * address pep8 failure Co-authored-by: Carlos Mocholi <carlossmocholi@gmail.com> 2021-07-13 11:35:10 +00:00			`from typing import Any, Iterable, Mapping, Union`
Fix isort failures in utilities (#5530) Remove from skipped module in pyproject.toml and fix failures on: - pytorch_lightning/utilities/*.py 2021-01-15 18:57:40 +00:00
Expose `extract_batch_size` method and add corresponding tests. (#8357) * expose extract_batch and make public * first pass * early return * add changelog * move to utilities/data.py * add test_data.py * tests are passing * precommit hook * address pep8 failure Co-authored-by: Carlos Mocholi <carlossmocholi@gmail.com> 2021-07-13 11:35:10 +00:00			`import torch`
Follow up of #2892 (#3202) * Follow up of #2892 * typo * iterabledataset 2020-08-27 19:28:29 +00:00			`from torch.utils.data import DataLoader, IterableDataset`

			`from pytorch_lightning.utilities import rank_zero_warn`

Replace `yapf` with `black` (#7783) Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> 2021-07-26 11:37:35 +00:00			`BType = Union[torch.Tensor, str, Mapping[Any, "BType"], Iterable["BType"]]`
Expose `extract_batch_size` method and add corresponding tests. (#8357) * expose extract_batch and make public * first pass * early return * add changelog * move to utilities/data.py * add test_data.py * tests are passing * precommit hook * address pep8 failure Co-authored-by: Carlos Mocholi <carlossmocholi@gmail.com> 2021-07-13 11:35:10 +00:00

			`def extract_batch_size(batch: BType) -> int:`
			`"""`
			`Recursively unpack a batch to find a torch.Tensor.`

			`Returns:`
			``len(tensor)`` when found, or ``1`` when it hits an empty or non iterable.
			`"""`
			`if isinstance(batch, torch.Tensor):`
			`return batch.size(0)`
			`if isinstance(batch, str):`
			`return len(batch)`
			`if isinstance(batch, dict):`
			`sample = next(iter(batch.values()), 1)`
			`return extract_batch_size(sample)`
			`if isinstance(batch, Iterable):`
			`sample = next(iter(batch), 1)`
			`return extract_batch_size(sample)`

			`return 1`

Follow up of #2892 (#3202) * Follow up of #2892 * typo * iterabledataset 2020-08-27 19:28:29 +00:00
Add tests for functions in utilities/data.py (#8785) * Add tests for utilities/data.py: test_has_iterable_dataset, test_has_len, test_get_len 2021-08-10 06:39:00 +00:00			`def has_iterable_dataset(dataloader: DataLoader) -> bool:`
Replace `yapf` with `black` (#7783) Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> 2021-07-26 11:37:35 +00:00			`return hasattr(dataloader, "dataset") and isinstance(dataloader.dataset, IterableDataset)`
Follow up of #2892 (#3202) * Follow up of #2892 * typo * iterabledataset 2020-08-27 19:28:29 +00:00

			`def has_len(dataloader: DataLoader) -> bool:`
document exceptions in utilities (#8122) Co-authored-by: ananthsub <ananth.subramaniam@gmail.com> Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> 2021-06-25 13:41:45 +00:00			`"""`
			Checks if a given Dataloader has ``__len__`` method implemented i.e. if
			`it is a finite dataloader or infinite dataloader.`

			`Raises:`
			`ValueError:`
			`If the length of Dataloader is 0, as it requires at least one batch`
			`"""`
Follow up of #2892 (#3202) * Follow up of #2892 * typo * iterabledataset 2020-08-27 19:28:29 +00:00
			`try:`
			`# try getting the length`
			`if len(dataloader) == 0:`
Replace `yapf` with `black` (#7783) Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> 2021-07-26 11:37:35 +00:00			raise ValueError("`Dataloader` returned 0 length. Please make sure that it returns at least 1 batch")
Follow up of #2892 (#3202) * Follow up of #2892 * typo * iterabledataset 2020-08-27 19:28:29 +00:00			`has_len = True`
			`except TypeError:`
			`has_len = False`
			`except NotImplementedError: # e.g. raised by torchtext if a batch_size_fn is used`
			`has_len = False`

Remove torch<=1.4.0 checks (#5998) * Remove torch<=1.4.0 checks * Update pytorch_lightning/utilities/data.py Co-authored-by: Kaushik B <45285388+kaushikb11@users.noreply.github.com> Co-authored-by: Kaushik B <45285388+kaushikb11@users.noreply.github.com> 2021-02-16 22:53:40 +00:00			`if has_len and has_iterable_dataset(dataloader):`
Follow up of #2892 (#3202) * Follow up of #2892 * typo * iterabledataset 2020-08-27 19:28:29 +00:00			`rank_zero_warn(`
Replace `yapf` with `black` (#7783) Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> 2021-07-26 11:37:35 +00:00			"Your `IterableDataset` has `__len__` defined."
			`" In combination with multi-process data loading (when num_workers > 1),"`
			" `__len__` could be inaccurate if each worker is not configured independently"
			`" to avoid having duplicate data."`
Follow up of #2892 (#3202) * Follow up of #2892 * typo * iterabledataset 2020-08-27 19:28:29 +00:00			`)`
			`return has_len`
Add Support for multiple train loaders (#1959) * add support for wrong dtype in apply_func * apply loader resetting to possible collection of loaders * add combined loader iter class * integrate combined loader iter to training loop * fix imports * fix imports * finish supporters * add tests for supporters * add test for model with multiple loaders * fix trainer integration * fix instance check * Train loaders (#4032) * patch for issues discussed in #1959, encapsulating underlying datastructures returned from train_dataloader * update data_loading.py to it uses patch discussed in #1959 * rename class * Separate CombinedLoaderIterator into two classes, and update related tests. (#4606) * Fix the bugs after rebasing. * Add custom get_len for apply_to_collection * Refactor MultiIterator to be as CombinedLoaderIterator * To get the right num_training_batches. Call the wrapper for multi trainloader in data_loading.py, instead of training_loop.py * Reload _loader_iters when calling __iter__ * Don't transform DataLoader to CombinedLoaderIterator when it's along * Updates test_fit_multiple_train_loaders for testing num_training_batches * Seperate CombinedLoaderIterator into CombinedLoaderIterator and CombinedDataLoader. Add CombinedDataset for unified DataLoader format. * Initialize CombinedDataLoader before calculating num_training_batches. Also updating self._worker_check for multiple loaders * Update tests for supporters * Update tests for multiple trainloaders. Add tests about few_workers for multiple loaders. * Fix pep8 issues * Add tests for train_loader_patch.py * Add descriptions to multiple_trainloader_mode * Remove unused variables * Add docstrings and typing * Add more tests for better converage * Remove unused commented codes * Add sampler property * Remove extract_dataset * Update typing * pep8 * Update train_loader_patch.py * Apply suggestions from code review Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com> * Update pytorch_lightning/trainer/supporters.py Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com> * reviewer comments * fix stupid import * add docs * add back line separator * fix line sep * pep8 * Apply suggestions from code review * fix * fix * Apply suggestions from code review Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * Apply suggestions from code review Co-authored-by: Nicki Skafte <skaftenicki@gmail.com> * flake8 Co-authored-by: Justus Schock <justusschock@justuss-mbp.fritz.box> Co-authored-by: Christofer Fransson <christofer_fransson@yahoo.com> Co-authored-by: YI-LIN SUNG <r06942076@ntu.edu.tw> Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com> Co-authored-by: Rohit Gupta <rohitgr1998@gmail.com> Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> Co-authored-by: Nicki Skafte <skaftenicki@gmail.com> Co-authored-by: Jirka Borovec <jirka.borovec@seznam.cz> 2021-01-04 19:57:53 +00:00

			`def get_len(dataloader: DataLoader) -> Union[int, float]:`
Replace `yapf` with `black` (#7783) Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> 2021-07-26 11:37:35 +00:00			"""Return the length of the given DataLoader. If ``__len__`` method is not implemented, return float('inf')."""
Add Support for multiple train loaders (#1959) * add support for wrong dtype in apply_func * apply loader resetting to possible collection of loaders * add combined loader iter class * integrate combined loader iter to training loop * fix imports * fix imports * finish supporters * add tests for supporters * add test for model with multiple loaders * fix trainer integration * fix instance check * Train loaders (#4032) * patch for issues discussed in #1959, encapsulating underlying datastructures returned from train_dataloader * update data_loading.py to it uses patch discussed in #1959 * rename class * Separate CombinedLoaderIterator into two classes, and update related tests. (#4606) * Fix the bugs after rebasing. * Add custom get_len for apply_to_collection * Refactor MultiIterator to be as CombinedLoaderIterator * To get the right num_training_batches. Call the wrapper for multi trainloader in data_loading.py, instead of training_loop.py * Reload _loader_iters when calling __iter__ * Don't transform DataLoader to CombinedLoaderIterator when it's along * Updates test_fit_multiple_train_loaders for testing num_training_batches * Seperate CombinedLoaderIterator into CombinedLoaderIterator and CombinedDataLoader. Add CombinedDataset for unified DataLoader format. * Initialize CombinedDataLoader before calculating num_training_batches. Also updating self._worker_check for multiple loaders * Update tests for supporters * Update tests for multiple trainloaders. Add tests about few_workers for multiple loaders. * Fix pep8 issues * Add tests for train_loader_patch.py * Add descriptions to multiple_trainloader_mode * Remove unused variables * Add docstrings and typing * Add more tests for better converage * Remove unused commented codes * Add sampler property * Remove extract_dataset * Update typing * pep8 * Update train_loader_patch.py * Apply suggestions from code review Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com> * Update pytorch_lightning/trainer/supporters.py Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com> * reviewer comments * fix stupid import * add docs * add back line separator * fix line sep * pep8 * Apply suggestions from code review * fix * fix * Apply suggestions from code review Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * Apply suggestions from code review Co-authored-by: Nicki Skafte <skaftenicki@gmail.com> * flake8 Co-authored-by: Justus Schock <justusschock@justuss-mbp.fritz.box> Co-authored-by: Christofer Fransson <christofer_fransson@yahoo.com> Co-authored-by: YI-LIN SUNG <r06942076@ntu.edu.tw> Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com> Co-authored-by: Rohit Gupta <rohitgr1998@gmail.com> Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> Co-authored-by: Nicki Skafte <skaftenicki@gmail.com> Co-authored-by: Jirka Borovec <jirka.borovec@seznam.cz> 2021-01-04 19:57:53 +00:00
			`if has_len(dataloader):`
			`return len(dataloader)`

Replace `yapf` with `black` (#7783) Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> 2021-07-26 11:37:35 +00:00			`return float("inf")`