lightning

Commit Graph

Author	SHA1	Message	Date
takahashi	52c07f2f03	Docs: Fix broken get-started link (#5960 )	2021-02-15 01:00:19 +01:00
Adrian Wälchli	c912c4b729	remove legacy accelerators (#5949 ) * remove legacy accelerators * update imports * formatting Co-authored-by: Jirka Borovec <jirka.borovec@seznam.cz>	2021-02-14 16:03:45 +00:00
Adrian Wälchli	a3d4e7c86a	move accelerator legacy tests (#5948 ) Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>	2021-02-13 19:42:18 -05:00
William Falcon	0345fcfaad	Update README.md	2021-02-13 14:44:19 -05:00
William Falcon	194f048263	Update README.md	2021-02-13 14:41:38 -05:00
William Falcon	d924dd6a41	Update README.md	2021-02-13 14:01:34 -05:00
William Falcon	11942558d0	Update README.md	2021-02-13 13:50:18 -05:00
William Falcon	e839d3bc22	Update README.md	2021-02-13 13:43:29 -05:00
William Falcon	ecf995bad9	Update README.md	2021-02-13 12:31:03 -05:00
William Falcon	891fc64af7	Update README.md	2021-02-13 12:27:44 -05:00
William Falcon	fcf894b621	Update README.md	2021-02-13 12:21:31 -05:00
William Falcon	68aac1f9df	Update README.md	2021-02-13 12:13:14 -05:00
William Falcon	7e69fbbc2e	Update README.md	2021-02-13 11:52:02 -05:00
William Falcon	d81851843e	Update README.md	2021-02-13 11:51:33 -05:00
Jirka Borovec	046ac714f6	v1.2.0rc1 (#5946 ) * v1.2.0rc0 * chlog * chlog * chlog * chlog Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>	2021-02-13 12:58:25 +00:00
Kaushik B	42dc5d2af1	Fix: Repeated .fit() calls ignore max_steps iteration bound (#5936 ) * fix repeated fit calls ignoring max_steps * fix fast dev progress bar	2021-02-13 07:36:22 +00:00
Adrian Wälchli	b8619a695f	new LightningModule hook "configure_callbacks" (#5621 ) Co-authored-by: Rohit Gupta <rohitgr1998@gmail.com> Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com> Co-authored-by: chaton <thomas@grid.ai> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>	2021-02-12 19:27:44 -05:00
Justus Schock	da6dbc8d1d	PoC: Accelerator refactor (#5743 ) * restoring the result from subprocess * fix queue.get() order for results * add missing "block_backward_sync" context manager * add missing "block_backward_sync" context manager * fix sync_batchnorm * fix supported gpu-ids for tuple * fix clip gradients and inf recursion * accelerator selection: added cluster_environment plugin * fix torchelastic test * fix reduce early stopping decision for DDP * fix tests: callbacks, conversion to lightning optimizer * fix lightning optimizer does not pickle * fix setting benchmark and deterministic option * fix slurm amp test * fix prepare_data test and determine node_rank * fix retrieving last path when testing * remove obsolete plugin argument * fix test: test_trainer_config * fix torchscript tests * fix trainer.model access * move properties * fix test_transfer_batch_hook * fix auto_select_gpus * fix omegaconf test * fix test that needs to simulate slurm ddp * add horovod plugin * fix test with named arguments * clean up whitespace * fix datamodules test * remove old accelerators * fix naming * move old plugins * move to plugins * create precision subpackage * create training_type subpackage * fix all new import errors * fix wrong arguments order passed to test * fix LR finder * Added sharded training type and amp plugin * Move clip grad to precision plugin * Added sharded spawn, select accelerators based on distributed_backend + enable custom fp16 plugin automatically * Fix import issue, attempting to fix tests * Fix initial test * Reflect hook logic from master, should wrap model after move to device * Optional state consolidation, since master has optimizers not wrapped * change attribute for instance test * reset optimizers optimizers are not used in main process, so state would be wrong. * legacy * imports in accel * legacy2 * trainer imports * fix import errors after rebase * move hook to new setup location * provide unwrapping logic * fix trainer callback system * added ddp2 implementation * fix imports .legacy * move plugins * restore legacy * drop test.py from root * add tpu accelerator and plugins * fixes * fix lightning optimizer merge * reset bugreportmodel * unwrapping * step routing forward * model access * unwrap * opt * integrate distrib_type * sync changes * sync * fixes * add forgotten generators * add missing logic * update * import * missed imports * import fixes * isort * mv f * changelog * format * move helper to parallel plugin * d * add world size * clean up * duplicate * activate ddp_sharded and tpu * set nvidia flags * remove unused colab var * use_tpu <-> on_tpu attrs * make some ddp_cpu and clusterplugin tests pass * Ref/accelerator connector (#5742) * final cleanup Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * connector cleanup Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * trainer cleanup Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * accelerator cleanup + missing logic in accelerator connector Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * add missing changes to callbacks Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * reflect accelerator changes to lightning module Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * clean cluster envs Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * cleanup plugins Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * add broadcasting Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * yapf * remove plugin connector Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * plugins * manual optimization * update optimizer routing * add rank to torchelastic * fix memory mixed precision * setstate on trainer for pickling in ddp spawn * add predict method * add back commented accelerator code * adapt test for sync_batch_norm to new plugin * fix deprecated tests * fix ddp cpu choice when no num_processes are given * yapf format * skip a memory test that cannot pass anymore * fix pickle error in spawn plugin * x * avoid * x * fix cyclic import in docs build * add support for sharded * update typing * add sharded and sharded_spawn to distributed types * make unwrap model default * refactor LightningShardedDataParallel similar to LightningDistributedDataParallel * update sharded spawn to reflect changes * update sharded to reflect changes * Merge 1.1.5 changes * fix merge * fix merge * yapf isort * fix merge * yapf isort * fix indentation in test * copy over reinit scheduler implementation from dev1.2 * fix apex tracking calls with dev_debugger * reduce diff to dev1.2, clean up * fix trainer config test when gpus>0 and num_processes >0 and ddp_cpu * sort plugin tests legacy/new * fix error handling for amp on cpu * fix merge fix merge fix merge * [Feat] Resolve manual_backward (#5837) * resolve manual_backward * resolve flake8 * update * resolve for ddp_spawn * resolve flake8 * resolve flake8 * resolve flake8 Co-authored-by: Ubuntu <ubuntu@ip-172-31-88-60.ec2.internal> * fix tests/accelerator tests on cpu * [BugFix] Resolve manual optimization (#5852) * resolve manual_optimization * update * update Co-authored-by: Ubuntu <ubuntu@ip-172-31-88-60.ec2.internal> * Remove copy trainer parameters to happen earlier within the loop and add safe guard to get ref model (#5856) * resovle a bug * Accelerator refactor sharded rpc (#5854) * rpc branch * merge * update handling of rpc * make devices etc. Optional in RPC * set devices etc. later if necessary * remove devices from sequential * make devices optional in rpc * fix import * uncomment everything * fix cluster selection Co-authored-by: Ubuntu <ubuntu@ip-172-31-88-60.ec2.internal> * resolve bug * fix assert in rpc test * resolve a test * fix docs compilation * accelerator refactor - fix for sharded parity test (#5866) * fix memory issue with ddp_spawn * x x x x x x x x x * x * Remove DDP2 as this does not apply * Add missing pre optimizer hook to ensure lambda closure is called * fix apex docstring * [accelerator][BugFix] Resolve some test for 1 gpu (#5863) * update * revert init * resolve a bug * update * resolve flake8 * update * update * update * revert init * resolve a bug * update * resolve flake8 * update * update * update * update * update * revert init * resolve a bug * update * resolve flake8 * update * update * update * revert init * update * resolve flake8 * update * update * update * update * update * all_gather * update * make plugins work, add misconfig for RPC * update * update * remove breaking test * resolve some tests * resolve flake8 * revert to ddp_spawn Co-authored-by: root <root@ip-172-31-88-60.ec2.internal> Co-authored-by: Ubuntu <ubuntu@ip-172-31-88-60.ec2.internal> Co-authored-by: Justus Schock <justus.schock@rwth-aachen.de> * yapf isort * resolve flake8 * fix apex doctests * fix apex doctests 2 * resolve docs * update drone * clean env * update * update * update * update * merge * Fix RPC related tests, clean out old API, update for new accelerator API [skip ci] (#5881) * Fix RPC related tests, clean out old API, update for new accelerator API * Move tests out of legacy folder, update paths and names * Update test_remove_1-4.py * Expose properties for tpu cores/gpus/num_gpus * Add root GPU property * Move properties to properties.py * move tests that were previously in drone * Fix root GPU property (#5908) * Move root GPU to property, remove horovod set as this is handled in horovod plugin, ensure we mock correctly to set GPU accelerator * Add missing tests back * fix best model path transfer when no checkpoint callback available * Fix setup hook order [wip] (#5858) * Call trainer setup hook before accelerator setup * Add test case * add new test * typo * fix callback order in test Co-authored-by: tchaton <thomas@grid.ai> Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * rename ddp sequential -> rpc sequential for special test * revert * fix stupid merge problem * Use property in connector for sampler (#5913) * merge the import conflicts * fix spawning of processes in slurm * [wip] Fix some bugs for TPU [skip ci] (#5878) * fixed for single tpu * fixed spawn * fixed spawn * update * update * wip * resolve bugs * resolve bug * update on comment * removed decorator * resolve comments * set to 4 * update * update * need cleaning * update * update * update * resolve flake8 * resolve bugs * exclude broadcast * resolve bugs * change test * update * update * skip if meet fails * properly raise trace * update * add catch * wrap test * resolve typo * update * typo Co-authored-by: Lezwon Castelino <lezwon@gmail.com> Co-authored-by: Your Name <you@example.com> * resolve some tests * update * fix imports * update * resolve flake8 * update azure pipeline * skip a sharded test on cpu that requires a gpu * resolve tpus * resolve bug * resolve flake8 * update * updat utils * revert permission change on files * suggestions from carlos Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com> * remove unrelated formatting changes * remove incomplete comment * Update pytorch_lightning/accelerators/__init__.py Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com> * remove unrelated formatting change * add types * warn 1.7 ddp manual backward only if ddp kwarg unset * yapf + isort * pep8 unused imports * fix cyclic import in docs * Apply suggestions from code review * typer in accelerator.py * typo * Apply suggestions from code review * formatting * update on comments * update typo * Update pytorch_lightning/trainer/properties.py Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> * update * suggestion from code review * suggestion from code review Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com> Co-authored-by: SeanNaren <sean@grid.ai> Co-authored-by: Jirka Borovec <jirka.borovec@seznam.cz> Co-authored-by: chaton <thomas@grid.ai> Co-authored-by: Ubuntu <ubuntu@ip-172-31-88-60.ec2.internal> Co-authored-by: Sean Naren <sean.narenthiran@gmail.com> Co-authored-by: root <root@ip-172-31-88-60.ec2.internal> Co-authored-by: Lezwon Castelino <lezwon@gmail.com> Co-authored-by: Your Name <you@example.com> Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com> Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com> Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>	2021-02-12 15:48:56 -05:00
Dusan Drevicky	309ce7a966	Fix: passing wrong strings for scheduler interval doesn't throw an error (#5923 ) * Raise if scheduler interval not 'step' or 'epoch' * Add test for unknown 'interval' value in scheduler * Use BoringModel instead of EvalModelTemplate Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com> * Fix import order * Apply yapf in test_datamodules * Add missing imports to test_datamodules * Fix too long comment * Update pytorch_lightning/trainer/optimizers.py * Fix unused imports and exception message * Fix failing test Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com> Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com>	2021-02-13 01:31:22 +05:30
Eric Cousineau	ae19c9723b	tests: Remove usage of --flake8 flag (#5909 ) * tests: Remove usage of --flake8 flag * Remove commented line Co-authored-by: Carlos Mocholi <carlossmocholi@gmail.com>	2021-02-12 12:25:08 -05:00
Nicki Skafte	979c879e45	drop DDP CLI test (#5938 ) * fix tests * = Co-authored-by: Jirka Borovec <jirka.borovec@seznam.cz>	2021-02-12 17:42:32 +01:00
Adrian Wälchli	4bdf2fe55f	remove executable bit on source files (#5929 ) * 644	2021-02-12 00:06:40 +01:00
Kaushik B	4857546c25	Fix: Failing test in data_modules(dp) (#5924 ) * Update test_datamodules.py * fix code format issue * fix test restore * fix code format issue	2021-02-11 17:32:46 +00:00
Jirka Borovec	e676ff96b1	Typing: callback base (#5919 ) * typing for callback base	2021-02-11 14:33:10 +00:00
Carlos Mocholí	9f12ca095a	More EpochResultStore refactors! 🎉 (#5522 ) Co-authored-by: chaton <thomas@grid.ai> Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com>	2021-02-11 14:32:45 +00:00
Teddy Koker	253e57c2c2	Feature: LightningDataModule.from_datasets(...) (#5133 ) * add class method * add tests * docstring * pep * Add type annotations Co-authored-by: Nicki Skafte <skaftenicki@gmail.com> * pep * fix import * remove num_workers inference * Update pytorch_lightning/core/datamodule.py Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com> * Update pytorch_lightning/core/datamodule.py Co-authored-by: Nicki Skafte <skaftenicki@gmail.com> * Update pytorch_lightning/core/datamodule.py Co-authored-by: Nicki Skafte <skaftenicki@gmail.com> * fix syntax * typing fix * list -> sequence * list -> sequence * missing import * fix test Co-authored-by: Nicki Skafte <skaftenicki@gmail.com> Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com>	2021-02-11 14:32:41 +00:00
Nicki Skafte	31da16344c	add docs (#5902 )	2021-02-11 14:32:32 +00:00
Carlos Mocholí	414aa5d345	Delete unused autopep8 config (#5904 )	2021-02-11 14:32:14 +00:00
Teddy Koker	44958ad964	forward cache fix (#5895 )	2021-02-11 14:32:12 +00:00
Nicki Skafte	0c80b9f890	fix metric docs (#5880 )	2021-02-11 14:32:12 +00:00
Rohit Gupta	cf30b956a2	update example (#5753 )	2021-02-11 14:32:11 +00:00
Rohit Gupta	8e9a026bc3	[tests/models] refactor with BoringModel (#5507 ) * update with BoringModel * update with BoringModel * step * try TPU * TPU * update tests * update tpu tests * self * fix * dp * update tests * ref * update tests * fix tpu tests * fix dp and run_prediction * dp * only dp * Apply suggestions from code review * Apply suggestions from code review * Apply suggestions from code review * Apply suggestions from code review Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com>	2021-02-11 14:32:07 +00:00
Jirka Borovec	b434c479e7	Quantisation (#5706 ) * empty * sq * obs * int * ts * helpers * chlog * yapf * avg * dupl * Apply suggestions from code review Co-authored-by: Nicki Skafte <skaftenicki@gmail.com> * Apply suggestions from code review Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com> * fixes * Apply suggestions from code review Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com> * fixes * note * warn * 45 * link Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com> * Apply suggestions from code review Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com> * yapf * flake8 * Apply suggestions from code review * Apply suggestions from code review Co-authored-by: Nicki Skafte <skaftenicki@gmail.com> Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com>	2021-02-11 07:04:57 -05:00
Jirka Borovec	9475c845cb	Docs/fixes (#5914 ) * wip * .. * ... * Apply suggestions from code review Co-authored-by: Justus Schock <12886177+justusschock@users.noreply.github.com> Co-authored-by: Justus Schock <12886177+justusschock@users.noreply.github.com>	2021-02-11 10:22:07 +00:00
Carlos Mocholí	e8190e8848	Convert progress bar metrics to float (#5692 ) * MetricsHolder(to_float=True) * Update CHANGELOG * Update tests/callbacks/test_progress_bar.py * flake8 Co-authored-by: Jirka Borovec <jirka.borovec@seznam.cz>	2021-02-10 19:16:53 -05:00
chaton	7b00894130	[feat] Add StochasticWeightAveragingCallback (#5640 ) * add swa callback * switch back to 1.6.0 * remove optimizer_step * move super * update * forgot update_parameters * update on comments * works for ddp * resolve flake8 * remove set_model * resolve flake8 * resolve cpu * resolve flake8 * resolve flake8 * update * update on comments	2021-02-11 00:05:59 +00:00
Carlos Mocholí	ecd3678a1f	Refactor utilities/imports.py (#5874 ) Co-authored-by: Jirka Borovec <jirka.borovec@seznam.cz> Co-authored-by: Nicki Skafte <skaftenicki@gmail.com>	2021-02-11 00:32:47 +01:00
Jirka Borovec	373a31e63e	add azure timeout (#5907 ) * add azure timeout * rework	2021-02-10 20:21:20 +00:00
Carlos Mocholí	a028171f26	Fix Pruning callback and add a few features (#5825 ) * Remove pruning check because it was added in 1.4.0 and that is our minimal torch version * Fixing many bugs * Fix misconfig test * Fix tests * Improve error message * Reduce whitespace * WIP * TODOs * _MODULE_CONTAINERS * Add LTH test * Allow resampling * Iterative pruning * Log pruning percentage * Properly make pruning permanent * Fix docstring * Minor changes * Test loading non-permanent model * corrent bugs * Revert "corrent bugs" This reverts commit `ffb8d47547`. * Add beta warning * Fix docs * 2 verbosity levels * OCD Co-authored-by: Your Name <you@example.com>	2021-02-10 15:03:23 +00:00
Jirka Borovec	388eeea52d	bye bye Drone (#5901 )	2021-02-10 09:13:38 -05:00
Jirka Borovec	c2c82dad62	CI: Azure (#5882 ) * add base Azure pipeline * skip	2021-02-10 04:43:26 -05:00
ananthsub	d26702bd66	Enable purely iteration-based training (#5726 ) Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com> Co-authored-by: Carlos Mocholí <carlossmocholi@gmail.com> Co-authored-by: Kaushik B <45285388+kaushikb11@users.noreply.github.com> Co-authored-by: Kaushik Bokka <kaushikbokka@gmail.com>	2021-02-10 08:51:08 +00:00
Jirka Borovec	a96137c035	codeowner: update test/boring model (#5869 ) * update test/boring model * reorder	2021-02-09 21:30:14 +00:00
Jirka Borovec	9dd56398e3	fixing some compatibility with PT 1.8 (#5864 ) * change default * . * p * 0.21.2 * . * fix * .	2021-02-09 18:25:57 +01:00
Jirka Borovec	1ac9164f91	create new Conda images (#5877 ) * create new Conda images * . * .	2021-02-09 15:30:48 +00:00
Jirka Borovec	a0f7831278	fix miss-leading imports in tests (#5873 ) * fix imorts * .	2021-02-09 05:10:52 -05:00
Jirka Borovec	937f11c05b	try fix: Docker with Conda & PT 1.8 (#5842 ) * ci * ver * list * pt * nk * ch * 4.9	2021-02-09 08:22:35 +00:00
Jirka Borovec	774e9bd0d5	prune Drone build (#5871 )	2021-02-08 14:35:36 -05:00
Jirka Borovec	58c90ba7b7	formatting 5/n: Core (#5721 ) * yapf core * Update pytorch_lightning/core/lightning.py Co-authored-by: Rohit Gupta <rohitgr1998@gmail.com>	2021-02-08 14:29:43 -05:00
Jirka Borovec	79d42d83e7	formatting 3/n: PL modules (#5716 ) * cb * log * prof * tune * flake8	2021-02-08 14:28:38 -05:00

1 2 3 4 5 ...

4331 Commits All Branches Search

4331 Commits

All Branches