Co-authored-by: Guillaume Lemaitre <g.lemaitre58@gmail.com>
Co-authored-by: Olivier Grisel <olivier.grisel@ensta.org>
Co-authored-by: Thomas J. Fan <thomasjpfan@gmail.com>
Co-authored-by: Julien Jerphanion <git@jjerphan.xyz>
Co-authored-by: Thomas J. Fan <thomasjpfan@gmail.com>
Co-authored-by: Chiara Marmo <cmarmo@users.noreply.github.com>
* Push scipy min version to 1.0.0
* Update all ubuntu images to 20.04 focal.
* Add ubuntu images 18.04 bionic and scipy fron conda-forge.
* Fix conditions.
* Pin python 3.6 for ubuntu bionic.
* Change pipeline name.
* Change matrix element name.
* Keep python 3.9 from system not conda in Ubuntu 20.04.
* Remove python directive when unnecessary.
* Cleanup.
* Downgrade to python 3.6 as scipy 1.0.0 is incompatible with 3.8.
* Fix comment.
* Fix comment.
* Pin pytest again as we are forced to use 3.6.
* Move to conda installer for 32bit linux.
* Install miniconda for ubuntu 32bit.
* Install wget for ubuntu 32bit.
* Revert 32bit OS to ubuntu bionic 18.04.
* Install scipy from pip in 32bit system.
* Fix doctest failures.
* Revert example rendering.
* Relax pytest version in ubuntu install.
* Skip failing tests.
* Put comment at the right place.
* Remove python3.6. Ubuntu32 still needs to be adapted.
* Push numpy and scipy min versions for compatibility with 3.7.
* Push matplotlib min version for compatibility with 3.7. Install numpy via pip in 32bit linux.
* Install numpy before scipy in Linux 32bit.
* Pass numpy version to linux32.
* Test 32bit architecture on debian buster (still exists for 32bit with python 3.7).
* Install matplotlib from distribution.
* Syntax error...
* Stick to the numpy debian version to avoid Expected 124 from C header, got 112 from PyObject error.
* Clean comments.
* Revert skip in doctest to check with new dependencies.
* Rename distrib.
* Skip again...
* Fix test on check_array.
* Remove comment and fix lint at the same time.
* Clean import.
* Increase atol in test_derivatives to make the test pass in py37_conda_openblas environment.
* Avoid sparse matrix dependent on scipy version.
* Skip docstring test for pandas versions less then 1.1.0.
* Fix lint error.
* Empty commit to force checks.
* Add minimal dependencies in changelog.
* Update to python 3.7 CircleCI and Travis builds.
* Move to debian buster for python3.7 dependencies.
* Fix the container tag.
* Lower the minimal pandas version for compatibility with python 3.7.
Co-authored-by: Sylvain MARIE <sylvain.marie@se.com>
Co-authored-by: Thomas J Fan <thomasjpfan@gmail.com>
Co-authored-by: Nicolas Hug <contact@nicolas-hug.com>
Co-authored-by: Joel Nothman <joel.nothman@gmail.com>
Co-authored-by: Olivier Grisel <olivier.grisel@ensta.org>
Co-authored-by: Olivier Grisel <olivier.grisel@gmail.com>
Co-authored-by: Tom Dupré la Tour <tom.dupre-la-tour@m4x.org>
* Initial implementation
* Forgot to add to second __add__ list
* Update split method parameter doc
* Added example; changed default test_size to 0.1; added to author list
* StratifiedGroupKFold impl and other improvements
* Add class to __all__ spec
* Remove random_state when no shuffle
* Tighter formatting
* Update the implementation of StratifiedGroupKFold
* Add StratifiedGroupKFold to __init__
* Add y checks to StartifiedGroupKFold
* Raise error if n_splits > max num samples in class
* Warn if n_splits > mn num samples in class
* Add SGKfold to general repr test
* Add SGKFold to 2d_y test case
* Add SGKfold to value erros test case
Parameters are the same as for StratifiedKFold
to ensure similar behavior given n_groups == n_samples
* Add SGKFold to StratifiedKFold test cases
The idea is to ensure similar behavior when groups are trivial
(n_groups == n_samples)
* Add SGKFold to reproducibility test case
* Add SGKFold to GroupKFold test case
* Add SGKFold to nested cv test case
* Add SGKFold to random_state with shuffle=False test case
* Add SGKFold to constant splits test case
* Fix repr test case
* Fix formatting issues
* Add samples to a fold with least num samples
Required to produce balanced size folds when the distribution of y is
more or less the same
* Remove GroupShuffleSplit impl
* Add notes to StratifiedGroupKFold
* Fix doctest
* Added stratified group kfold tests
* Better variable naming
* Add section to documentation
* Remove leftover StratifiedGroupShuffleSplit import
* Add changelist and reference to original kernel
* Better naming for least populated class check
* Better expression for number of labels
* Remove use of Counter
We already have this data in output of np.unique
* Add tests for homogeneous groups
* Add StratifiedGroupKFold test against GroupKFold
* Add changes to changelist in docstring
* Add StratifiedGroupKFold to classes.rst
* Fix description of StratifiedGroupKFold
* Move license notice out of docstring
* Disambiguate labels to classes in doc
* Add changelog entry
* Fix changelog author entry
* Fix StratifiedGroupKFold docstring
* Better variable names
* Remove defaultdict in favor of numpy indexing
* Extracted best_fold search into a separate method
* Make use of numpy broadcasting instead of for loop
* Encode groups and use arrays instead of dicts
* Use numpy sort instead of python
* Clarify shuffling behavior of StratifiedGroupKF in docs
* Switch name from label_idx to class_idx
* Remove accidentally leftover comment
* Fix np.sort keyword to support numpy < 1.15
* Fix typo in docstring
* Add StratifiedGroupKFold to visualization doc
* Add visualization for uneven group as an example
* Fix image numbers to match updated example
* Add author
* Add SGKF visualization to docs
* Add comments for groups in stratified CV tests
Co-authored-by: Leandro Hermida <hermidal@cs.umd.edu>
Co-authored-by: marrodion <rodion_martynov@epam.com>