* Initial implementation
* Forgot to add to second __add__ list
* Update split method parameter doc
* Added example; changed default test_size to 0.1; added to author list
* StratifiedGroupKFold impl and other improvements
* Add class to __all__ spec
* Remove random_state when no shuffle
* Tighter formatting
* Update the implementation of StratifiedGroupKFold
* Add StratifiedGroupKFold to __init__
* Add y checks to StartifiedGroupKFold
* Raise error if n_splits > max num samples in class
* Warn if n_splits > mn num samples in class
* Add SGKfold to general repr test
* Add SGKFold to 2d_y test case
* Add SGKfold to value erros test case
Parameters are the same as for StratifiedKFold
to ensure similar behavior given n_groups == n_samples
* Add SGKFold to StratifiedKFold test cases
The idea is to ensure similar behavior when groups are trivial
(n_groups == n_samples)
* Add SGKFold to reproducibility test case
* Add SGKFold to GroupKFold test case
* Add SGKFold to nested cv test case
* Add SGKFold to random_state with shuffle=False test case
* Add SGKFold to constant splits test case
* Fix repr test case
* Fix formatting issues
* Add samples to a fold with least num samples
Required to produce balanced size folds when the distribution of y is
more or less the same
* Remove GroupShuffleSplit impl
* Add notes to StratifiedGroupKFold
* Fix doctest
* Added stratified group kfold tests
* Better variable naming
* Add section to documentation
* Remove leftover StratifiedGroupShuffleSplit import
* Add changelist and reference to original kernel
* Better naming for least populated class check
* Better expression for number of labels
* Remove use of Counter
We already have this data in output of np.unique
* Add tests for homogeneous groups
* Add StratifiedGroupKFold test against GroupKFold
* Add changes to changelist in docstring
* Add StratifiedGroupKFold to classes.rst
* Fix description of StratifiedGroupKFold
* Move license notice out of docstring
* Disambiguate labels to classes in doc
* Add changelog entry
* Fix changelog author entry
* Fix StratifiedGroupKFold docstring
* Better variable names
* Remove defaultdict in favor of numpy indexing
* Extracted best_fold search into a separate method
* Make use of numpy broadcasting instead of for loop
* Encode groups and use arrays instead of dicts
* Use numpy sort instead of python
* Clarify shuffling behavior of StratifiedGroupKF in docs
* Switch name from label_idx to class_idx
* Remove accidentally leftover comment
* Fix np.sort keyword to support numpy < 1.15
* Fix typo in docstring
* Add StratifiedGroupKFold to visualization doc
* Add visualization for uneven group as an example
* Fix image numbers to match updated example
* Add author
* Add SGKF visualization to docs
* Add comments for groups in stratified CV tests
Co-authored-by: Leandro Hermida <hermidal@cs.umd.edu>
Co-authored-by: marrodion <rodion_martynov@epam.com>
* More flexible grid search interface
* added info dict parameter
* Put back removed test
* renamed info into more_results
* Passed grroups as well since we need n_to use get_n_splits(X, y, groups)
* port
* pep8
* dabl -> sklearn
* add _required_parameters
* skipping check in rst file if pandas not installed
* Update sklearn/model_selection/_search_successive_halving.py
Co-Authored-By: Joel Nothman <joel.nothman@gmail.com>
* renamed into GridHalvingSearchCV and RandomHalvingSearchCV
* Addressed thomas' comments
* repr
* removed passing group as a parameter to evaluate_candidates
* Joels comments
* pep8
* reorganized user user guide
* renaming
* update user guide
* remove groups support + pass fit_params
* parameter renaming
* pep8
* r_i -> resource_iter
* fixed r_i issues
* examples + removed use of word budget
* Added inpute checking tests
* added cv_resutlts_ user guide
* minor title change
* fixed doc layout
* Addressed some comments
* properly pass down fit_params
* change default value of force_exhaust_resources and update doc
* should fix doc
* Used check_fit_params
* Update section about min_resources and number of candidates
* Clarified ratio section
* Use ~ to refer to classes
* fixed doc checks
* Apply suggestions from code review
Co-authored-by: Joel Nothman <joel.nothman@gmail.com>
* Addressed easy comments from Joel
* missed some
* updated docstring of run_search
* Used f strings instead of format
* remove candidate duplication checks
* fix example
* Addressed easy comments
* rotate ticks labels
* Added discussion in the intro as suggested by Joel
* Split examples into sections
* minor changes
* remove force_exhaust_budget and introduce min_resources=exhaust
* some minor validation
* Added a n_resources_ attribute
* update examples
* Addressed comments
* passing CV instead of X,y
* minor revert for handling fit_params
* updated docs
* fix len
* whatsnew
* Add test for sampling when all_list
* minor change to top-k
* Force CV splits to be consistent across calls
* reorder parameters
* reduced diff
* added tests for top_k
* put back doc for groups
* not sure what went wrong
* put import at its place
* some comment
* Addressed comments
* Added tests for cv_results_ and base estimator inputs
* pep8
* avoid monkeypatching
* rename df
* use Joel's suggestions for testing masks
* Made it experimental
* Should fix docs
* whats new entry
* Apply suggestions from code review
Co-authored-by: Andreas Mueller <t3kcit@gmail.com>
* Addressed comments to docs
* Addressed comments in examples
* minor doc update
* minor renaming in UG
* forgot some
* some sad note about splitter statefulness :'(
* Addressed comments
* ratio -> factor
Co-authored-by: Joel Nothman <joel.nothman@gmail.com>
Co-authored-by: Andreas Mueller <t3kcit@gmail.com>
* ENH cross_val_score now supports multiple metrics
* DOCFIX permutation_test_score
* ENH validate multiple metric scorers
* ENH Move validation of multimetric scoring param out
* ENH GridSearchCV and RandomizedSearchCV now support multiple metrics
* EXA Add an example demonstrating the multiple metric in GridSearchCV
* ENH Let check_multimetric_scoring tell if its multimetric or not
* FIX For single metric name of scorer should remain 'score'
* ENH validation_curve and learning_curve now support multiple metrics
* MNT move _aggregate_score_dicts helper into _validation.py
* TST More testing/ Fixing scores to the correct values
* EXA Add cross_val_score to multimetric example
* Rename to multiple_metric_evaluation.py
* MNT Remove scaffolding
* FIX doctest imports
* FIX wrap the scorer and unwrap the score when using _score() in rfe
* TST Cleanup the tests. Test for is_multimetric too
* TST Make sure it registers as single metric when scoring is of that type
* PEP8
* Don't use dict comprehension to make it work in python2.6
* ENH/FIX/TST grid_scores_ should not be available for multimetric evaluation
* FIX+TST delegated methods NA when multimetric is enabled...
TST Add general tests to GridSearchCV and RandomizedSearchCV
* ENH add option to disable delegation on multimetric scoring
* Remove old function from __all__
* flake8
* FIX revert disable_on_multimetric
* stash
* Fix incorrect rebase
* [ci skip]
* Make sure refit works as expected and remove irrelevant tests
* Allow passing standard scorers by name in multimetric scorers
* Fix example
* flake8
* Address reviews
* Fix indentation
* Ensure {'acc': 'accuracy'} and ['precision'] are valid inputs
* Test that for single metric, 'score' is a key
* Typos
* Fix incorrect rebase
* Compare multimetric grid search with multiple single metric searches
* Test X, y list and pandas input; Test multimetric for unsupervised grid search
* Fix tests; Unsupervised multimetric gs will not pass until #8117 is merged
* Make a plot of Precision vs ROC AUC for RandomForest varying the n_estimators
* Add example to grid_search.rst
* Use the classic tuning of C param in SVM instead of estimators in RF
* FIX Remove scoring arg in deafult scorer test
* flake8
* Search for min_samples_split in DTC; Also show f-score
* REVIEW Make check_multimetric_scoring private
* FIX Add more samples to see if 3% mismatch on 32 bit systems gets fixed
* REVIEW Plot best score; Shorten legends
* REVIEW/COSMIT multimetric --> multi-metric
* REVIEW Mark the best scores of P/R scores too
* Revert "FIX Add more samples to see if 3% mismatch on 32 bit systems gets fixed"
This reverts commit ba766d98353380a186fbc3dade211670ee72726d.
* ENH Use looping for iid testing
* FIX use param grid as scipy's stats dist in 0.12 do not accept seed
* ENH more looping less code; Use small non-noisy dataset
* FIX Use named arg after expanded args
* TST More testing of the refit parameter
* Test that in multimetric search refit to single metric, the delegated methods
work as expected.
* Test that setting probability=False works with multimetric too
* Test refit=False gives sensible error
* COSMIT multimetric --> multi-metric
* REV Correct example doc
* COSMIT
* REVIEW Make tests stronger; Fix bugs in _check_multimetric_scorer
* REVIEW refit param: Raise for empty strings
* TST Invalid refit params
* REVIEW Use <scorer_name> alone; recall --> Recall
* REV specify when we expect scorers to not be None
* FLAKE8
* REVERT multimetrics in learning_curve and validation_curve
* REVIEW Simpler coding style
* COSMIT
* COSMIT
* REV Compress example a bit. Move comment to top
* FIX fit_grid_point's previous API must be preserved
* Flake8
* TST Use loop; Compare with single-metric
* REVIEW Use dict-comprehension instead of helper
* REVIEW Remove redundant test
* Fix tests incorrect braces
* COSMIT
* REVIEW Use regexp
* REV Simplify aggregation of score dicts
* FIX precision and accuracy test
* FIX doctest and flake8
* TST the best_* attributes multimetric with single metric
* Address @jnothman's review
* Address more comments \o/
* DOCFIXES
* Fix use the validated fit_param from fit's arguments
* Revert alpha to a lower value as before
* Using def instead of lambda
* Address @jnothman's review batch 1: Fix tests / Doc fixes
* Remove superfluous tests
* Remove more superfluous testing
* TST/FIX loop over refit and check found n_clusters
* Cosmetic touches
* Use zip instead of manually listing the keys
* Fix inverse_transform
* FIX bug in fit_grid_point; Allow only single score
TST if fit_grid_point works as intended
* ENH Use only ROC-AUC and F1-score
* Fix typos and flake8; Address Andy's reviews
MNT Add a comment on why we do such a transpose + some fixes
* ENH Better error messages for incorrect multimetric scoring values +...
ENH Avoid exception traceback while using incorrect scoring string
* Dict keys must be of string type only
* 1. Better error message for invalid scoring 2...
Internal functions return single score for single metric scoring
* Fix test failures and shuffle tests
* Avoid wrapping scorer as dict in learning_curve
* Remove doc example as asked for
* Some leftover ones
* Don't wrap scorer in validation_curve either
* Add a doc example and skip it as dict order fails doctest
* Import zip from six for python2.7 compat
* Make cross_val_score return a cv_results-like dict
* Add relevant sections to userguide
* Flake8 fixes
* Add whatsnew and fix broken links
* Use AUC and accuracy instead of f1
* Fix failing doctests cross_validation.rst
* DOC add the wrapper example for metrics that return multiple return values
* Address andy's comments
* Be less weird
* Address more of andy's comments
* Make a separate cross_validate function to return dict and a cross_val_score
* Update the docs to reflect the new cross_validate function
* Add cross_validate to toc-tree
* Add more tests on type of cross_validate return and time limits
* FIX failing doctests
* FIX ensure keys are not plural
* DOC fix
* Address some pending comments
* Remove the comment as it is irrelevant now
* Remove excess blank line
* Fix flake8 inconsistencies
* Allow fit_times to be 0 to conform with windows precision
* DOC specify how refit param is to be set in multiple metric case
* TST ensure cross_validate works for string single metrics + address @jnothman's reviews
* Doc fixes
* Remove the shape and transform parameter of _aggregate_score_dicts
* Address Joel's doc comments
* Fix broken doctest
* Fix the spurious file
* Address Andy's comments
* MNT Remove erroneous entry
* Address Andy's comments
* FIX broken links
* Update whats_new.rst
missing newline
* Add _RepeatedSplits and RepeatedKFold class
* Add RepeatedStratifiedKFold and doc for repeated cvs
* Change default value of n_repeats
* Change input parameters of repeated cv constructor to n_splits, n_repeats, random_state
* Generate random states in split function rather than store it beforehand
* Doc changes, inheriting RepeatedKFold, RepeatedStratifiedKFold from _RepeatedSplits and other review changes
* Remove blank line, put testcases for deterministic split in loop and add StopIteration check in testcase
* Using rng directly as random_state param to create cv instance and added a check for cvargs
* Fix pep8 warnings
* Changing default values for n_splits and n_repeats and add entry in changelog
* Adding name to the feature
* Missing space
--------------------
* ENH Reogranize classes/fn from grid_search into search.py
* ENH Reogranize classes/fn from cross_validation into split.py
* ENH Reogranize cls/fn from cross_validation/learning_curve into validate.py
* MAINT Merge _check_cv into check_cv inside the model_selection module
* MAINT Update all the imports to point to the model_selection module
* FIX use iter_cv to iterate throught the new style/old style cv objs
* TST Add tests for the new model_selection members
* ENH Wrap the old-style cv obj/iterables instead of using iter_cv
* ENH Use scipy's binomial coefficient function comb for calucation of nCk
* ENH Few enhancements to the split module
* ENH Improve check_cv input validation and docstring
* MAINT _get_test_folds(X, y, labels) --> _get_test_folds(labels)
* TST if 1d arrays for X introduce any errors
* ENH use 1d X arrays for all tests;
* ENH X_10 --> X (global var)
Minor
-----
* ENH _PartitionIterator --> _BaseCrossValidator;
* ENH CVIterator --> CVIterableWrapper
* TST Import the old SKF locally
* FIX/TST Clean up the split module's tests.
* DOC Improve documentation of the cv parameter
* COSMIT consistently hyphenate cross-validation/cross-validator
* TST Calculate n_samples from X
* COSMIT Use separate lines for each import.
* COSMIT cross_validation_generator --> cross_validator
Commits merged manually
-----------------------
* FIX Document the random_state attribute in RandomSearchCV
* MAINT Use check_cv instead of _check_cv
* ENH refactor OVO decision function, use it in SVC for sklearn-like
decision_function shape
* FIX avoid memory cost when sampling from large parameter grids
ENH Major to Minor incremental enhancements to the model_selection
Squashed commit messages - (For reference)
Major
-----
* ENH p --> n_labels
* FIX *ShuffleSplit: all float/invalid type errors at init and int error at split
* FIX make PredefinedSplit accept test_folds in constructor; Cleanup docstrings
* ENH+TST KFold: make rng to be generated at every split call for reproducibility
* FIX/MAINT KFold: make shuffle a public attr
* FIX Make CVIterableWrapper private.
* FIX reuse len_cv instead of recalculating it
* FIX Prevent adding *SearchCV estimators from the old grid_search module
* re-FIX In all_estimators: the sorting to use only the 1st item (name)
To avoid collision between the old and the new GridSearch classes.
* FIX test_validate.py: Use 2D X (1D X is being detected as a single sample)
* MAINT validate.py --> validation.py
* MAINT make the submodules private
* MAINT Support old cv/gs/lc until 0.19
* FIX/MAINT n_splits --> get_n_splits
* FIX/TST test_logistic.py/test_ovr_multinomial_iris:
pass predefined folds as an iterable
* MAINT expose BaseCrossValidator
* Update the model_selection module with changes from master
- From #5161
- - MAINT remove redundant p variable
- - Add check for sparse prediction in cross_val_predict
- From #5201 - DOC improve random_state param doc
- From #5190 - LabelKFold and test
- From #4583 - LabelShuffleSplit and tests
- From #5300 - shuffle the `labels` not the `indxs` in LabelKFold + tests
- From #5378 - Make the GridSearchCV docs more accurate.
- From #5458 - Remove shuffle from LabelKFold
- From #5466(#4270) - Gaussian Process by Jan Metzen
- From #4826 - Move custom error / warnings into sklearn.exception
Minor
-----
* ENH Make the KFold shuffling test stronger
* FIX/DOC Use the higher level model_selection module as ref
* DOC in check_cv "y : array-like, optional"
* DOC a supervised learning problem --> supervised learning problems
* DOC cross-validators --> cross-validation strategies
* DOC Correct Olivier Grisel's name ;)
* MINOR/FIX cv_indices --> kfold
* FIX/DOC Align the 'See also' section of the new KFold, LeaveOneOut
* TST/FIX imports on separate lines
* FIX use __class__ instead of classmethod
* TST/FIX import directly from model_selection
* COSMIT Relocate the random_state documentation
* COSMIT remove pass
* MAINT Remove deprecation warnings from old tests
* FIX correct import at test_split
* FIX/MAINT Move P_sparse, X, y defns to top; rm unused W_sparse, X_sparse
* FIX random state to avoid doctest failure
* TST n_splits and split wrapping of _CVIterableWrapper
* FIX/MAINT Use multilabel indicator matrix directly
* TST/DOC clarify why we conflate classes 0 and 1
* DOC add comment that this was taken from BaseEstimator
* FIX use of labels is not needed in stratified k fold
* Fix cross_validation reference
* Fix the labels param doc
FIX/DOC/MAINT Addressing the review comments by Arnaud and Andy
COSMIT Sort the members alphabetically
COSMIT len_cv --> n_splits
COSMIT Merge 2 if; FIX Use kwargs
DOC Add my name to the authors :D
DOC make labels parameter consistent
FIX Remove hack for boolean indices; + COSMIT idx --> indices; DOC Add Returns
COSMIT preds --> predictions
DOC Add Returns and neatly arrange X, y, labels
FIX idx(s)/ind(s)--> indice(s)
COSMIT Merge if and else to elif
COSMIT n --> n_samples
COSMIT Use bincount only once
COSMIT cls --> class_i / class_i (ith class indices) -->
perm_indices_class_i
FIX/ENH/TST Addressing the final reviews
COSMIT c --> count
FIX/TST make check_cv raise ValueError for string cv value
TST nested cv (gs inside cross_val_score) works for diff cvs
FIX/ENH Raise ValueError when labels is None for label based cvs;
TST if labels is being passed correctly to the cv and that the
ValueError is being propagated to the cross_val_score/predict and grid
search
FIX pass labels to cross_val_score
FIX use make_classification
DOC Add Returns; COSMIT Remove scaffolding
TST add a test to check the _build_repr helper
REVERT the old GS/RS should also be tested by the common tests.
ENH Add a tuple of all/label based CVS
FIX raise VE even at get_n_splits if labels is None
FIX Fabian's comments
PEP8