In order to fix#11408, this swaps `joblib` and `_joblib`. It however, allows users to access joblib's `Memory` or `Parallel` functionality without accessing `sklearn.externals._joblib` by importing `Memory`, `Parallel`, etc. into `sklearn.utils`.
MissingIndicator transformer for the missing values indicator mask.
see #6556
#### What does this implement/fix? Explain your changes.
The current implementation returns a indicator mask for the missing values.
#### Any other comments?
It is a very initial attempt and currently no tests are present. Please do have a look and give suggestions on the design. Thanks !
- [X] Implementation
- [x] Documentation
- [x] Tests
* OPTICS clustering algorithm
Equivalent results to DBSCAN, but allows execution on arbitrarily large datasets. After initial construction, allows multiple 'scans' to quickly extract DBSCAN clusters at variable epsilon distances
* Create plot_optics
Shows example usage of OPTICS to extract clustering structure
* pep8 fixes
Mainly issues in long lines, long comments
* fixed conditional to be pep8
* updated to match sklearn API
OPTICS object, with fit method. Documentation updates.
extract method
* removed extra files
* plotting example updated, small changes
new plot example that matches the updated API
updated n_cluster attribute with reruns of extract
removed scaling factor on first ‘fit’ run
* updated OPTICS.labels to OPTICS.labels_
should pass unit test now?
* additional labels_ changes
* added stability warning
Scales eps to give stable results from first input distance. Extraction
above scaled eps is not allowed; extraction below scaled eps but
greater than input distance prints stability warning but still will run
clustering extraction. All distances below initial input distance are
stable; no warning printed
* Noise fix; updated example plot
Fixed noise points from being initialized as type ‘core point’
Fixed initialization for first ‘fit’ call
Decoupled eps and eps_prime (deep copy)
Matched plot example to same random state as dbscan
Added second plot to show ‘extract’ usage
* Changed to match Sklearn API
eps is not modified in init; kwargs fix
* Forcing 2 parameters
Why do I have to do this?
* Conforming to API
Fit now returns self; test fix for unit test fail in ball tree (asarray
problem)
* Fixing plot example
labels to labels_
* Fixed issue with sparse matrices
* Another attempt at fixing the sparse matrix error
Temporary fix until balltree can be updated to deal with sparse
matrices.
* Better checking of sparse arrays
Using ‘check_array’
* General cleanup
Removed extraneous lines and comments (old commented out code has been
removed)
* Added unit tests for extract function
Added the unit tests in test_optics for fit, extract, and declaration.
Should bring coverage to ~100%. Additionally, fixed a small bug that
cropped up in the extract function during testing.
* Attempting for near 100% coverage
Removed unused imports, added to get warning.
* Fixed error in unit tests
* Trimmed extraneous 'if-else' check
see title
* forcing to check for a warning.
Result should be 100% coverage
* Updates to doc strings
All public methods now have doc strings; forcing a rebuild of OPTICS so
that build tests pass (last round failed due to external module)
* Style / pep8 changes
99% pep8 now… line 138 isn’t, but reads better with the long variable
names
* Added Narrative Documentation
Includes general description, discussion of algorithm output, and
comparison with DBSCAN. References and implementation notes are included
* Vectorized nneighbors lookups
Following suggestion from jnothman for doing nneighbors queries enmass.
Added OPTICS to cluster init
* fixing init build error
* reverting init
* All code now vectorized
…at least all code that can be ;)
Some general pruning and cleanup as well
* Style changes
Now 100% pep8
* Changing parameter style
matching DBSCAN
* Extraction change; Authors update
—initialize all points as ‘not core’ (fixes bug when plotting at
epsPrime larger than eps)
—added Sean Freeman to authors list
* Changed eps scaling to 5x instead of 10x
10x scaling is too conservative…eps scaling at 5x is perfectly stable,
and much faster as well.
* Fixing unit test
Should be ‘None’ for this; initialization previously at 1 for ‘is_core’
was incorrect
* Actually fixing unit test
Null comparison doesn’t work, using size
* Making ordering_ and other attributes public
renamed core_samples to core_sample_indices_ (as in DBSCAN). Used
attribute naming conventions (trailing _ character), and made
ordering_, reachability_, and core_dists_ public.
* Updates for Documentation
includes attribute strings for now public attributes
* Pep8 cleanup
Minor pep8 and pyflakes fixes
* updating plot example to match new attribute name
* CamelCase fixes
conforming to sklearn API on CamelCase
* adding hierarchical extraction function
Added hierarchical extraction from #2043#2043 is BSD licensed and hasn’t had any activity for 10 months, so
this seems pretty kosher; authors are cited in the code
* added hierarchical switch to extract
Additional style and documentation changes as well
* cluster order bug fix
ensured that ordered list is always same length as input data
* removed hierarchical cluster extraction
Code from FredrikAppelros is totally unable to handle noise— as
currently written every point is assigned to a cluster, except the
first point of that cluster. May include later as a third method
* initial import of automatic cluster extraction code
Adding the excellent (and working) automatic cluster extraction code
from Amy X. Zhang. Some minor formatting changes on import to conform
to pep8
* wrapper for 'auto' extraction
additional fixes to style (camelCase, etc.), comments; made all helper
functions private
* test and example updates
Much better example data to showcase ‘auto’ extraction. Unit tests now
test both extraction methods. Set ‘auto’ as default, as it doesn’t
require any parameters and gives a better result. Pruned references to
hierarchical clustering
* fixing unit coverage; pruning unused functions
probably could still get a better unit test for ‘auto’ clustering…
* Added 'filter' fuction
Allows density-based filtering. Useful for cases where only a
‘noise’/‘no noise’ classification is desired. Function does not require
fitting prior to running, although it can be run after a fit if desired
* Vectorizing auto_cluster
generalizes input to multiple dimensions as well…
* updated filter function
* fixing test error
setting minPts to < data size returns None (with error message)
* removing exception handling in favor of conditional check
* Updated unit tests
Coverage to 90%. PEP8 fixes
* Additional unit test
Now at 94%; auto extract method is very hard to test… this new unit
test adds a more robust dataset to trigger more branches for testing.
Some of the remaining conditionals are pretty rare … :-/
* Fix unit test bug / python 3 compat
None type comparison problem…
* Fixed annoying deprecation warning for 1d array in BallTree
PEP8 Fixes
* More PEP8 and remove print statements fromt est
* 70 to 80% faster, fixed distance metrics
Modified to remove extraneous sorts, nneighbors query, and reduced
pairwise distance calculations to the upper triangle of a distance
matrix (instead of the full matrix). It appears that in ‘full’ scans
(i.e., when epsilon is set to inf, or the width of the data set),
OPTICS now actually outperforms DBSCAN… also, should be easy to run
distance calculations in parallel now for large datasets, with proper
heuristics.
* Fixing unit test failure
* Exposed auto_cluster parameters as public
documentation and API update so that users can tweek the auto method
* Fixed def with missing ':'
* Fix bugs from api change...
make sure that arguments are being called correctly with the new
extract_auto() method
* pep8 / pyflakes changes
* Updated example / plot
Correctly generates figure with subplots for DBSCAN / OPTICS comparison
* Tuning plot example / pep8 change
* Bug fix for commit 0b4cbdd (enforce stable sort)
the returned index from “sp.argmin(setofobjects.reachability_[n_pr])”
assumes that entires are stably sorted by distance from query point
(i.e., that ties in the argmin return the closest point to the input
point).
Still overall faster compared to the pre 0b4cbdd version, since we’re
filtering processed points before sorting by distance… and also only
calculating distances for non-processed points, instead of all within
the epsilon query.
* Code review fixes (style)
Fixes coding style (test comments, camel case, author list, relative
imports, etc). Included a new unit test and .npy file to explicitly
test reachability distances (test coverage at 99%).
* fixing new unit test
can’t import reach_values.npy in current directory….just placed testing
values in script directly (~200 lines for 1500 testing values)
* refactored min_heap to c extension
reduced optics file by about 100 lines of code… new c extension for
speedup (needs further optimization…6-10X speedup still possible)
* minor fixes..
* small cython optimizations
* cython fixes
* Add foo.txt
* Remove foo.txt
* fix compilation error
* last optimizations cython/numpy
MinHeap is only called once…so it’s faster to do a simple linear scan
here. The np.argmin() function does *almost* the exact same thing, but
the custom quick scan function is needed for cases where reachability
distances are tied (and the next point is selected based on which of
the points tied in reachability are closest to the querying point).
OPTICS is now faster than DBSCAN for medium-to-large number of input
points, and has better worst case run time (eps=inf).
* fix pyflakes errors; change default eps value
* _
* API changes from agramfort
Removed ‘filter method’, changed print statements to exceptions and
warnings as needed, remove ‘processed’ flag and replaced with fit_check
method, changed array copies to parameters, updated unit tests, changed
inline documentation to match proper doc string formatting, made
private class private. Probably some other changes too…
* fixed core samples bug
* added fit_predict
conforming to scipy api
* updates to variable names; update plot
* refactor to remove balltree specific code
* major refactor
finished decoupling balltree; lots of changes
* fixed bugs; test all pass again
:-)
* fixed weird cython bug
…not at all sure why pairwise_distances doesn’t automatically return
np.float arrays. Makes no sense to me. I could understand cases in
which the metric call returns int’s (i.e., city block)… but why it
would return float.32 instead of float.64 seems super odd.
* major refactor
deleted extract and auto_extract methods; added optics function; added
extract_dbscan method; added extract_dbscan and extract_optics
functions; updated unit tests; renamed `eps` to `max_bounds`; flake8
corrections; added types: int —> labels, bool —> is_core, enforced X to
be ‘float’ (fixes cython type errors)
* Updated Documentation!
Updated plot with reachability plot, as well as documentation :)
* fix flake8 error
* added optics to cluster comparison
Don't like the figure since it doesn't use black for noise, but added OPTICS for consistency
* Updated comparison plot to transpose
...kinda kludgy fix for the transpose
* small fix
* flake8 error
* reverting transpose of cluster comparison (seperate PR #9739)
* fixes from agramfort's review
public/private changes, numpydocstring fixes, a few pep8 fixes that flake8 didn't flag for some reason...
* fix for error message
unit test should pass now
* force cluster_id's to start at 0
* Fix sp. error and flake8 warning(s)
* Updated documentation
responses to reviews
* Removed extraneous files
also fixed small typo
* fixing lgtm alert
* changes from jnothman
small fixes for docs, plots, and tests
* Fixes from jnothman's review
Reorded parameters, updated documentation, removed unneeded else statement, changes to varible names.
* Fixing flake8 error
* Removed neighbors / balltree inheritance
Also decouples n_jobs -- can set n_jobs for just the kneighbors lookup, while keeping pairwise lookups to single job
* Made nbrs private and moved initiation to fit()
also renamed core_dists to core_distances.
* fixed non-standard characters
* Response to TomDLT review
narrative changes to tests. Normalized reachability distances for significant_min parameter. Cleaned up plots; small changes to documentation with optics_.py
* Fixed labeling bug
also minor documentation updates
* update unit test
since labels are 0 indexed, max of labels is (1 - total number of clusters). We can't take len(set(clusters)) because of noise (will be 4, not 3).
* Simple fixes per jnothman
Fixed float division, condensed variable names to be shorter, renamed bools to True and False, removed un-needed code block. Still need to add tests for cluster tree extraction :-(
* Auto-cluster tests
coverage should be complete now; fixed minor bug; removed un-needed check.
* fixing test error
* removed python loop
also fixed documentation link
* Fixing test error
* Fix typo in unit test
entry was supposed to be '1.0' not '10'; the test is supposed to posit 3 clusters, 2 of which are too small and are merged. Old version posited 4 clusters, two of which were merged as intended, and two of which were discarded (cluster merging requires one of the clusters to be large enough to be an independent cluster; with 4 instead of three, this case did not happen for either of the first two clusters).
* documentation updates
* Post-merge doctest fix
merge conflict in clustering.rst in previous commit; this push updates the doctest values to current correct values, and resolves the conflict
* DBSCAN / OPTICS invariant test
Restructured documentation. Small unit test fixes. Added test to ensure clustering metrics between OPTICS dbscan_extract and DBSCAN are within 2% of each other.
* Update _auto_cluster docstring
renamed reachability_ordering --> ordering for consistency
* changes fro jnothman
changes unit tests to check for specific error message (instead of 'a' error); minor updates to documentation. This also fixes a bug in the extract_dbscan function whereby some core points were erroneously marked as periphery... this is fixed by reverting to a previous extraction code block that initalizes all points as core and then demotes noise and periphery points during the extract scan. Parameterized unit test.
* fix spelling error in tests
* contingency_matrix test
New invarient test between optics and dbscan
* small unit test updates per jnothman
* unit test typo fix
* extract dbscan updates
Vectorized extract dbscan function to see if would improve performance of periphery point labeling; it did not, but the function is vectorized. Changed unit test with min_samples=1 to min_samples=3, as at min_samples=1 the test isn't meaningful (no noise is possible, all points are marked core). Parameterized parity test. Changed parity test to assert ~5% or better mismatch, instead of 5 points (this is needed for larger clusters, as the starting point mismatch effect scales with cluster size).
* updated documentation comparing OPTICS/DBSCAN
* DOC: phrasing and whats_new
* MISC: small mem footprint in OPTICS
Skeleton for a glossary of concepts and API elements.
This responds to at least three issues:
* Many aspects of scikit-learn API for users and developers are known
tacitly by core contributors (and the stack overflow crowd), but are
not written down in a consistent place.
* What is written is in an ad-hoc narrative style which may be useful
for introduction, but is difficult to refer to and to maintain.
* Parameters such as `n_jobs` and methods like `decision_function` are
described repeatedly in documentation giving sometimes more sometimes
less information. This glossary allows us to use "See :term:`the
glossary <n_jobs>`." so that parameter descriptions in the
API reference can remain brief (just as not every numpy operation
needs to describe broadcasting).
* Update example to clarify difference with Confusion Matrix, added link to Confusion Matrix documentation
* fixed 80 character limit issue
* changed indentation for automated test
* fixed formatting
* fixed formatting
* fixed formatting
* address requested changes in Contingency Matrix docs
* basic formatting
* use less technical and correct definition of contingency matrix
* arrange in alphabetical order
* add 'advantages' and 'disadvantages' of contingency matrix
* use better 'str' labels for example
* address the language of documentation
* add `ref` in documentation
* add a better 'advantage'
* fix
* address review
* small change
* add full stop.
* add function computing balanced accuracy
* documentation for the balanced_accuracy_score
* apply common tests to balanced_accuracy_score
* constrained to binary classification problems only
* add balanced_accuracy_score for CLF test
* add scorer for balanced_accuracy
* reorder the place of importing balanced_accuracy_score to be consistent with others
* eliminate an accidentally added non-ascii character
* remove balanced_accuracy_score from METRICS_WITH_LABELS
* eliminate all non-ascii charaters in the doc of balanced_accuracy_score
* fix doctest for nonexistent scoring function
* fix documentation, clarify linkages to recall and auc
* FIX: added changes as per last review See #6752, fixes#6747
* FIX: fix typo
* FIX: remove flake8 errors
* DOC: merge fixes
* DOC: remove unwanted files
* DOC update what's new
* DOC cleaning up what's new for 0.19
* More cleaning up
* More cleaning up
* Deprecations
* Clean up merge
* Update
* TODOs to prose and minor changes
* Changed models and minor fixes
* sort
* Merge in 0.18.2 docs
* Missing entry from 0.18 logs
* Optimistically add some features to highlights
* Forgotten user directive
* Fix alignment
* Cleaning up for Andy's comments
* Mention beta_loss=0 speedup
* Update
* Clean up new what's new entries
* DOC Add changes missed from what's new
And other minor things.
This took lots of effort which I would have not committed where I not home sick...
* ENH cross_val_score now supports multiple metrics
* DOCFIX permutation_test_score
* ENH validate multiple metric scorers
* ENH Move validation of multimetric scoring param out
* ENH GridSearchCV and RandomizedSearchCV now support multiple metrics
* EXA Add an example demonstrating the multiple metric in GridSearchCV
* ENH Let check_multimetric_scoring tell if its multimetric or not
* FIX For single metric name of scorer should remain 'score'
* ENH validation_curve and learning_curve now support multiple metrics
* MNT move _aggregate_score_dicts helper into _validation.py
* TST More testing/ Fixing scores to the correct values
* EXA Add cross_val_score to multimetric example
* Rename to multiple_metric_evaluation.py
* MNT Remove scaffolding
* FIX doctest imports
* FIX wrap the scorer and unwrap the score when using _score() in rfe
* TST Cleanup the tests. Test for is_multimetric too
* TST Make sure it registers as single metric when scoring is of that type
* PEP8
* Don't use dict comprehension to make it work in python2.6
* ENH/FIX/TST grid_scores_ should not be available for multimetric evaluation
* FIX+TST delegated methods NA when multimetric is enabled...
TST Add general tests to GridSearchCV and RandomizedSearchCV
* ENH add option to disable delegation on multimetric scoring
* Remove old function from __all__
* flake8
* FIX revert disable_on_multimetric
* stash
* Fix incorrect rebase
* [ci skip]
* Make sure refit works as expected and remove irrelevant tests
* Allow passing standard scorers by name in multimetric scorers
* Fix example
* flake8
* Address reviews
* Fix indentation
* Ensure {'acc': 'accuracy'} and ['precision'] are valid inputs
* Test that for single metric, 'score' is a key
* Typos
* Fix incorrect rebase
* Compare multimetric grid search with multiple single metric searches
* Test X, y list and pandas input; Test multimetric for unsupervised grid search
* Fix tests; Unsupervised multimetric gs will not pass until #8117 is merged
* Make a plot of Precision vs ROC AUC for RandomForest varying the n_estimators
* Add example to grid_search.rst
* Use the classic tuning of C param in SVM instead of estimators in RF
* FIX Remove scoring arg in deafult scorer test
* flake8
* Search for min_samples_split in DTC; Also show f-score
* REVIEW Make check_multimetric_scoring private
* FIX Add more samples to see if 3% mismatch on 32 bit systems gets fixed
* REVIEW Plot best score; Shorten legends
* REVIEW/COSMIT multimetric --> multi-metric
* REVIEW Mark the best scores of P/R scores too
* Revert "FIX Add more samples to see if 3% mismatch on 32 bit systems gets fixed"
This reverts commit ba766d98353380a186fbc3dade211670ee72726d.
* ENH Use looping for iid testing
* FIX use param grid as scipy's stats dist in 0.12 do not accept seed
* ENH more looping less code; Use small non-noisy dataset
* FIX Use named arg after expanded args
* TST More testing of the refit parameter
* Test that in multimetric search refit to single metric, the delegated methods
work as expected.
* Test that setting probability=False works with multimetric too
* Test refit=False gives sensible error
* COSMIT multimetric --> multi-metric
* REV Correct example doc
* COSMIT
* REVIEW Make tests stronger; Fix bugs in _check_multimetric_scorer
* REVIEW refit param: Raise for empty strings
* TST Invalid refit params
* REVIEW Use <scorer_name> alone; recall --> Recall
* REV specify when we expect scorers to not be None
* FLAKE8
* REVERT multimetrics in learning_curve and validation_curve
* REVIEW Simpler coding style
* COSMIT
* COSMIT
* REV Compress example a bit. Move comment to top
* FIX fit_grid_point's previous API must be preserved
* Flake8
* TST Use loop; Compare with single-metric
* REVIEW Use dict-comprehension instead of helper
* REVIEW Remove redundant test
* Fix tests incorrect braces
* COSMIT
* REVIEW Use regexp
* REV Simplify aggregation of score dicts
* FIX precision and accuracy test
* FIX doctest and flake8
* TST the best_* attributes multimetric with single metric
* Address @jnothman's review
* Address more comments \o/
* DOCFIXES
* Fix use the validated fit_param from fit's arguments
* Revert alpha to a lower value as before
* Using def instead of lambda
* Address @jnothman's review batch 1: Fix tests / Doc fixes
* Remove superfluous tests
* Remove more superfluous testing
* TST/FIX loop over refit and check found n_clusters
* Cosmetic touches
* Use zip instead of manually listing the keys
* Fix inverse_transform
* FIX bug in fit_grid_point; Allow only single score
TST if fit_grid_point works as intended
* ENH Use only ROC-AUC and F1-score
* Fix typos and flake8; Address Andy's reviews
MNT Add a comment on why we do such a transpose + some fixes
* ENH Better error messages for incorrect multimetric scoring values +...
ENH Avoid exception traceback while using incorrect scoring string
* Dict keys must be of string type only
* 1. Better error message for invalid scoring 2...
Internal functions return single score for single metric scoring
* Fix test failures and shuffle tests
* Avoid wrapping scorer as dict in learning_curve
* Remove doc example as asked for
* Some leftover ones
* Don't wrap scorer in validation_curve either
* Add a doc example and skip it as dict order fails doctest
* Import zip from six for python2.7 compat
* Make cross_val_score return a cv_results-like dict
* Add relevant sections to userguide
* Flake8 fixes
* Add whatsnew and fix broken links
* Use AUC and accuracy instead of f1
* Fix failing doctests cross_validation.rst
* DOC add the wrapper example for metrics that return multiple return values
* Address andy's comments
* Be less weird
* Address more of andy's comments
* Make a separate cross_validate function to return dict and a cross_val_score
* Update the docs to reflect the new cross_validate function
* Add cross_validate to toc-tree
* Add more tests on type of cross_validate return and time limits
* FIX failing doctests
* FIX ensure keys are not plural
* DOC fix
* Address some pending comments
* Remove the comment as it is irrelevant now
* Remove excess blank line
* Fix flake8 inconsistencies
* Allow fit_times to be 0 to conform with windows precision
* DOC specify how refit param is to be set in multiple metric case
* TST ensure cross_validate works for string single metrics + address @jnothman's reviews
* Doc fixes
* Remove the shape and transform parameter of _aggregate_score_dicts
* Address Joel's doc comments
* Fix broken doctest
* Fix the spurious file
* Address Andy's comments
* MNT Remove erroneous entry
* Address Andy's comments
* FIX broken links
* Update whats_new.rst
missing newline
* Add deprecation message and test.
* Adding performance warning and ignore_warnings in test
* Add deprecation to whatsnew and remove LSHForest references from docs.
Removing benchmark for lsh
* resurrect quantile scaler
* move the code in the pre-processing module
* first draft
* Add tests.
* Fix bug in QuantileNormalizer.
* Add quantile_normalizer.
* Implement pickling
* create a specific function for dense transform
* Create a fit function for the dense case
* Create a toy examples
* First draft with sparse matrices
* remove useless functions and non-negative sparse compatibility
* fix slice call
* Fix tests of QuantileNormalizer.
* Fix estimator compatibility
* List of functions became tuple of functions
* Check X consistency at transform and inverse transform time
* fix doc
* Add negative ValueError tests for QuantileNormalizer.
* Fix cosmetics
* Fix compatibility numpy <= 1.8
* Add n_features tests and correct ValueError.
* PEP8
* fix fill_value for early scipy compatibility
* simplify sampling
* Fix tests.
* removing last pring
* Change choice for permutation
* cosmetics
* fix remove remaining choice
* DOC
* Fix inconsistencies
* pep8
* Add checker for init parameters.
* hack bounds and make a test
* FIX/TST bounds are provided by the fitting and not X at transform
* PEP8
* FIX/TST axis should be <= 1
* PEP8
* ENH Add parameter ignore_implicit_zeros
* ENH match output distribution
* ENH clip the data to avoid infinity due to output PDF
* FIX ENH restraint to uniform and norm
* [MRG] ENH Add example comparing the distribution of all scaling preprocessor (#2)
* ENH Add example comparing the distribution of all scaling preprocessor
* Remove Jupyter notebook convert
* FIX/ENH Select feat before not after; Plot interquantile data range for all
* Add heatmap legend
* Remove comment maybe?
* Move doc from robust_scaling to plot_all_scaling; Need to update doc
* Update the doc
* Better aesthetics; Better spacing and plot colormap only at end
* Shameless author re-ordering ;P
* Use env python for she-bang
* TST Validity of output_pdf
* EXA Use OrderedDict; Make it easier to add more transformations
* FIX PEP8 and replace scipy.stats by str in example
* FIX remove useless import
* COSMET change variable names
* FIX change output_pdf occurence to output_distribution
* FIX partial fixies from comments
* COMIT change class name and code structure
* COSMIT change direction to inverse
* FIX factorize transform in _transform_col
* PEP8
* FIX change the magic 10
* FIX add interp1d to fixes
* FIX/TST allow negative entries when ignore_implicit_zeros is True
* FIX use np.interp instead of sp.interpolate.interp1d
* FIX/TST fix tests
* DOC start checking doc
* TST add test to check the behaviour of interp numpy
* TST/EHN Add the possibility to add noise to compute quantile
* FIX factorize quantile computation
* FIX fixes issues
* PEP8
* FIX/DOC correct doc
* TST/DOC improve doc and add random state
* EXA add examples to illustrate the use of smoothing_noise
* FIX/DOC fix some grammar
* DOC fix example
* DOC/EXA make plot titles more succint
* EXA improve explanation
* EXA improve the docstring
* DOC add a bit more documentation
* FIX advance review
* TST add subsampling test
* DOC/TST better example for the docstring
* DOC add ellipsis to docstring
* FIX address olivier comments
* FIX remove random_state in sparse.rand
* FIX spelling doc
* FIX cite example in user guide and docstring
* FIX olivier comments
* EHN improve the example comparing all the pre-processing methods
* FIX/DOC remove title
* FIX change the scaling of the figure
* FIX plotting layout
* FIX ratio w/h
* Reorder and reword the plot_all_scaling example
* Fix aspect ratio and better explanations in the plot_all_scaling.py example
* Fix broken link and remove useless sentence
* FIX fix couples of spelling
* FIX comments joel
* FIX/DOC address documentation comments
* FIX address comments joel
* FIX inline sparse and dense transform
* PEP8
* TST/DOC temporary skipping test
* FIX raise an error if n_quantiles > subsample
* FIX wording in smoothing_noise example
* EXA Denis comments
* FIX rephrasing
* FIX make smoothing_noise to be a boolearn and change doc
* FIX address comments
* FIX verbose the doc slightly more
* PEP8/DOC
* ENH: 2-ways interpolation to avoid smoothing_noise
Simplifies also the code, examples, and documentation