Commit Graph

54 Commits

Author SHA1 Message Date
Thomas J. Fan 21121f5099
TST Fixes column transformer tests (#18452) 2020-09-30 12:39:05 +02:00
Thomas J. Fan 2b655efaf2
FIX Adds remainder in column transformer repr_html (#18167) 2020-08-19 18:40:34 +10:00
Thomas J. Fan 3a796a5873
ENH Adds support for list of bools in column transformer (#17616) 2020-08-07 14:11:05 +02:00
Juan Carlos Alfaro Jiménez 547ead6c0c
MNT More replacements of numpy aliases with built-in types (#17707)
* MNT More replacements of numpy aliases with built-in types [scipy-dev]

* MNT More (manual) replacements of numpy aliases

* MNT More (manual) replacements of numpy aliases

* FIX Minor change

* MNT Trigger build [scipy-dev]

* FIX Minor change

* MNT Trigger build [scipy-dev]

* FIX More deprecations of numpy aliases [scipy-dev]

* MNT Silent numpy aliases warnings [scipy-dev]

* FIX Fix silent deprecations

* Trigger build [scipy-dev]

* FIX Add scape characters [scipy-dev]

* FIX Remove package specification from silent

* Trigger build [scipy-dev]

* Revert
2020-06-26 12:45:03 +02:00
Thomas J. Fan 7cc0177f8e
MNT Replaces numpy alias with builtin typse (#17687)
* MNT Replaces numpy alias with builtin typse

* STY Lint error
2020-06-24 16:51:51 +02:00
Jake Tae 493f167c76
Replace *kwargs with named arguments in make_column_transformer (#17623) 2020-06-17 21:09:13 +02:00
lrjball 670b85c9e9
ENH ColumnTransformer.get_feature_names() handles passthrough (#14048) 2020-04-20 01:24:20 +10:00
lrjball 91d5ac882c
TST Fixes test so that whole test isn't skipped if pandas not… (#16627) 2020-03-03 20:03:57 -05:00
Nicolas Hug d205638475
MNT Introduction of n_features_in_ attr with _validate_data mtd (#16112) 2020-02-29 09:05:11 -05:00
Christian Lorentzen 21686b7147
ENH Sample weights for ElasticNet (#15436) 2020-02-16 14:48:01 +01:00
Roman Feldbauer 9b39c4c4d2
TST Fix unreachable code in tests (#16110) 2020-02-16 14:41:26 +01:00
Roman Yurchak 528b044a22
FIX ColumnTransformer.get_feature_names with for empty list of (#15963)
columns
2020-01-31 14:05:07 +01:00
Thomas J Fan 37ac3fd125 FEA Add make_column_selector for ColumnTransformer (#12371) 2019-11-05 09:22:26 -05:00
Nicolas Hug 19ad136223 MNT Replace DeprecationWarning with FutureWarning (#15080)
* bruteforce renaming

* WIP

* WIP

* some more

* removed weird line

* update -Werror

* testiforest

* again

* Fixed some tests

* fixed some tests

* removed -Werror

* fixed test_docstring_param issue

* fixed some tests

* some more

* renamed to SklearnDeprecationWarning

* pep8

* updated docs

* pep8

* merge

* changed to FutureWarning

* fixes

* Update doc/developers/tips.rst

Co-Authored-By: Adrin Jalali <adrin.jalali@gmail.com>

* avoid duplicates

* fixed warning for deprecations

* Still make CI break if DeprecationWarning isn't caught

* updated one warning

* fixed test

* fixed some renamings

* updated new dep warnings

* fixed bad import

* ignore warnings

* update again

* ignore futurewarning when walking packages

* Fixed test

* pep8

* Added whatsnew
2019-10-29 15:39:26 +01:00
Nicolas Hug b92455a6b2 MAINT Deprecate all of utils.testing except all_estimators (#15367) 2019-10-28 17:28:56 +01:00
Miguel Cabrera a1f514f2e1 [MGR] TransformedTargetRegressor passes fit_params to regressor (#14890) 2019-09-06 12:26:04 +02:00
Andreas Mueller 03ea20db0f Fix mixin inheritance order, allow overwriting tags (#14884) 2019-09-05 10:13:30 +02:00
Samesh Lakhotia db23c4ece0 MAINT Fix assert raises in sklearn/compose/tests/ (#14670) 2019-08-17 01:24:08 -04:00
Adrin Jalali da6614fe42 ColumnTransformer input feature name and count validation (#14544) 2019-08-07 12:19:13 -04:00
Guillaume Lemaitre 992ed41ecc FIX change boolean array-likes indexing in old NumPy ver… (#14510) 2019-08-02 10:46:16 -04:00
Andreas Schuderer 9115ab0ed3 FIX ColumnTransformer: raise error on reordered columns with remainder (#14237)
* FIX Raise error on reordered columns in ColumnTransformer with remainder

* FIX Check for different length of X.columns to avoid exception

* FIX linter, line too long

* FIX import _check_key_type from its new location utils

* ENH Adjust doc, allow added columns

* Fix comment typo as suggested, remove non-essential exposition in doc

* Add PR 14237 to what's new

* Avoid AttributeError in favor of ValueError "column names only for DF"

* ENH Add check for n_features_ for array-likes and DataFrames

* Rename self.n_features to self._n_features

* Replaced backslash line continuation with parenthesis

* Style changes
2019-07-22 15:36:46 +02:00
Adrin Jalali 19c068a2ec MNT towards removing assert_equal, etc (#14222) 2019-07-01 09:13:32 -04:00
lrjball 0c110701f9 TST Fixed typo in test_column_transformer (#14128)
In test_column_transformer_dataframe(), X_df2 was created but never used. X_df was mistakely being used in its place. The PR fixes that typo.
2019-06-20 10:43:19 +10:00
Guillaume Lemaitre 76ce7c5b63 DEP change default validate in FunctionTransformer to False(#13817) 2019-06-13 12:06:01 -04:00
Guillaume Lemaitre 9ee164baa3 [MRG] DEP remove legacy mode from OneHotEncoder (#13855) 2019-05-29 21:03:29 +10:00
Guillaume Lemaitre 7fdac52c19 MNT Remove backward compatibility of param order in make_column_transformer (#13831) 2019-05-09 20:54:30 +08:00
Thomas J Fan 8edd9f9f15 ENH Add verbose option to Pipeline, FeatureUnion, and ColumnTransformer (#11364) 2019-04-22 00:51:21 +10:00
Andreas Mueller 8d815af732
MRG don't warn on changing dtypes in scalers (#13306)
* don't warn on changing dtypes in scalers

* remove tests for dtype warnings

* len(False) > len(True)

pep8
2019-02-28 09:42:50 +01:00
Guillaume Lemaitre c5075f5e46 MNT do not call fit twice in TransformedTargetetRegressor (#11641) 2019-02-03 10:07:44 +08:00
pierretallotte ce9dedb1a3 FIX Convert the negative indices to positive ones in ColumnTransformer._get_column_indices (#13017) 2019-01-22 22:22:40 +08:00
Rüdiger Busche d300f406ae MAINT Simplify super() calls (#12812) 2019-01-10 22:27:06 +01:00
Andreas Mueller 952ef6637a MRG Drop legacy python / remove six dependencies (#12639) 2019-01-03 15:50:05 +02:00
Hanmin Qin aae4e337d2 FIX Fix error in make_column_transformer when columns is pandas.Index (#12704)
Fixes #12703
2018-12-03 21:22:32 +11:00
Bartosz Michałowski fa98a72dcc MNT Replaced all occurrences of assert_true and assert_false with assert (#12588) 2018-11-28 09:16:26 +08:00
Adrin Jalali a816af7d1d API Change the tuple order in make_column_transformer (#12626) 2018-11-21 11:07:12 +08:00
Hanmin Qin 43e3a02085
MNT Remove unused assert_true imports (#12560) 2018-11-11 11:08:37 +08:00
Yaroslav Halchenko 362cb3bcab TST autoreplace assert_true(...==...) with plain assert (#12547) 2018-11-11 09:05:34 +08:00
Andreas Mueller 0ad873646b [MRG] ENH apply sparse_threshold even if all columns are sparse (#12304) 2018-10-14 17:18:23 +11:00
Adrin Jalali 58228cb6d3 [MRG+1] ColumnTransformer fix having mixed types in a single passthrough (#12200)
* fix having mixed types in a single passthrough

* use check_array

* add comment

* Reference the relevant issue in the test

* don't force finite on check_array

* add the whats_new entry

* modify whats_new entry for a [probably] better description.

* raise a value error if columns can't be stacked.

* test the custom ValueError

* take sparse.hstack out of try/catch

* improve whats_new entry

* add comment for the dtype conversion

* change the tests to use an np of dtype 'O' instead
2018-10-04 09:37:39 -04:00
Joris Van den Bossche 09851aca9e TST update make_column_transformer test + add comment (#12156)
Follow-up on https://github.com/scikit-learn/scikit-learn/pull/12152
And added comment why transformer_weights is not passed through, see
https://github.com/scikit-learn/scikit-learn/pull/11183#pullrequestreview-125539051
for more discussion
2018-09-25 23:08:21 +08:00
janvanrijn e58f366e03 ColumnTransformer generalization to work on empty lists (#12084) 2018-09-25 17:03:32 +02:00
Jan Koch 661a8b42bd add sparse_threshold to make_columns_transformer (#12152)
fixes issue #12149
2018-09-25 10:43:05 -04:00
Vinayak Mehta f15ebb953b [MRG] Convert ColumnTransformer input list to numpy array (#12104)
<!--
Thanks for contributing a pull request! Please ensure you have taken a look at
the contribution guidelines: https://github.com/scikit-learn/scikit-learn/blob/master/CONTRIBUTING.md#pull-request-checklist
-->

#### Reference Issues/PRs
<!--
Example: Fixes #1234. See also #3456.
Please use keywords (e.g., Fixes) to create link to the issues or pull requests
you resolved, so that they will automatically be closed when your pull request
is merged. See https://github.com/blog/1506-closing-issues-via-pull-requests
-->
Fixes #12096.

#### What does this implement/fix? Explain your changes.
Converts the input list for ColumnTransformer to a numpy array.

Added a check inside `transform` and `fit_transform` to check if the input `X` is a list, if it is then it gets converted to a numpy array.

#### Any other comments?
Should this conversion be documented in the docstrings for ColumnTransfomer's `fit`, `transform` and `fit_transform`?

<!--
Please be aware that we are a loose team of volunteers so patience is
necessary; assistance handling other issues is very welcome. We value
all user contributions, no matter how minor they are. If we are slow to
review, either the pull request needs some benchmarking, tinkering,
convincing, etc. or more likely the reviewers are simply busy. In either
case, we ask for your understanding during the review process.
For more information, see our FAQ on this topic:
http://scikit-learn.org/dev/faq.html#why-is-my-pull-request-not-getting-any-attention.

Thanks for contributing!
-->
2018-09-25 10:37:16 -04:00
Joris Van den Bossche 4035e60a6f [MRG +1] ColumnTransformer: store evaluated function column specifier during fit (#12107) 2018-09-21 13:22:34 +02:00
Olivier Grisel 40e6c43cb4
Joblib 0.12.2 (#11741)
* joblib 0.12.2

* Export _joblib's register_parallel_backend

* Use latest version of coverage
2018-08-03 12:34:25 +02:00
Joris Van den Bossche cf897de0ab ENH Add sparse_threshold keyword to ColumnTransformer (#11614)
Reasoning: when eg OneHotEncoder is used as part of ColumnTransformer it would cause the final result to be a sparse matrix. As this is a typical case when we have mixed dtype data, it means that many pipeline will have to deal with sparse data implicitly, even if you have only some categorical features with low cardinality. 
Idea was first to change default of `OneHotEncoder` sparse to False, but based on gitter discussion (https://gitter.im/scikit-learn/dev?at=5b4e5a69a94c5255523bc9fc) we decided to let ColumnTransformer switch between both based on a threshold. The user still has full control if he/she wants always or never sparse.
2018-07-25 22:03:04 +10:00
Thomas Fan 06ac22d06f BUG: Fixes _BaseCompostion._set_params broken where there are no estimators (#11333) 2018-07-20 04:13:55 -05:00
Joris Van den Bossche 277e1250c2 Change default of ColumnTransformer remainder from passthrough to drop (#11603) 2018-07-18 09:57:15 +02:00
Joris Van den Bossche 9caa982605 [MRG + 1] ENH: allow to pass callable as column specifier in ColumnTransformer (#11592)
* define select_types callable factory

* include dtype column selector example

* support selector function for remainder=passthrough (already worked for drop)

* add docstring to example file. apparently sphinx fails if missing

* remove example and select_dtype factory

* generalize callable case (all specification types) + add tests

* add docstring
2018-07-17 22:27:52 +02:00
Thomas Fan 895dfd3f5c ENH Adds transformer support in ColumnTransformer.remainder (#11315) 2018-06-26 23:27:47 +10:00