scikit-learn/sklearn/ensemble/setup.py

55 lines
2.1 KiB
Python
Raw Normal View History

import numpy
from numpy.distutils.misc_util import Configuration
def configuration(parent_package="", top_path=None):
config = Configuration("ensemble", parent_package, top_path)
config.add_extension("_gradient_boosting",
sources=["_gradient_boosting.pyx"],
include_dirs=[numpy.get_include()])
config.add_subpackage("tests")
# Histogram-based gradient boosting files
config.add_extension(
"_hist_gradient_boosting._gradient_boosting",
sources=["_hist_gradient_boosting/_gradient_boosting.pyx"],
include_dirs=[numpy.get_include()])
config.add_extension("_hist_gradient_boosting.histogram",
sources=["_hist_gradient_boosting/histogram.pyx"],
include_dirs=[numpy.get_include()])
config.add_extension("_hist_gradient_boosting.splitting",
sources=["_hist_gradient_boosting/splitting.pyx"],
include_dirs=[numpy.get_include()])
config.add_extension("_hist_gradient_boosting._binning",
sources=["_hist_gradient_boosting/_binning.pyx"],
include_dirs=[numpy.get_include()])
config.add_extension("_hist_gradient_boosting._predictor",
sources=["_hist_gradient_boosting/_predictor.pyx"],
include_dirs=[numpy.get_include()])
config.add_extension("_hist_gradient_boosting._loss",
sources=["_hist_gradient_boosting/_loss.pyx"],
include_dirs=[numpy.get_include()])
ENH Native support for missing values in GBDTs (#13911) * Added NaN support in mapper * pep * WIP * some more * WIP * WIP * bug fix * basic tests * some doc * avoid some interactions * Added tag * better test * decent test + fix bug * add missing_fraction param to benchmark * bin training and validation data separately * shorter test * Map missing values to first bin instead of last * pep8 * Added whats new entry * avoid some python interactions * make predict_binned work * fixed bug due to offset in bin_thresholds_ attribute * more sensible binning strat * typo * user name * Add small test * convert to fortran array in tests * some doc * Added function test * pep8 * Bin validation data using binmaper of training data * Allocate first bin for missing entries based on the whole data, not just training data. * Addressed Thomas' comments * Update sklearn/ensemble/_hist_gradient_boosting/tests/test_grower.py * Addressed Guillaume's comments * always allocate first bin for missing values * reduce diff * minor more consistent test * typo * WIP * some doc * reduce diff * pep8 * minor * remove prints * towards nan only splits * don't check right to left on split_on_nan * cleaups * format and comment * Fixed bug + added more tests * refactor tests * put back n_threads to max value * minor changes * minor cleaning * Add (failing) test that checks equivalence with min max imputation * Decrease the likelihood of ties when training the trees * More robust test * Fix pytest parametrization * Check bin thresholds in test * Try to make the test even easier to see if the Linux 32bit build would pass in this case * Don't check last non-missing bin if there's no nan * Improve min-max imputation test * FIX: _find_best_bin_to_split_right_to_left is still required even when left to right wants to split on nans * comments * remove split_on_nan * ooops deleted useless files * Got rid of individual checks in predictor code +inf thresholds are only allowed in a split on nan situation. Thresholds that are computed as +inf are capped to a very high constant value * can also remove special case in binning code * minor typos + more consistent test * renamed types -> common * 1e300 -> almost inf * added user guide section on missing values * Addressed Olivier's comment + updated whatsnew * addressed comments * Fix doctest formatting * Fix nan predictive doctest
2019-08-21 17:22:00 +08:00
config.add_extension("_hist_gradient_boosting.common",
sources=["_hist_gradient_boosting/common.pyx"],
include_dirs=[numpy.get_include()])
config.add_extension("_hist_gradient_boosting.utils",
sources=["_hist_gradient_boosting/utils.pyx"],
include_dirs=[numpy.get_include()])
config.add_subpackage("_hist_gradient_boosting.tests")
return config
if __name__ == "__main__":
from numpy.distutils.core import setup
setup(**configuration().todict())