In previous commit, I used n_features to set the number of features, and plotted based on that variable. Based on feedback, I removed n_features, and now plotting is based on X.shape[1]. This makes the code general and easy to port. I copied this into an iPython notebook to verify it still outputs the correct graph.
I ported this code into my own, and when I tried plotting, I found the number of features was hard coded to 10. By switching to a set variable, the number of features is no longer a magic number, and the code is more readable.
This example trains several tree based ensemble methods and uses
them to transform the data into a high dimensional, sparse space.
The trains a linear model on this new feature space. The idea is
taken from:
Practical Lessons from Predicting Clicks on Ads at Facebook Junfeng Pan,
He Xinran, Ou Jin, Tianbing XU, Bo Liu, Tao Xu, Yanxin Shi, Antoine
Atallah, Ralf Herbrich, Stuart Bowers, Joaquin Quiñonero Candela
International Workshop on Data Mining for Online Advertising (ADKDD)
https://www.facebook.com/publications/329190253909587/
A number of further amendments to the plot_ensemble_oob.py example
script were suggested in the PR thread and addressed accordingly:
- The ExtraTreesClassifier models were removed from the example, since
they don't use bootstrapping by default (but can be using bootstrap=True).
- Included the OOB errors for RandomForestClassifier models with various
max_features values.
- Changed the sample datasets to make for a nicer looking plot.
- Changed "cross-validated" to "validated" in the docstring.
- Added the relevant page numbers to the Hastie et al. reference.
- PEP8 compliance, fixed line > 80 chars.
@amueller provided feedback on improving my original PR (#4665) of the
plot_ensemble_oob.py script.
A number of major changes were made accordingly:
- Used `matplotlib.pyplot` instead of `pylab`.
- To improve the run-time to <10secs, I reduced the dimensionality of
the sample dataset and set the max. number of estimators to 150.
- To avoid OOB warnings, the min. number of estimators was set to 15.
Values <15 would still raise the warnings.
- The script is PEP8-compliant via the `pep8` command-line script. I
needed to move `print(__doc__)` and author list comments.
- Re-added @amueller to the author list (had mistakenly been removed).
- Added a link to this example to the user-guide under the `Ensemble
Methods` section.