<oml:flow xmlns:oml="http://openml.org/openml">
  <oml:id>17493</oml:id>
<oml:uploader>5348</oml:uploader>
<oml:name>sklearn.ensemble._forest.RandomForestClassifier</oml:name>
<oml:custom_name>sklearn.RandomForestClassifier</oml:custom_name>
<oml:class_name>sklearn.ensemble._forest.RandomForestClassifier</oml:class_name>
<oml:version>1</oml:version>
<oml:external_version>openml==0.11.0dev,sklearn==0.22.1</oml:external_version>
<oml:description>A random forest classifier.

A random forest is a meta estimator that fits a number of decision tree
classifiers on various sub-samples of the dataset and uses averaging to
improve the predictive accuracy and control over-fitting.
The sub-sample size is always the same as the original
input sample size but the samples are drawn with replacement if
`bootstrap=True` (default).</oml:description>
<oml:upload_date>2020-01-16T15:55:06</oml:upload_date>
<oml:language>English</oml:language>
<oml:dependencies>sklearn==0.22.1
numpy&gt;=1.6.1
scipy&gt;=0.9</oml:dependencies>
<oml:parameter>
	<oml:name>bootstrap</oml:name>
	<oml:data_type>boolean</oml:data_type>
	<oml:default_value>true</oml:default_value>
	<oml:description>Whether bootstrap samples are used when building trees. If False, the
    whole datset is used to build each tree</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>ccp_alpha</oml:name>
	<oml:data_type>non</oml:data_type>
	<oml:default_value>0.0</oml:default_value>
	<oml:description>Complexity parameter used for Minimal Cost-Complexity Pruning. The
    subtree with the largest cost complexity that is smaller than
    ``ccp_alpha`` will be chosen. By default, no pruning is performed. See
    :ref:`minimal_cost_complexity_pruning` for details

    .. versionadded:: 0.22</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>class_weight</oml:name>
	<oml:data_type>dict</oml:data_type>
	<oml:default_value>null</oml:default_value>
	<oml:description>Weights associated with classes in the form ``{class_label: weight}``
    If not given, all classes are supposed to have weight one. For
    multi-output problems, a list of dicts can be provided in the same
    order as the columns of y

    Note that for multioutput (including multilabel) weights should be
    defined for each class of every column in its own dict. For example,
    for four-class multilabel classification weights should be
    [{0: 1, 1: 1}, {0: 1, 1: 5}, {0: 1, 1: 1}, {0: 1, 1: 1}] instead of
    [{1:1}, {2:5}, {3:1}, {4:1}]

    The &quot;balanced&quot; mode uses the values of y to automatically adjust
    weights inversely proportional to class frequencies in the input data
    as ``n_samples / (n_classes * np.bincount(y))``

    The &quot;balanced_subsample&quot; mode is the same as &quot;balanced&quot; except that
    weights are computed based on the bootstrap sample for every tree
    grown

    For multi-output, the weights of each column of y will be multiplied

    Note that these weights will be multiplied...</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>criterion</oml:name>
	<oml:data_type>string</oml:data_type>
	<oml:default_value>&quot;gini&quot;</oml:default_value>
	<oml:description>The function to measure the quality of a split. Supported criteria are
    &quot;gini&quot; for the Gini impurity and &quot;entropy&quot; for the information gain
    Note: this parameter is tree-specific</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>max_depth</oml:name>
	<oml:data_type>integer or None</oml:data_type>
	<oml:default_value>null</oml:default_value>
	<oml:description>The maximum depth of the tree. If None, then nodes are expanded until
    all leaves are pure or until all leaves contain less than
    min_samples_split samples</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>max_features</oml:name>
	<oml:data_type>int</oml:data_type>
	<oml:default_value>&quot;auto&quot;</oml:default_value>
	<oml:description>The number of features to consider when looking for the best split:

    - If int, then consider `max_features` features at each split
    - If float, then `max_features` is a fraction and
      `int(max_features * n_features)` features are considered at each
      split
    - If &quot;auto&quot;, then `max_features=sqrt(n_features)`
    - If &quot;sqrt&quot;, then `max_features=sqrt(n_features)` (same as &quot;auto&quot;)
    - If &quot;log2&quot;, then `max_features=log2(n_features)`
    - If None, then `max_features=n_features`

    Note: the search for a split does not stop until at least one
    valid partition of the node samples is found, even if it requires to
    effectively inspect more than ``max_features`` features</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>max_leaf_nodes</oml:name>
	<oml:data_type>int or None</oml:data_type>
	<oml:default_value>null</oml:default_value>
	<oml:description>Grow trees with ``max_leaf_nodes`` in best-first fashion
    Best nodes are defined as relative reduction in impurity
    If None then unlimited number of leaf nodes</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>max_samples</oml:name>
	<oml:data_type>int or float</oml:data_type>
	<oml:default_value>null</oml:default_value>
	<oml:description>If bootstrap is True, the number of samples to draw from X
    to train each base estimator

    - If None (default), then draw `X.shape[0]` samples
    - If int, then draw `max_samples` samples
    - If float, then draw `max_samples * X.shape[0]` samples. Thus,
      `max_samples` should be in the interval `(0, 1)`

    .. versionadded:: 0.22</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>min_impurity_decrease</oml:name>
	<oml:data_type>float</oml:data_type>
	<oml:default_value>0.0</oml:default_value>
	<oml:description>A node will be split if this split induces a decrease of the impurity
    greater than or equal to this value

    The weighted impurity decrease equation is the following::

        N_t / N * (impurity - N_t_R / N_t * right_impurity
                            - N_t_L / N_t * left_impurity)

    where ``N`` is the total number of samples, ``N_t`` is the number of
    samples at the current node, ``N_t_L`` is the number of samples in the
    left child, and ``N_t_R`` is the number of samples in the right child

    ``N``, ``N_t``, ``N_t_R`` and ``N_t_L`` all refer to the weighted sum,
    if ``sample_weight`` is passed

    .. versionadded:: 0.19</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>min_impurity_split</oml:name>
	<oml:data_type>float</oml:data_type>
	<oml:default_value>null</oml:default_value>
	<oml:description>Threshold for early stopping in tree growth. A node will split
    if its impurity is above the threshold, otherwise it is a leaf

    .. deprecated:: 0.19
       ``min_impurity_split`` has been deprecated in favor of
       ``min_impurity_decrease`` in 0.19. The default value of
       ``min_impurity_split`` will change from 1e-7 to 0 in 0.23 and it
       will be removed in 0.25. Use ``min_impurity_decrease`` instead</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>min_samples_leaf</oml:name>
	<oml:data_type>int</oml:data_type>
	<oml:default_value>1</oml:default_value>
	<oml:description>The minimum number of samples required to be at a leaf node
    A split point at any depth will only be considered if it leaves at
    least ``min_samples_leaf`` training samples in each of the left and
    right branches.  This may have the effect of smoothing the model,
    especially in regression

    - If int, then consider `min_samples_leaf` as the minimum number
    - If float, then `min_samples_leaf` is a fraction and
      `ceil(min_samples_leaf * n_samples)` are the minimum
      number of samples for each node

    .. versionchanged:: 0.18
       Added float values for fractions</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>min_samples_split</oml:name>
	<oml:data_type>int</oml:data_type>
	<oml:default_value>2</oml:default_value>
	<oml:description>The minimum number of samples required to split an internal node:

    - If int, then consider `min_samples_split` as the minimum number
    - If float, then `min_samples_split` is a fraction and
      `ceil(min_samples_split * n_samples)` are the minimum
      number of samples for each split

    .. versionchanged:: 0.18
       Added float values for fractions</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>min_weight_fraction_leaf</oml:name>
	<oml:data_type>float</oml:data_type>
	<oml:default_value>0.0</oml:default_value>
	<oml:description>The minimum weighted fraction of the sum total of weights (of all
    the input samples) required to be at a leaf node. Samples have
    equal weight when sample_weight is not provided</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>n_estimators</oml:name>
	<oml:data_type>integer</oml:data_type>
	<oml:default_value>100</oml:default_value>
	<oml:description>The number of trees in the forest

    .. versionchanged:: 0.22
       The default value of ``n_estimators`` changed from 10 to 100
       in 0.22</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>n_jobs</oml:name>
	<oml:data_type>int or None</oml:data_type>
	<oml:default_value>null</oml:default_value>
	<oml:description>The number of jobs to run in parallel. :meth:`fit`, :meth:`predict`,
    :meth:`decision_path` and :meth:`apply` are all parallelized over the
    trees. ``None`` means 1 unless in a :obj:`joblib.parallel_backend`
    context. ``-1`` means using all processors. See :term:`Glossary
    &lt;n_jobs&gt;` for more details</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>oob_score</oml:name>
	<oml:data_type>bool</oml:data_type>
	<oml:default_value>false</oml:default_value>
	<oml:description>Whether to use out-of-bag samples to estimate
    the generalization accuracy</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>random_state</oml:name>
	<oml:data_type>int</oml:data_type>
	<oml:default_value>3058</oml:default_value>
	<oml:description>Controls both the randomness of the bootstrapping of the samples used
    when building trees (if ``bootstrap=True``) and the sampling of the
    features to consider when looking for the best split at each node
    (if ``max_features &lt; n_features``)
    See :term:`Glossary &lt;random_state&gt;` for details</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>verbose</oml:name>
	<oml:data_type>int</oml:data_type>
	<oml:default_value>0</oml:default_value>
	<oml:description>Controls the verbosity when fitting and predicting</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>warm_start</oml:name>
	<oml:data_type>bool</oml:data_type>
	<oml:default_value>false</oml:default_value>
	<oml:description>When set to ``True``, reuse the solution of the previous call to fit
    and add more estimators to the ensemble, otherwise, just fit a whole
    new forest. See :term:`the Glossary &lt;warm_start&gt;`</oml:description>
</oml:parameter>
<oml:tag>openml-python</oml:tag>
<oml:tag>python</oml:tag>
<oml:tag>scikit-learn</oml:tag>
<oml:tag>sklearn</oml:tag>
<oml:tag>sklearn_0.22.1</oml:tag>
</oml:flow>