<oml:flow xmlns:oml="http://openml.org/openml">
  <oml:id>17456</oml:id>
<oml:uploader>10776</oml:uploader>
<oml:name>sklearn.tree._classes.DecisionTreeClassifier</oml:name>
<oml:custom_name>sklearn.DecisionTreeClassifier</oml:custom_name>
<oml:class_name>sklearn.tree._classes.DecisionTreeClassifier</oml:class_name>
<oml:version>2</oml:version>
<oml:external_version>openml==0.10.2,sklearn==0.22</oml:external_version>
<oml:description>A decision tree classifier.</oml:description>
<oml:upload_date>2019-12-13T20:07:47</oml:upload_date>
<oml:language>English</oml:language>
<oml:dependencies>sklearn==0.22
numpy&gt;=1.6.1
scipy&gt;=0.9</oml:dependencies>
<oml:parameter>
	<oml:name>ccp_alpha</oml:name>
	<oml:data_type>non</oml:data_type>
	<oml:default_value>0.0</oml:default_value>
	<oml:description>Complexity parameter used for Minimal Cost-Complexity Pruning. The
    subtree with the largest cost complexity that is smaller than
    ``ccp_alpha`` will be chosen. By default, no pruning is performed. See
    :ref:`minimal_cost_complexity_pruning` for details

    .. versionadded:: 0.22</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>class_weight</oml:name>
	<oml:data_type>dict</oml:data_type>
	<oml:default_value>null</oml:default_value>
	<oml:description>Weights associated with classes in the form ``{class_label: weight}``
    If not given, all classes are supposed to have weight one. For
    multi-output problems, a list of dicts can be provided in the same
    order as the columns of y

    Note that for multioutput (including multilabel) weights should be
    defined for each class of every column in its own dict. For example,
    for four-class multilabel classification weights should be
    [{0: 1, 1: 1}, {0: 1, 1: 5}, {0: 1, 1: 1}, {0: 1, 1: 1}] instead of
    [{1:1}, {2:5}, {3:1}, {4:1}]

    The &quot;balanced&quot; mode uses the values of y to automatically adjust
    weights inversely proportional to class frequencies in the input data
    as ``n_samples / (n_classes * np.bincount(y))``

    For multi-output, the weights of each column of y will be multiplied

    Note that these weights will be multiplied with sample_weight (passed
    through the fit method) if sample_weight is specified</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>criterion</oml:name>
	<oml:data_type>str</oml:data_type>
	<oml:default_value>&quot;gini&quot;</oml:default_value>
	<oml:description>The function to measure the quality of a split. Supported criteria are
    &quot;gini&quot; for the Gini impurity and &quot;entropy&quot; for the information gain</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>max_depth</oml:name>
	<oml:data_type>int or None</oml:data_type>
	<oml:default_value>null</oml:default_value>
	<oml:description>The maximum depth of the tree. If None, then nodes are expanded until
    all leaves are pure or until all leaves contain less than
    min_samples_split samples</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>max_features</oml:name>
	<oml:data_type>int</oml:data_type>
	<oml:default_value>null</oml:default_value>
	<oml:description>The number of features to consider when looking for the best split:

        - If int, then consider `max_features` features at each split
        - If float, then `max_features` is a fraction and
          `int(max_features * n_features)` features are considered at each
          split
        - If &quot;auto&quot;, then `max_features=sqrt(n_features)`
        - If &quot;sqrt&quot;, then `max_features=sqrt(n_features)`
        - If &quot;log2&quot;, then `max_features=log2(n_features)`
        - If None, then `max_features=n_features`

    Note: the search for a split does not stop until at least one
    valid partition of the node samples is found, even if it requires to
    effectively inspect more than ``max_features`` features</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>max_leaf_nodes</oml:name>
	<oml:data_type>int or None</oml:data_type>
	<oml:default_value>null</oml:default_value>
	<oml:description>Grow a tree with ``max_leaf_nodes`` in best-first fashion
    Best nodes are defined as relative reduction in impurity
    If None then unlimited number of leaf nodes</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>min_impurity_decrease</oml:name>
	<oml:data_type>float</oml:data_type>
	<oml:default_value>0.0</oml:default_value>
	<oml:description>A node will be split if this split induces a decrease of the impurity
    greater than or equal to this value

    The weighted impurity decrease equation is the following::

        N_t / N * (impurity - N_t_R / N_t * right_impurity
                            - N_t_L / N_t * left_impurity)

    where ``N`` is the total number of samples, ``N_t`` is the number of
    samples at the current node, ``N_t_L`` is the number of samples in the
    left child, and ``N_t_R`` is the number of samples in the right child

    ``N``, ``N_t``, ``N_t_R`` and ``N_t_L`` all refer to the weighted sum,
    if ``sample_weight`` is passed

    .. versionadded:: 0.19</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>min_impurity_split</oml:name>
	<oml:data_type>float</oml:data_type>
	<oml:default_value>null</oml:default_value>
	<oml:description>Threshold for early stopping in tree growth. A node will split
    if its impurity is above the threshold, otherwise it is a leaf

    .. deprecated:: 0.19
       ``min_impurity_split`` has been deprecated in favor of
       ``min_impurity_decrease`` in 0.19. The default value of
       ``min_impurity_split`` will change from 1e-7 to 0 in 0.23 and it
       will be removed in 0.25. Use ``min_impurity_decrease`` instead</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>min_samples_leaf</oml:name>
	<oml:data_type>int</oml:data_type>
	<oml:default_value>1</oml:default_value>
	<oml:description>The minimum number of samples required to be at a leaf node
    A split point at any depth will only be considered if it leaves at
    least ``min_samples_leaf`` training samples in each of the left and
    right branches.  This may have the effect of smoothing the model,
    especially in regression

    - If int, then consider `min_samples_leaf` as the minimum number
    - If float, then `min_samples_leaf` is a fraction and
      `ceil(min_samples_leaf * n_samples)` are the minimum
      number of samples for each node

    .. versionchanged:: 0.18
       Added float values for fractions</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>min_samples_split</oml:name>
	<oml:data_type>int</oml:data_type>
	<oml:default_value>2</oml:default_value>
	<oml:description>The minimum number of samples required to split an internal node:

    - If int, then consider `min_samples_split` as the minimum number
    - If float, then `min_samples_split` is a fraction and
      `ceil(min_samples_split * n_samples)` are the minimum
      number of samples for each split

    .. versionchanged:: 0.18
       Added float values for fractions</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>min_weight_fraction_leaf</oml:name>
	<oml:data_type>float</oml:data_type>
	<oml:default_value>0.0</oml:default_value>
	<oml:description>The minimum weighted fraction of the sum total of weights (of all
    the input samples) required to be at a leaf node. Samples have
    equal weight when sample_weight is not provided</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>presort</oml:name>
	<oml:data_type>deprecated</oml:data_type>
	<oml:default_value>&quot;deprecated&quot;</oml:default_value>
	<oml:description>This parameter is deprecated and will be removed in v0.24

    .. deprecated:: 0.22</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>random_state</oml:name>
	<oml:data_type>int</oml:data_type>
	<oml:default_value>null</oml:default_value>
	<oml:description>If int, random_state is the seed used by the random number generator;
    If RandomState instance, random_state is the random number generator;
    If None, the random number generator is the RandomState instance used
    by `np.random`</oml:description>
</oml:parameter>
<oml:parameter>
	<oml:name>splitter</oml:name>
	<oml:data_type>str</oml:data_type>
	<oml:default_value>&quot;best&quot;</oml:default_value>
	<oml:description>The strategy used to choose the split at each node. Supported
    strategies are &quot;best&quot; to choose the best split and &quot;random&quot; to choose
    the best random split</oml:description>
</oml:parameter>
<oml:tag>openml-python</oml:tag>
<oml:tag>python</oml:tag>
<oml:tag>scikit-learn</oml:tag>
<oml:tag>sklearn</oml:tag>
<oml:tag>sklearn_0.22</oml:tag>
</oml:flow>