LinearSVC#

class cuml.svm.LinearSVC(
*,
penalty='l2',
loss='squared_hinge',
C=1.0,
fit_intercept=True,
penalized_intercept=False,
class_weight=None,
tol=0.0001,
max_iter=1000,
linesearch_max_iter=100,
lbfgs_memory=5,
n_streams=1,
multi_class='ovr',
verbose=False,
output_type=None,
)[source]#

Linear Support Vector Classification.

Similar to SVC with parameter kernel=’linear’, but implemented using a linear solver. This enables flexibility in penalties and loss functions, and can scale better for larger problems.

Parameters:
penalty{‘l1’, ‘l2’}, default = ‘l2’

The norm used in the penalization.

loss{‘hinge’, ‘squared_hinge’}, default=’squared_hinge’

The loss function.

Cfloat, default=1.0

Regularization parameter. The strength of the regularization is inversely proportional to C. Must be strictly positive.

fit_interceptbool, default=True

Whether to fit the bias term. Set to False if you expect that the data is already centered.

penalized_interceptbool, default=False

When true, the bias term is treated the same way as other features; i.e. it’s penalized by the regularization term of the target function. Enabling this feature forces an extra copying the input data X.

class_weightdict or string, default=None

Weights to modify the parameter C for class i to class_weight[i]*C. The string ‘balanced’ is also accepted, in which case class_weight[i] = n_samples / (n_classes * n_samples_of_class[i])

tolfloat, default=1e-4

Tolerance for the stopping criterion.

max_iterint, default=1000

Maximum number of iterations for the underlying solver.

linesearch_max_iterint, default=100

Maximum number of linesearch (inner loop) iterations for the underlying (QN) solver.

lbfgs_memoryint, default=5

Number of vectors approximating the hessian for the underlying QN solver (l-bfgs).

n_streamsint (default = 1)

Number of parallel streams used for fitting.

multi_class{‘ovr’}, default=’ovr’

Multiclass classification strategy. Currently only ‘ovr’ is supported.

verboseint or boolean, default=False

Sets logging level. It must be one of cuml.common.logger.level_*. See Verbosity Levels for more info.

output_type{None, ‘input’, ‘cupy’, ‘numpy’, ‘cudf’, ‘pandas’}, default=None

Return results and set estimator attributes to the indicated output type. If None, the output type set at the module level (cuml.global_settings.output_type) will be used. See Output Data Type Configuration for more info.

Attributes:
coef_array, shape (1, n_features) if n_classes == 2 else (n_classes, n_features)

Weights assigned to the features (coefficients in the primal problem).

intercept_array or float, shape (1,) if n_classes == 2 else (n_classes,)

The constant factor in the decision function. If fit_intercept=False this is instead a float with value 0.0.

classes_np.ndarray, shape=(n_classes,)

A sorted array of the class labels.

n_iter_int

The maximum number of iterations run across all classes during the fit.

Methods

fit(X, y[, sample_weight])

Fit the model according to the given training data.

predict(X)

Predict class labels for samples in X.

Notes

The model uses the quasi-newton (QN) solver to find the solution in the primal space. Thus, in contrast to generic SVC model, it does not compute the support coefficients/vectors.

Check the solver’s documentation for more details Quasi-Newton (L-BFGS/OWL-QN).

For additional docs, see scikitlearn’s LinearSVC.

Examples

>>> import cupy as cp
>>> from cuml.svm import LinearSVC
>>> X = cp.array([[1,1], [2,1], [1,2], [2,2], [1,3], [2,3]],
...              dtype=cp.float32);
>>> y = cp.array([0, 0, 1, 0, 1, 1], dtype=cp.float32)
>>> clf = LinearSVC(penalty='l1', C=1).fit(X, y)
>>> print("Predicted labels:", clf.predict(X))
Predicted labels: [0 0 1 0 1 1]
as_sklearn()[source]#

Convert this estimator into an equivalent scikit-learn (or scikit-learn extension) estimator.

Returns:
sklearn.base.BaseEstimator

A scikit-learn compatible estimator instance that mirrors the trained state of the current estimator.

decision_function(X)[source]#

Predict confidence scores for samples.

Parameters:
Xarray-like (device or host) shape = (n_samples, n_features)

Dense or sparse matrix with dtype float32 or float64. Acceptable dense formats: CUDA array interface compliant objects like CuPy, cuDF DataFrame/Series, NumPy ndarray and Pandas DataFrame/Series.

Returns:
scorescuDF, CuPy or NumPy object depending on cuML’s output type configuration, shape = (n_samples,) or (n_samples, n_classes)

Confidence scores

For more information on how to configure cuML’s output type, refer to: Output Data Type Configuration.

fit(
X,
y,
sample_weight=None,
) LinearSVC[source]#

Fit the model according to the given training data.

Parameters:
Xarray-like (device or host) shape = (n_samples, n_features)

Dense matrix with dtype float32 or float64. Acceptable formats: CUDA array interface compliant objects like CuPy, cuDF DataFrame/Series, NumPy ndarray and Pandas DataFrame/Series.

yarray-like (device or host) shape = (n_samples, 1)

Dense matrix with dtype float32 or float64. Acceptable formats: CUDA array interface compliant objects like CuPy, cuDF DataFrame/Series, NumPy ndarray and Pandas DataFrame/Series.

sample_weightarray-like (device or host) shape = (n_samples,), default=None

The weights for each observation in X. If None, all observations are assigned equal weight. Acceptable formats: CUDA array interface compliant objects like CuPy, cuDF DataFrame/Series, NumPy ndarray and Pandas DataFrame/Series.

classmethod from_sklearn(model)[source]#

Create a cuml estimator from a scikit-learn estimator.

Parameters:
modelsklearn.base.BaseEstimator

A compatible scikit-learn (or scikit-learn extension) estimator.

Returns:
cls

A new instance of this cuml estimator class that mirrors the state of the input estimator.

Notes

output_type of the estimator is set to “numpy” by default, as these cannot be inferred from training arguments. If something different is required, then please use cuml’s output_type configuration utilities.

get_params(deep=True)[source]#

Returns a dict of all params owned by this class. If the child class has appropriately overridden the _get_param_names method and does not need anything other than what is there in this method, then it doesn’t have to override this method

predict(X)[source]#

Predict class labels for samples in X.

Parameters:
Xarray-like (device or host) shape = (n_samples, n_features)

Dense matrix with dtype float32 or float64. Acceptable formats: CUDA array interface compliant objects like CuPy, cuDF DataFrame/Series, NumPy ndarray and Pandas DataFrame/Series.

Returns:
y_predcuDF, CuPy or NumPy object depending on cuML’s output type configuration, shape = (n_samples,)

Predicted class labels.

For more information on how to configure cuML’s output type, refer to: Output Data Type Configuration.

score(X, y, sample_weight=None, **kwargs)[source]#

Scoring function for classifier estimators based on mean accuracy.

Parameters:
Xarray-like (device or host) shape = (n_samples, n_features)

Dense matrix with dtype float32 or float64. Acceptable formats: CUDA array interface compliant objects like CuPy, cuDF DataFrame/Series, NumPy ndarray and Pandas DataFrame/Series.

yarray-like (device or host) shape = (n_samples, 1)

Dense matrix with dtype float32 or float64. Acceptable formats: CUDA array interface compliant objects like CuPy, cuDF DataFrame/Series, NumPy ndarray and Pandas DataFrame/Series.

sample_weightarray-like (device or host) shape = (n_samples,), default=None

The weights for each observation in X. If None, all observations are assigned equal weight. Acceptable formats: CUDA array interface compliant objects like CuPy, cuDF DataFrame/Series, NumPy ndarray and Pandas DataFrame/Series.

Returns:
scorefloat

Accuracy of self.predict(X) wrt. y (fraction where y == pred_y)

set_params(**params)[source]#

Accepts a dict of params and updates the corresponding ones owned by this class. If the child class has appropriately overridden the _get_param_names method and does not need anything other than what is, there in this method, then it doesn’t have to override this method