HashingVectorizer#

class cuml.feature_extraction.text.HashingVectorizer(
*,
lowercase=True,
preprocessor=None,
tokenizer=None,
delimiter=None,
stop_words=None,
ngram_range=(1,
1),
analyzer='word',
alternate_sign=True,
n_features=1048576,
dtype=<class 'numpy.float32'>,
binary=False,
norm='l2',
verbose=False,
output_type=None,
)[source]#

Convert a collection of text documents to a matrix of token occurrences.

It turns a collection of text documents into a sparse matrix holding token occurrence counts (or binary occurrence information), possibly normalized as token frequencies if norm=’l1’ or projected on the euclidean unit sphere if norm=’l2’.

This text vectorizer implementation uses the hashing trick to find the token string name to feature integer index mapping.

This strategy has several advantages:

  • it is very low memory scalable to large datasets as there is no need to store a vocabulary dictionary in memory.

  • it is fast to pickle and un-pickle as it holds no state besides the constructor parameters.

  • it can be used in a streaming (partial fit) or parallel pipeline as there is no state computed during fit.

There are also a couple of cons (vs using a CountVectorizer with an in-memory vocabulary):

  • there is no way to compute the inverse transform (from feature indices to string feature names) which can be a problem when trying to introspect which features are most important to a model.

  • there can be collisions: distinct tokens can be mapped to the same feature index. However in practice this is rarely an issue if n_features is large enough (e.g. 2 ** 18 for text classification problems).

  • no IDF weighting as this would render the transformer stateful.

The hash function employed is the signed 32-bit version of Murmurhash3.

Parameters:
lowercasebool, default=True

Convert all characters to lowercase before tokenizing.

preprocessorcallable, default=None

Override the preprocessing (string transformation) stage while preserving the tokenizing and n-grams generation steps. This function receives a cudf.Series of strings and should return a cudf.Series of strings.

tokenizercallable, default=None

Override the string tokenization step while preserving the preprocessing and n-grams generation steps. This function receives a cudf.Series of strings and should return a cudf.Series of lists of strings. Only applies if analyzer == 'word'.

delimiterstr, default=None

String used to delimit tokens in the document. If None, then any non-alphanumeric (or “ “) character is treated as a delimiter. Only applies if analyzer == "word'.

stop_words{‘english’}, list, default=None

If ‘english’, a built-in stop word list for English is used. If a list, that list is assumed to contain stop words, all of which will be removed from the resulting tokens. If None, no stop words will be used. Only applies if analyzer == 'word'.

ngram_rangetuple (min_n, max_n), default=(1, 1)

The lower and upper boundary of the range of n-values for different n-grams to be extracted. All values of n such that min_n <= n <= max_n will be used. For example an ngram_range of (1, 1) means only unigrams, (1, 2) means unigrams and bigrams, and (2, 2) means only bigrams.

analyzer{‘word’, ‘char’, ‘char_wb’}, default=’word’

Whether the feature should be made of word or character n-grams. Option ‘char_wb’ creates character n-grams only from text inside word boundaries; n-grams at the edges of words are padded with space.

n_featuresint, default=(2 ** 20)

The number of features (columns) in the output matrices. Small numbers of features are likely to cause hash collisions, but large numbers will cause larger coefficient dimensions in linear learners.

binarybool, default=False

If True, all non zero counts are set to 1. This is useful for discrete probabilistic models that model binary events rather than integer counts.

norm{‘l1’, ‘l2’, None}, default=’l2’

Norm used to normalize term vectors. None for no normalization.

alternate_signbool, default=True

When True, an alternating sign is added to the features as to approximately conserve the inner product in the hashed space even for small n_features. This approach is similar to sparse random projection.

dtypetype, default=np.float32

Type of the matrix returned by fit_transform() or transform().

verboseint or boolean, default=False

Sets logging level. It must be one of cuml.common.logger.level_*. See Verbosity Levels for more info.

output_type{None, ‘input’, ‘cupy’, ‘numpy’, ‘cudf’, ‘pandas’}, default=None

Return results and set estimator attributes to the indicated output type. If None, the output type set at the module level (cuml.global_settings.output_type) will be used. See Output Data Type Configuration for more info.

Methods

fit(X[, y])

Only validate's the estimator's parameters.

fit_transform(X[, y])

Transform a sequence of documents to a document-term matrix.

partial_fit(X[, y])

Only validate's the estimator's parameters.

transform(X)

Transform a sequence of documents to a document-term matrix.

Examples

>>> from cuml.feature_extraction.text import HashingVectorizer
>>> corpus = [
...     'This is the first document.',
...     'This document is the second document.',
...     'And this is the third one.',
...     'Is this the first document?',
... ]
>>> vectorizer = HashingVectorizer(n_features=2**4)
>>> X = vectorizer.fit_transform(corpus)
>>> X.shape
(4, 16)
fit(X, y=None)[source]#

Only validate’s the estimator’s parameters.

This estimator is stateless, fit is a no-op.

Parameters:
XIterable[str]

Training samples. Each sample must be a text document which will be tokenized and hashed.

yNone

Ignored. Exists for API compatibility only.

Returns:
selfobject

The instance itself.

fit_transform(X, y=None)[source]#

Transform a sequence of documents to a document-term matrix.

Parameters:
XIterable[str]

Training samples. Each sample must be a text document which will be tokenized and hashed.

yNone

Ignored. Exists for API compatibility only.

Returns:
Xsparse matrix of shape (n_samples, n_features)

Document-term matrix.

get_params(deep=True)[source]#

Returns a dict of all params owned by this class. If the child class has appropriately overridden the _get_param_names method and does not need anything other than what is there in this method, then it doesn’t have to override this method

partial_fit(X, y=None)[source]#

Only validate’s the estimator’s parameters.

This estimator is stateless, fit is a no-op.

Parameters:
XIterable[str]

Training samples. Each sample must be a text document which will be tokenized and hashed.

yNone

Ignored. Exists for API compatibility only.

Returns:
selfobject

The instance itself.

set_params(**params)[source]#

Accepts a dict of params and updates the corresponding ones owned by this class. If the child class has appropriately overridden the _get_param_names method and does not need anything other than what is, there in this method, then it doesn’t have to override this method

transform(X)[source]#

Transform a sequence of documents to a document-term matrix.

Parameters:
XIterable[str]

Training samples. Each sample must be a text document which will be tokenized and hashed.

Returns:
Xsparse matrix of shape (n_samples, n_features)

Document-term matrix.