HashingVectorizer#
- class cuml.feature_extraction.text.HashingVectorizer(
- *,
- lowercase=True,
- preprocessor=None,
- tokenizer=None,
- delimiter=None,
- stop_words=None,
- ngram_range=(1,
- 1),
- analyzer='word',
- alternate_sign=True,
- n_features=1048576,
- dtype=<class 'numpy.float32'>,
- binary=False,
- norm='l2',
- verbose=False,
- output_type=None,
Convert a collection of text documents to a matrix of token occurrences.
It turns a collection of text documents into a sparse matrix holding token occurrence counts (or binary occurrence information), possibly normalized as token frequencies if norm=’l1’ or projected on the euclidean unit sphere if norm=’l2’.
This text vectorizer implementation uses the hashing trick to find the token string name to feature integer index mapping.
This strategy has several advantages:
it is very low memory scalable to large datasets as there is no need to store a vocabulary dictionary in memory.
it is fast to pickle and un-pickle as it holds no state besides the constructor parameters.
it can be used in a streaming (partial fit) or parallel pipeline as there is no state computed during fit.
There are also a couple of cons (vs using a CountVectorizer with an in-memory vocabulary):
there is no way to compute the inverse transform (from feature indices to string feature names) which can be a problem when trying to introspect which features are most important to a model.
there can be collisions: distinct tokens can be mapped to the same feature index. However in practice this is rarely an issue if n_features is large enough (e.g. 2 ** 18 for text classification problems).
no IDF weighting as this would render the transformer stateful.
The hash function employed is the signed 32-bit version of Murmurhash3.
- Parameters:
- lowercasebool, default=True
Convert all characters to lowercase before tokenizing.
- preprocessorcallable, default=None
Override the preprocessing (string transformation) stage while preserving the tokenizing and n-grams generation steps. This function receives a
cudf.Seriesof strings and should return acudf.Seriesof strings.- tokenizercallable, default=None
Override the string tokenization step while preserving the preprocessing and n-grams generation steps. This function receives a
cudf.Seriesof strings and should return acudf.Seriesof lists of strings. Only applies ifanalyzer == 'word'.- delimiterstr, default=None
String used to delimit tokens in the document. If
None, then any non-alphanumeric (or “ “) character is treated as a delimiter. Only applies ifanalyzer == "word'.- stop_words{‘english’}, list, default=None
If ‘english’, a built-in stop word list for English is used. If a list, that list is assumed to contain stop words, all of which will be removed from the resulting tokens. If None, no stop words will be used. Only applies if
analyzer == 'word'.- ngram_rangetuple (min_n, max_n), default=(1, 1)
The lower and upper boundary of the range of n-values for different n-grams to be extracted. All values of n such that min_n <= n <= max_n will be used. For example an
ngram_rangeof(1, 1)means only unigrams,(1, 2)means unigrams and bigrams, and(2, 2)means only bigrams.- analyzer{‘word’, ‘char’, ‘char_wb’}, default=’word’
Whether the feature should be made of word or character n-grams. Option ‘char_wb’ creates character n-grams only from text inside word boundaries; n-grams at the edges of words are padded with space.
- n_featuresint, default=(2 ** 20)
The number of features (columns) in the output matrices. Small numbers of features are likely to cause hash collisions, but large numbers will cause larger coefficient dimensions in linear learners.
- binarybool, default=False
If True, all non zero counts are set to 1. This is useful for discrete probabilistic models that model binary events rather than integer counts.
- norm{‘l1’, ‘l2’, None}, default=’l2’
Norm used to normalize term vectors. None for no normalization.
- alternate_signbool, default=True
When True, an alternating sign is added to the features as to approximately conserve the inner product in the hashed space even for small n_features. This approach is similar to sparse random projection.
- dtypetype, default=np.float32
Type of the matrix returned by fit_transform() or transform().
- verboseint or boolean, default=False
Sets logging level. It must be one of
cuml.common.logger.level_*. See Verbosity Levels for more info.- output_type{None, ‘input’, ‘cupy’, ‘numpy’, ‘cudf’, ‘pandas’}, default=None
Return results and set estimator attributes to the indicated output type. If None, the output type set at the module level (
cuml.global_settings.output_type) will be used. See Output Data Type Configuration for more info.
Methods
fit(X[, y])Only validate's the estimator's parameters.
fit_transform(X[, y])Transform a sequence of documents to a document-term matrix.
partial_fit(X[, y])Only validate's the estimator's parameters.
transform(X)Transform a sequence of documents to a document-term matrix.
Examples
>>> from cuml.feature_extraction.text import HashingVectorizer >>> corpus = [ ... 'This is the first document.', ... 'This document is the second document.', ... 'And this is the third one.', ... 'Is this the first document?', ... ] >>> vectorizer = HashingVectorizer(n_features=2**4) >>> X = vectorizer.fit_transform(corpus) >>> X.shape (4, 16)
- fit(X, y=None)[source]#
Only validate’s the estimator’s parameters.
This estimator is stateless,
fitis a no-op.- Parameters:
- XIterable[str]
Training samples. Each sample must be a text document which will be tokenized and hashed.
- yNone
Ignored. Exists for API compatibility only.
- Returns:
- selfobject
The instance itself.
- fit_transform(X, y=None)[source]#
Transform a sequence of documents to a document-term matrix.
- Parameters:
- XIterable[str]
Training samples. Each sample must be a text document which will be tokenized and hashed.
- yNone
Ignored. Exists for API compatibility only.
- Returns:
- Xsparse matrix of shape (n_samples, n_features)
Document-term matrix.
- get_params(deep=True)[source]#
Returns a dict of all params owned by this class. If the child class has appropriately overridden the
_get_param_namesmethod and does not need anything other than what is there in this method, then it doesn’t have to override this method
- partial_fit(X, y=None)[source]#
Only validate’s the estimator’s parameters.
This estimator is stateless,
fitis a no-op.- Parameters:
- XIterable[str]
Training samples. Each sample must be a text document which will be tokenized and hashed.
- yNone
Ignored. Exists for API compatibility only.
- Returns:
- selfobject
The instance itself.
- set_params(**params)[source]#
Accepts a dict of params and updates the corresponding ones owned by this class. If the child class has appropriately overridden the
_get_param_namesmethod and does not need anything other than what is, there in this method, then it doesn’t have to override this method