cugraph.all_pairs_sorensen#

cugraph.all_pairs_sorensen(
input_graph: Graph,
vertices: Series = None,
use_weight: bool = False,
topk: int = None,
) DataFrame[source]#

Compute the All Pairs Sorensen coefficient between each pair of vertices connected by an edge, or between arbitrary pairs of vertices specified by the user. Sorensen coefficient is defined between two sets as the ratio of twice the volume of their intersection over the volume of each set. If first is specified but second is not, or vice versa, an exception will be thrown.

cugraph.all_pairs_sorensen, in the absence of specified vertices, will compute the two_hop_neighbors of the entire graph to construct a vertex pair list and will return the sorensen coefficient for all the vertex pairs in the graph. This is not advisable as the vertex_pairs can grow exponentially with respect to the size of the datasets.

If the topk parameter is specified then the result will only contain the top k highest scoring results.

Parameters:
input_graphcugraph.Graph

cuGraph Graph instance, should contain the connectivity information as an edge list. The graph should be undirected where an undirected edge is represented by a directed edge in both direction.The adjacency list will be computed if not already present.

This implementation only supports undirected, non-multi Graphs.

verticesint or list or cudf.Series or cudf.DataFrame, optional (default=None)

A GPU Series containing the input vertex list. If the vertex list is not provided then the current implementation computes the sorensen coefficient for all vertices that are two hops apart in the graph.

use_weightbool, optional (default=False)

Flag to indicate whether to compute weighted sorensen (if use_weight==True) or un-weighted sorensen (if use_weight==False). ‘input_graph’ must be weighted if ‘use_weight=True’.

topkint, optional (default=None)

Specify the number of answers to return otherwise returns the entire solution

Returns:
dfcudf.DataFrame

GPU data frame of size E (the default) or the size of the given pairs (first, second) containing the Sorensen weights. The ordering is relative to the adjacency list, or that given by the specified vertex pairs.

df[‘first’]cudf.Series

The first vertex ID of each pair (will be identical to first if specified).

df[‘second’]cudf.Series

The second vertex ID of each pair (will be identical to second if specified).

df[‘sorensen_coeff’]cudf.Series

The computed Sorensen coefficient between the first and the second vertex ID.

Examples

>>> from cugraph.datasets import karate
>>> from cugraph import all_pairs_sorensen
>>> input_graph = karate.get_graph(download=True, ignore_weights=True)
>>> df = all_pairs_sorensen(input_graph)