cuSOLVERMp C API#
Library Management#
cusolverMpCreate#
cusolverStatus_t cusolverMpCreate(
cusolverMpHandle_t *handle,
int deviceId,
cudaStream_t stream)
deviceId and the CUDA stream stream.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
Out |
cuSOLVERMp library handle. |
deviceId |
Host |
In |
Device that will be assigned to the handle. |
stream |
Host |
In |
Stream that will be assigned to the handle. |
cusolverMpDestroy#
cusolverStatus_t cusolverMpDestroy(
cusolverMpHandle_t handle)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In/Out |
cuSOLVERMp library handle. |
cusolverMpSetStream#
cusolverStatus_t cusolverMpSetStream(
cusolverMpHandle_t handle,
cudaStream_t stream)
stream associated to the handle.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
stream |
Host |
In |
New stream associated with the handle. |
cusolverMpGetStream#
cusolverStatus_t cusolverMpGetStream(
cusolverMpHandle_t handle,
cudaStream_t *stream)
stream associated to the handle.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
stream |
Host |
Out |
Stream associated with the handle. |
cusolverMpGetVersion#
cusolverStatus_t cusolverMpGetVersion(
cusolverMpHandle_t handle,
int *version)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
version |
Host |
Out |
cuSOLVERMp library version. Value is |
cusolverMpSetMathMode#
cusolverStatus_t cusolverMpSetMathMode(
cusolverMpHandle_t handle,
cusolverMathMode_t mode)
Note
Please note that the workspace sizes returned by *_bufferSize APIs may depend on the math mode.
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
mode |
Host |
In |
Math mode to be set for the handle. Available options are |
cusolverMpGetMathMode#
cusolverStatus_t cusolverMpGetMathMode(
cusolverMpHandle_t handle,
cusolverMathMode_t *mode)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
mode |
Host |
Out |
Current math mode of the handle. |
cusolverMpSetEmulationStrategy#
cusolverStatus_t cusolverMpSetEmulationStrategy(
cusolverMpHandle_t handle,
cudaEmulationStrategy_t strategy)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
strategy |
Host |
In |
Emulation strategy to be set for the handle. Available options are |
CUSOLVER_FP32_EMULATED_BF16X9_MATH.cusolverMpGetEmulationStrategy#
cusolverStatus_t cusolverMpGetEmulationStrategy(
cusolverMpHandle_t handle,
cudaEmulationStrategy_t *strategy)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
strategy |
Host |
Out |
Current emulation strategy of the handle. |
Grid Management#
cusolverMpCreateDeviceGrid#
cusolverStatus_t cusolverMpCreateDeviceGrid(
cusolverMpHandle_t handle,
cusolverMpGrid_t *grid,
ncclComm_t comm,
int32_t numRowDevices,
int32_t numColDevices,
cusolverMpGridMapping_t mapping)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
grid |
Host |
Out |
Grid object to be initialized. |
comm |
Host |
In |
Communicator that will be associated with the grid. |
numRowDevices |
Host |
In |
How many process rows the grid will contain. |
numColDevices |
Host |
In |
How many process columns the grid will contain. |
mapping |
Host |
In |
How to map processes to the grid. See description of cusolverMpGrid_t for further details. |
cusolverMpDestroyGrid#
cusolverStatus_t cusolverMpDestroyGrid(
cusolverMpGrid_t grid)
grid object.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
grid |
Host |
In/Out |
Grid object to be destroyed. |
Memory Management#
cusolverMpBufferRegister#
cusolverStatus_t cusolverMpBufferRegister(
cusolverMpGrid_t grid,
void *ptr,
size_t size)
ncclMemAlloc. Registration is idempotent: re-registering the same pointer with the same size is a no-op.size value on every rank. Re-registering the same pointer with a different size is invalid.cusolverMp<routine>_bufferSize) return each rank’s local requirement, which may differ across ranks. To register a workspace buffer, reduce the queried size to the grid-wide maximum (e.g., with MPI_Allreduce using MPI_MAX, or ncclAllReduce using ncclMax) before allocating and registering.ptr is not compatible with NCCL symmetric memory registration on any rank, this function returns CUSOLVER_STATUS_NOT_SUPPORTED on all ranks and registers nothing.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
grid |
Host |
In |
Grid object. |
ptr |
Device |
In |
Device buffer compatible with NCCL symmetric memory registration. |
size |
Host |
In |
Buffer size in bytes. |
cusolverMpBufferDeregister#
cusolverStatus_t cusolverMpBufferDeregister(
cusolverMpGrid_t grid,
void *ptr)
CUSOLVER_STATUS_INVALID_VALUE.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
grid |
Host |
In |
Grid object. |
ptr |
Device |
In |
Device buffer to deregister. |
cusolverMpMalloc#
cusolverStatus_t cusolverMpMalloc(
cusolverMpGrid_t grid,
void **ptr,
size_t size)
ncclMemAlloc and cusolverMpBufferRegister() into a single call. See cusolverMpBufferRegister() for details on how buffer registration affects the library’s communication paths.size value. If allocation or registration fails on any rank, *ptr is set to NULL on every rank, the same status is returned on every rank, and no memory is leaked.cusolverMp<routine>_bufferSize) return each rank’s local requirement, which may differ across ranks. To allocate a workspace with this function, reduce the queried size to the grid-wide maximum (e.g., with MPI_Allreduce using MPI_MAX, or ncclAllReduce using ncclMax) first.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
grid |
Host |
In |
Grid object. |
ptr |
Host |
Out |
Receives the allocated device buffer. |
size |
Host |
In |
Allocation size in bytes. |
cusolverMpFree#
cusolverStatus_t cusolverMpFree(
cusolverMpGrid_t grid,
void *ptr)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
grid |
Host |
In |
Grid object. |
ptr |
Device |
In |
Device buffer allocated with cusolverMpMalloc. |
Matrix Management#
cusolverMpCreateMatrixDesc#
cusolverStatus_t cusolverMpCreateMatrixDesc(
cusolverMpMatrixDescriptor_t *desc,
cusolverMpGrid_t grid,
cudaDataType dataType,
int64_t M_A,
int64_t N_A,
int64_t MB_A,
int64_t NB_A,
uint32_t RSRC_A,
uint32_t CSRC_A,
int64_t LLD_A)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
desc |
Host |
Out |
Matrix descriptor object initialized by this function. |
grid |
Host |
In |
Grid object associated with the global matrix A. |
dataType |
Host |
In |
Data type of the matrix A. |
M_A |
Host |
In |
Number of rows in the global matrix A. |
N_A |
Host |
In |
Number of columns in the global matrix A. |
MB_A |
Host |
In |
Blocking factor used to distribute the rows of the global matrix A. |
NB_A |
Host |
In |
Blocking factor used to distribute the columns of the global matrix A. |
RSRC_A |
Host |
In |
Process row over which the first row of the matrix A is distributed. Only the value of |
CSRC_A |
Host |
In |
Process column over which the first column of the matrix A is distributed. Only the value of |
LLD_A |
Host |
In |
Leading dimension of the local matrix. |
dataType argument are listed below:Data Type of A |
Description |
|---|---|
CUDA_R_16F |
Half precision real values. |
CUDA_R_16BF |
bfloat16 real values. |
CUDA_R_32I |
32-bit integer values. |
CUDA_R_64I |
64-bit integer values. |
CUDA_R_32F |
Single precision real values. |
CUDA_R_64F |
Double precision real values. |
CUDA_C_32F |
Single precision complex values. |
CUDA_C_64F |
Double precision complex values. |
cusolverMpDestroyMatrixDesc#
cusolverStatus_t cusolverMpDestroyMatrixDesc(
cusolverMpMatrixDescriptor_t desc)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
desc |
Host |
In/Out |
Matrix descriptor object destroyed by this function. |
Newton-Schulz Properties#
cusolverMpNewtonSchulzDescriptorCreate#
cusolverStatus_t cusolverMpNewtonSchulzDescriptorCreate(
cusolverMpNewtonSchulzDescriptor_t *nsDesc)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
nsDesc |
Host |
Out |
cusolverMpNewtonSchulzDescriptor_t descriptor to be created. |
cusolverMpNewtonSchulzDescriptorDestroy#
cusolverStatus_t cusolverMpNewtonSchulzDescriptorDestroy(
cusolverMpNewtonSchulzDescriptor_t nsDesc)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
nsDesc |
Host |
In/Out |
Newton-Schulz descriptor to be destroyed. |
cusolverMpNewtonSchulzDescriptorSetAttribute#
cusolverStatus_t cusolverMpNewtonSchulzDescriptorSetAttribute(
cusolverMpNewtonSchulzDescriptor_t nsDesc,
cusolverMpNewtonSchulzDescriptorAttribute_t attr,
const void *buf,
size_t sizeInBytes)
CUSOLVERMP_NEWTON_SCHULZ_DESCRIPTOR_ATTRIBUTE_NORMALIZE(int, default1): When set to1, the input matrix is normalized by its Frobenius norm before the iterations begin. Normalization is required for convergence. Set to0only when the input is already normalized (e.g., to avoid redundant normalization in a pipeline that pre-normalizes the matrix).
CUSOLVERMP_NEWTON_SCHULZ_DESCRIPTOR_ATTRIBUTE_REDUCE_VIA_COMPUTE_TYPE(int, default0): When set to1, the distributed Gram-matrix reduction path may communicate/reduce intermediateX^T Xdata using the compute type when the value type differs from the compute type.
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
nsDesc |
Host |
In/Out |
Newton-Schulz descriptor. |
attr |
Host |
In |
cusolverMpNewtonSchulzDescriptorAttribute_t attribute to set. |
buf |
Host |
In |
Pointer to the attribute value. |
sizeInBytes |
Host |
In |
Size of the attribute value in bytes. |
cusolverMpNewtonSchulzDescriptorGetAttribute#
cusolverStatus_t cusolverMpNewtonSchulzDescriptorGetAttribute(
cusolverMpNewtonSchulzDescriptor_t nsDesc,
cusolverMpNewtonSchulzDescriptorAttribute_t attr,
void *buf,
size_t sizeInBytes,
size_t *sizeInBytesWritten)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
nsDesc |
Host |
In |
Newton-Schulz descriptor. |
attr |
Host |
In |
cusolverMpNewtonSchulzDescriptorAttribute_t attribute to query. |
buf |
Host |
Out |
Buffer to receive the attribute value. |
sizeInBytes |
Host |
In |
Size of the output buffer in bytes. |
sizeInBytesWritten |
Host |
Out |
Number of bytes actually written to |
Polar Decomposition Properties#
cusolverMpPolarDescriptorCreate#
cusolverStatus_t cusolverMpPolarDescriptorCreate(
cusolverMpPolarDescriptor_t *polarDesc)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
polarDesc |
Host |
Out |
cusolverMpPolarDescriptor_t descriptor to be created. |
cusolverMpPolarDescriptorDestroy#
cusolverStatus_t cusolverMpPolarDescriptorDestroy(
cusolverMpPolarDescriptor_t polarDesc)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
polarDesc |
Host |
In/Out |
Polar decomposition descriptor to be destroyed. |
cusolverMpPolarDescriptorSetAttribute#
cusolverStatus_t cusolverMpPolarDescriptorSetAttribute(
cusolverMpPolarDescriptor_t polarDesc,
cusolverMpPolarDescriptorAttribute_t attr,
const void *buf,
size_t sizeInBytes)
CUSOLVERMP_POLAR_DESCRIPTOR_ATTRIBUTE_REQUESTED_KSI is a double input attribute. buf must point to a double and sizeInBytes must equal sizeof(double). Values greater than 0 request a diagonal perturbation in original unscaled units. Values less than or equal to 0 request no perturbation.double requested_ksi = 1.0e-6;
cusolverMpPolarDescriptorSetAttribute(
polarDesc,
CUSOLVERMP_POLAR_DESCRIPTOR_ATTRIBUTE_REQUESTED_KSI,
&requested_ksi,
sizeof(requested_ksi));
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
polarDesc |
Host |
In/Out |
Polar descriptor. |
attr |
Host |
In |
cusolverMpPolarDescriptorAttribute_t attribute to set. |
buf |
Host |
In |
Pointer to the attribute value. |
sizeInBytes |
Host |
In |
Size of the attribute value in bytes. |
cusolverMpPolarDescriptorGetAttribute#
cusolverStatus_t cusolverMpPolarDescriptorGetAttribute(
cusolverMpPolarDescriptor_t polarDesc,
cusolverMpPolarDescriptorAttribute_t attr,
void *buf,
size_t sizeInBytes,
size_t *sizeInBytesWritten)
double values. buf must provide at least sizeof(double) bytes, and sizeInBytesWritten is set to sizeof(double) on success.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
polarDesc |
Host |
In |
Polar descriptor. |
attr |
Host |
In |
cusolverMpPolarDescriptorAttribute_t attribute to query. |
buf |
Host |
Out |
Buffer to receive the attribute value. |
sizeInBytes |
Host |
In |
Size of the output buffer in bytes. |
sizeInBytesWritten |
Host |
Out |
Number of bytes actually written to |
Singular Value Decomposition Properties#
cusolverMpGesvdDescriptorCreate#
cusolverStatus_t cusolverMpGesvdDescriptorCreate(
cusolverMpGesvdDescriptor_t *gesvdDesc)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
gesvdDesc |
Host |
Out |
cusolverMpGesvdDescriptor_t descriptor to be created. |
cusolverMpGesvdDescriptorDestroy#
cusolverStatus_t cusolverMpGesvdDescriptorDestroy(
cusolverMpGesvdDescriptor_t gesvdDesc)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
gesvdDesc |
Host |
In/Out |
Singular value decomposition descriptor to be destroyed. |
cusolverMpGesvdDescriptorSetAttribute#
cusolverStatus_t cusolverMpGesvdDescriptorSetAttribute(
cusolverMpGesvdDescriptor_t gesvdDesc,
cusolverMpGesvdDescriptorAttribute_t attr,
const void *buf,
size_t sizeInBytes)
CUSOLVERMP_GESVD_DESCRIPTOR_ATTRIBUTE_COMPUTE_RESIDUAL(int, default0): When set to a nonzero value, successful non-empty calls that request bothUandV^HcomputeCUSOLVERMP_GESVD_DESCRIPTOR_ATTRIBUTE_RESIDUAL_FROBENIUS_ESTIMATEas the absolute, unnormalized reconstruction residual||A_original - U * Sigma * V^H||_F. Empty calls and calls that do not request both vector factors leave that output attribute as NaN.
CUSOLVERMP_GESVD_DESCRIPTOR_ATTRIBUTE_SHAPE(cusolverMpGesvdOutputShape_t, defaultCUSOLVERMP_GESVD_OUTPUT_SHAPE_THIN): Selects THIN or FULL singular-vector output shape.
CUSOLVER_STATUS_INVALID_VALUE.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
gesvdDesc |
Host |
In/Out |
Singular value decomposition descriptor. |
attr |
Host |
In |
cusolverMpGesvdDescriptorAttribute_t attribute to set. |
buf |
Host |
In |
Pointer to the attribute value. |
sizeInBytes |
Host |
In |
Size of the attribute value in bytes. |
cusolverMpGesvdDescriptorGetAttribute#
cusolverStatus_t cusolverMpGesvdDescriptorGetAttribute(
cusolverMpGesvdDescriptor_t gesvdDesc,
cusolverMpGesvdDescriptorAttribute_t attr,
void *buf,
size_t sizeInBytes,
size_t *sizeInBytesWritten)
double attributes) or 0 (int64_t attributes) at the entry of every cusolverMpGesvd() call, and may be queried after that call returns.sizeInBytesWritten is set to 0 and the routine returns CUSOLVER_STATUS_INVALID_VALUE.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
gesvdDesc |
Host |
In |
Singular value decomposition descriptor. |
attr |
Host |
In |
cusolverMpGesvdDescriptorAttribute_t attribute to query. |
buf |
Host |
Out |
Buffer to receive the attribute value. |
sizeInBytes |
Host |
In |
Size of the output buffer in bytes. |
sizeInBytesWritten |
Host |
Out |
Number of bytes actually written to |
Utility#
cusolverMpNUMROC#
int64_t cusolverMpNUMROC(
int64_t n,
int64_t nb,
uint32_t iproc,
uint32_t isrcproc,
uint32_t nprocs)
iproc argument.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
n |
Host |
In |
Number of rows or columns in the global distributed matrix. |
nb |
Host |
In |
Row or column blocking size of the global matrix. |
iproc |
Host |
In |
The coordinate of the process whose local array row or column is to be determined. |
isrcproc |
Host |
In |
The coordinate of the process that owns the first row or column of the distributed matrix. |
nprocs |
Host |
In |
The total number of row or column processes over which the matrix is distributed. |
iproc argument.cusolverMpMatrixGatherD2H#
cusolverStatus_t cusolverMpMatrixGatherD2H(
cusolverMpHandle_t handle,
int64_t M,
int64_t N,
const void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
int root,
void *h_dst,
int64_t h_lddst)
A on a buffer provided on process root. The input matrix A is originally distributed using 2D block-cyclic format, on output h_dst contains the matrix in column-major format.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
M |
Host |
In |
Number of rows of the global distributed matrix A. |
N |
Host |
In |
Number of columns of the global distributed matrix A. |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index in the global matrix A indicating the first row of sub(A). This function does not make any assumptions on the alignment of |
JA |
Host |
In |
Column index in the global matrix A indicating the first column of sub(A). This function does not make any assumptions on the alignment of |
descA |
Host |
In |
Matrix descriptor of the global matrix A. |
root |
Host |
In |
Process ID on which the matrix A will be gathered. |
h_dst |
Host |
Out |
Destination host buffer on |
h_lddst |
Host |
In |
Leading dimension of the |
Warning
This function is meant as a utility function to verify correctness of the data layouts and it is not intended to achieve high performance on large inputs.
cusolverMpMatrixScatterH2D#
cusolverStatus_t cusolverMpMatrixScatterH2D(
cusolverMpHandle_t handle,
int64_t M,
int64_t N,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
int root,
const void *h_src,
int64_t h_ldsrc)
h_src from root process to a distributed global matrix A.h_src is stored in column-major format. On output, d_A contains the local portions of the global matrix A distributed in 2D block-cyclic format.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
M |
Host |
In |
Number of rows of the global distributed matrix A. |
N |
Host |
In |
Number of columns of the global distributed matrix A. |
d_A |
Device |
Out |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index in the global matrix A indicating the first row of sub(A). This function does not make any assumptions on the alignment of |
JA |
Host |
In |
Column index in the global matrix A indicating the first column of sub(A). This function does not make any assumptions on the alignment of |
descA |
Host |
In |
Matrix descriptor of the global matrix A. |
root |
Host |
In |
Process ID which the matrix A will be scattered from. |
h_src |
Host |
In |
Source buffer on |
h_ldsrc |
Host |
In |
Leading dimension of the |
Warning
This function is meant as a utility function to verify correctness of the data layouts and it is not intended to achieve high performance on large inputs.
Logging#
cusolverMpLoggerSetCallback#
cusolverStatus_t cusolverMpLoggerSetCallback(
cusolverMpLoggerCallback_t callback)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
callback |
Host |
In |
Pointer to a callback function. See cusolverMpLoggerCallback_t. |
Warning
This is an experimental feature.
cusolverMpLoggerSetFile#
cusolverStatus_t cusolverMpLoggerSetFile(
FILE *file)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
file |
Host |
In |
Pointer to an open file. File should have write permission. |
Warning
This is an experimental feature.
cusolverMpLoggerOpenFile#
cusolverStatus_t cusolverMpLoggerOpenFile(
const char* logFile)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
logFile |
Host |
In |
Path of the logging output file. |
Warning
This is an experimental feature.
cusolverMpLoggerSetLevel#
cusolverStatus_t cusolverMpLoggerSetLevel(
int level)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
level |
Host |
In |
Value of the logging level. See |
Warning
This is an experimental feature.
cusolverMpLoggerSetMask#
cusolverStatus_t cusolverMpLoggerSetMask(
int mask)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
mask |
Host |
In |
Value of the logging mask. See |
Warning
This is an experimental feature.
cusolverMpLoggerForceDisable#
cusolverStatus_t cusolverMpLoggerForceDisable()
Warning
This is an experimental feature.
Dense Linear Algebra APIs#
Note
For every routine in this section, the returned cusolverStatus_t is the
authoritative API status and must always be checked. Runtime device info
outputs are optional; pass NULL to skip device-side info reporting. When a
non-NULL info is supplied, successful calls reset it to 0 unless
the routine reports a routine-specific positive value. For routines that
define positive info values, a successful host status can still accompany
info > 0; passing NULL suppresses those routine-specific reports, so
provide info when singularity or convergence diagnostics are required.
Malformed arguments return CUSOLVER_STATUS_INVALID_VALUE and, with a
non-NULL info output, write a negative value identifying the offending
API argument after handle. Valid but unsupported configurations return
CUSOLVER_STATUS_NOT_SUPPORTED. Invalid handles are reported by the
returned status. Use the routine-specific info parameter description for
routine-specific positive values and whether those values accompany
CUSOLVER_STATUS_SUCCESS or an error status.
cusolverMpGetrf#
cusolverStatus_t cusolverMpGetrf(
cusolverMpHandle_t handle,
int64_t M,
int64_t N,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
int64_t *d_ipiv,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *info)
d_ipiv=NULL.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
M |
Host |
In |
Number of rows of sub(A). |
N |
Host |
In |
Number of columns of sub(A). |
d_A |
Device |
In/Out |
Pointer to the first entry of the local portion of the global matrix A. On output, the sub(A) is overwritten with the L and U factors. |
IA |
Host |
In |
Row index of the first row of the sub(A). This function does not make any assumptions on the alignment of |
JA |
Host |
In |
Column index of the first column of the sub(A). This function does not make any assumptions on the alignment of |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_ipiv |
Device |
Out |
Local array of dimension |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpGetrf_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpGetrf_bufferSize(). |
info |
Device |
Out |
Optional. |
Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpGetrf_bufferSize#
cusolverStatus_t cusolverMpGetrf_bufferSize(
cusolverMpHandle_t handle,
int64_t M,
int64_t N,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
int64_t *d_ipiv,
cudaDataType_t computeType,
size_t *workspaceInBytesOnDevice,
size_t *workspaceInBytesOnHost)
d_ipiv=NULL so cusolverMpGetrf() will compute the LU factorization of the input matrix A without pivoting.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
M |
Host |
In |
Number of rows of sub(A). |
N |
Host |
In |
Number of columns of sub(A). |
d_A |
Device |
In |
Pointer to the first entry of the local portion of the global matrix A. |
IA |
Host |
In |
Row index of the first row of the sub(A). This function does not make any assumptions on the alignment of |
JA |
Host |
In |
Column index of the first column of the sub(A). This function does not make any assumptions on the alignment of |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_ipiv |
Device |
In |
Indicates a pointer to a distributed integer array. When it is not |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
On output, contains the size in bytes of the local device workspace needed by cusolverMpGetrf(). |
workspaceInBytesOnHost |
Host |
Out |
On output, contains the size in bytes of the local host workspace needed by cusolverMpGetrf(). |
Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpGetrs#
cusolverStatus_t cusolverMpGetrs(
cusolverMpHandle_t handle,
cublasOperation_t trans,
int64_t N,
int64_t NRHS,
const void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
const int64_t *d_ipiv,
void *d_B,
int64_t IB,
int64_t JB,
cusolverMpMatrixDescriptor_t descB,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *d_info)
trans, which allows to solve linear systems of the form:trans |
Form of the linear system |
|---|---|
CUBLAS_OP_N |
\(sub(A) \cdot X = sub(B)\) |
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
trans |
Host |
In |
Specifies the form of the linear system. Only |
N |
Host |
In |
Number of rows of sub(A). |
NRHS |
Host |
In |
Number of columns of sub(B). |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). This function does not make any assumptions on the alignment of |
JA |
Host |
In |
Column index of the first column of the sub(A). This function does not make any assumptions on the alignment of |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_ipiv |
Device |
In |
Local array of dimension |
d_B |
Device |
In/Out |
Pointer into the local memory to an array of dimension |
IB |
Host |
In |
Row index of the first row of the sub(B). This function does not require |
JB |
Host |
In |
Column index of the first column of the sub(B). This function does not make any assumptions on the alignment of |
descB |
Host |
In |
Matrix descriptor associated to the global matrix B. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpGetrs_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpGetrs_bufferSize(). |
info |
Device |
Out |
Optional. |
Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpGetrs_bufferSize#
cusolverStatus_t cusolverMpGetrs_bufferSize(
cusolverMpHandle_t handle,
cublasOperation_t trans,
int64_t N,
int64_t NRHS,
const void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
const int64_t *d_ipiv,
void *d_B,
int64_t IB,
int64_t JB,
cusolverMpMatrixDescriptor_t descB,
cudaDataType_t computeType,
size_t *workspaceInBytesOnDevice,
size_t *workspaceInBytesOnHost)
d_ipiv=NULL.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
trans |
Host |
In |
Specifies the form of the linear system. Only |
N |
Host |
In |
Number of rows of sub(A). |
NRHS |
Host |
In |
Number of columns of sub(B). |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). This function does not make any assumptions on the alignment of |
JA |
Host |
In |
Column index of the first column of the sub(A). This function does not make any assumptions on the alignment of |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_ipiv |
Device |
In |
Local array of dimension |
d_B |
Device |
In |
Pointer to the first entry of the local portion of the global matrix B. On output, B is overwritten the solution of the linear system. |
IB |
Host |
In |
Row index of the first row of the sub(B). The corresponding cusolverMpGetrs() call does not require |
JB |
Host |
In |
Column index of the first column of the sub(B). This function does not make any assumptions on the alignment of |
descB |
Host |
In |
Matrix descriptor associated to the global matrix B. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
On output, contains the size in bytes of the local device workspace needed by cusolverMpGetrs(). |
workspaceInBytesOnHost |
Host |
Out |
On output, contains the size in bytes of the local host workspace needed by cusolverMpGetrs(). |
Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpPotrf#
cusolverStatus_t cusolverMpPotrf(
cusolverMpHandle_t handle,
cublasFillMode_t uplo,
int64_t N,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *info)
A(IA:IA+N-1, JA:JA+N-1).uplo=CUBLAS_FILL_MODE_UPPER, the factorization has the formuplo is set to CUBLAS_FILL_MODE_LOWER, the factorization has the formParameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
uplo |
Host |
In |
Specifies if A is upper ( |
N |
Host |
In |
Number of rows and columns of sub(A). |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). |
JA |
Host |
In |
Column index of the first column of the sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpPotrf_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpPotrf_bufferSize(). |
info |
Device |
Out |
Optional. |
(MB_A == NB_A).Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpPotrf_bufferSize#
cusolverStatus_t cusolverMpPotrf_bufferSize(
cusolverMpHandle_t handle,
cublasFillMode_t uplo,
int64_t N,
const void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
cudaDataType_t computeType,
size_t* workspaceInBytesOnDevice,
size_t* workspaceInBytesOnHost)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
uplo |
Host |
In |
Specifies if A is upper ( |
N |
Host |
In |
Number of rows and columns of sub(A). |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). This function does not make any assumptions on the alignment of |
JA |
Host |
In |
Column index of the first column of the sub(A). This function does not make any assumptions on the alignment of |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
On output, contains the size in bytes of the local device workspace needed by cusolverMpPotrf(). |
workspaceInBytesOnHost |
Host |
Out |
On output, contains the size in bytes of the local host workspace needed by cusolverMpPotrf(). |
Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpPotrs#
cusolverStatus_t cusolverMpPotrs(
cusolverMpHandle_t handle,
cublasFillMode_t uplo,
int64_t N,
int64_t NRHS,
const void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
void *d_B,
int64_t IB,
int64_t JB,
cusolverMpMatrixDescriptor_t descB,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *info)
A(IA:IA+N-1,JA:JA+N-1) and is a N-by-N symmetric or hermitian positive definite distributed matrix using the Cholesky factorization:\[sub(A) = U^H \cdot U\]
B(IB:IB+N-1,JB:JB+NRHS-1).Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
uplo |
Host |
In |
Specifies if A is upper ( |
N |
Host |
In |
Number of rows and columns of sub(A). |
NRHS |
Host |
In |
Number of columns of sub(B). |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). |
JA |
Host |
In |
Column index of the first column of the sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_B |
Device |
In/Out |
Pointer into the local memory to an array of dimension |
IB |
Host |
In |
Row index of the first row of the sub(B). This function does not make any assumptions on the alignment of |
JB |
Host |
In |
Column index of the first column of the sub(B). This function does not make any assumptions on the alignment of |
descB |
Host |
In |
Matrix descriptor associated to the global matrix B. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpPotrs_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpPotrs_bufferSize(). |
info |
Device |
Out |
Optional. |
(MB_A == NB_A) and alignment of sub(A) and sub(B) matrices, meaning (MB_A == MB_B) and (IA == IB).Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
—
cusolverMpPotrs_bufferSize#
cusolverStatus_t cusolverMpPotrs_bufferSize(
cusolverMpHandle_t handle,
cublasFillMode_t uplo,
int64_t n,
int64_t nrhs,
const void *a,
int64_t ia,
int64_t ja,
cusolverMpMatrixDescriptor_t descA,
const void *b,
int64_t ib,
int64_t jb,
cusolverMpMatrixDescriptor_t descB,
cudaDataType_t computeType,
size_t* workspaceInBytesOnDevice,
size_t* workspaceInBytesOnHost)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
uplo |
Host |
In |
Specifies if A is upper ( |
N |
Host |
In |
Number of rows and columns of sub(A). |
NRHS |
Host |
In |
Number of columns of sub(B). |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). |
JA |
Host |
In |
Column index of the first column of the sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_B |
Device |
In |
Pointer into the local memory to an array of dimension |
IB |
Host |
In |
Row index of the first row of the sub(B). This function does not make any assumptions on the alignment of |
JB |
Host |
In |
Column index of the first column of the sub(B). This function does not make any assumptions on the alignment of |
descB |
Host |
In |
Matrix descriptor associated to the global matrix B. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
On output, contains the size in bytes of the local device workspace needed by cusolverMpPotrs(). |
workspaceInBytesOnHost |
Host |
Out |
On output, contains the size in bytes of the local host workspace needed by cusolverMpPotrs(). |
(MB_A == NB_A) and alignment of sub(A) and sub(B) matrices, meaning (MB_A == MB_B) and (IA == IB).Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
—
cusolverMpGeqrf#
cusolverStatus_t cusolverMpGeqrf(
cusolverMpHandle_t handle,
int64_t M,
int64_t N,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
void *d_tau,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *info)
A(IA:IA+M-1, JA:JA+N-1).tau and R is upper triangular matrix.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
M |
Host |
In |
Number of rows of sub(A). |
N |
Host |
In |
Number of columns of sub(A). |
d_A |
Device |
In/Out |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). |
JA |
Host |
In |
Column index of the first column of the sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_tau |
Device |
Out |
Pointer into the local memory to an array of dimension |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpGeqrf_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpGeqrf_bufferSize(). |
info |
Device |
Out |
Optional. |
(MB_A == NB_A).Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpGeqrf_bufferSize#
cusolverStatus_t cusolverMpGeqrf_bufferSize(
cusolverMpHandle_t handle,
int64_t M,
int64_t N,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
cudaDataType_t computeType,
size_t* workspaceInBytesOnDevice,
size_t* workspaceInBytesOnHost)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
M |
Host |
In |
Number of rows of sub(A). |
N |
Host |
In |
Number of columns of sub(A). |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index in the global matrix A indicating the first row of sub(A). This function does not make any assumptions on the alignment of |
JA |
Host |
In |
Column index in the global matrix A indicating the first column of sub(A). This function does not make any assumptions on the alignment of |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
On output, contains the size in bytes of the local device workspace needed by cusolverMpGeqrf(). |
workspaceInBytesOnHost |
Host |
Out |
On output, contains the size in bytes of the local host workspace needed by cusolverMpGeqrf(). |
Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
—
cusolverMpOrmqr#
cusolverStatus_t cusolverMpOrmqr(
cusolverMpHandle_t handle,
cublasSideMode_t side,
cublasOperation_t trans,
int64_t M,
int64_t N,
int64_t K,
const void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
const void *d_tau,
void *d_C,
int64_t IC,
int64_t JC,
cusolverMpMatrixDescriptor_t descC,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *info)
C(IC:IC+M-1, JC:JC+N-1) by the orthogonal matrix Q can be given from cusolverMpGeqrf().side of CUBLAS_SIDE_LEFT and CUBLAS_SIDE_RIGHT respectively. Currently, only CUBLAS_SIDE_LEFT is supported.K <= M and K <= N for CUBLAS_SIDE_LEFT and CUBLAS_SIDE_RIGHT respectively.op can be translated to \(Q\), \(Q^T\), \(Q^H\) based on the trans argument. Real types support CUBLAS_OP_N and CUBLAS_OP_T; complex types support CUBLAS_OP_N and CUBLAS_OP_C.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
side |
Host |
In |
Indicate that Q is applied from left or right side. |
trans |
Host |
In |
Indicate that Q is applied with no-transpose or (conj)transpose. Real types support |
M |
Host |
In |
Number of rows of sub(C). |
N |
Host |
In |
Number of columns of sub(C). |
K |
Host |
In |
Number of Householder reflectors defining Q. |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
1-based row index of the first row of sub(A). |
JA |
Host |
In |
1-based column index of the first column of sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_tau |
Device |
In |
Pointer into the local memory to an array of dimension |
d_C |
Device |
In/Out |
Pointer into the local memory to an array of dimension |
IC |
Host |
In |
1-based row index of the first row of sub(C). |
JC |
Host |
In |
1-based column index of the first column of sub(C). |
descC |
Host |
In |
Matrix descriptor associated to the global matrix C. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpOrmqr_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpOrmqr_bufferSize(). |
info |
Device |
Out |
Optional. |
(MB_A == NB_A) and alignment of sub(A) and sub(C) matrices, meaning (MB_A == MB_C) and (IA == IC).Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpOrmqr_bufferSize#
cusolverStatus_t cusolverMpOrmqr_bufferSize(
cusolverMpHandle_t handle,
cublasSideMode_t side,
cublasOperation_t trans,
int64_t M,
int64_t N,
int64_t K,
const void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
const void *d_tau,
void *d_C,
int64_t IC,
int64_t JC,
cusolverMpMatrixDescriptor_t descC,
cudaDataType_t computeType,
size_t* workspaceInBytesOnDevice,
size_t* workspaceInBytesOnHost)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
side |
Host |
In |
Indicate that Q is applied from left or right side. |
trans |
Host |
In |
Indicate that Q is applied with no-transpose or (conj)transpose. Real types support |
M |
Host |
In |
Number of rows of sub(C). |
N |
Host |
In |
Number of columns of sub(C). |
K |
Host |
In |
Number of Householder reflectors defining Q. |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
1-based row index of the first row of sub(A). |
JA |
Host |
In |
1-based column index of the first column of sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_tau |
Device |
In |
Pointer into the local memory to an array of dimension |
d_C |
Device |
In |
Pointer into the local memory to an array of dimension |
IC |
Host |
In |
1-based row index of the first row of sub(C). |
JC |
Host |
In |
1-based column index of the first column of sub(C). |
descC |
Host |
In |
Matrix descriptor associated to the global matrix C. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
The size in bytes of the local device workspace needed by cusolverMpOrmqr(). |
workspaceInBytesOnHost |
Host |
Out |
The size in bytes of the local host workspace needed by cusolverMpOrmqr(). |
Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpOrgqr#
cusolverStatus_t cusolverMpOrgqr(
cusolverMpHandle_t handle,
int64_t M,
int64_t N,
int64_t K,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
const void *d_tau,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *d_info)
M-by-N matrix Q with orthonormal columns from the QR factorization computed by cusolverMpGeqrf(). Q is defined as the product of K elementary Householder reflectors of order M:H(i) are the elementary reflectors stored in the lower triangular part of A(IA:IA+M-1, JA:JA+K-1) as returned by cusolverMpGeqrf(), with corresponding scalar factors in d_tau.M >= N >= K >= 0. When K = 0, the routine sets Q to the identity matrix.d_A contains the Householder reflectors and d_tau as output by cusolverMpGeqrf(). On output, the submatrix A(IA:IA+M-1, JA:JA+N-1) is overwritten with the first N columns of Q.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
M |
Host |
In |
Number of rows of the matrix Q. |
N |
Host |
In |
Number of columns of the matrix Q. |
K |
Host |
In |
Number of elementary reflectors. |
d_A |
Device |
In/Out |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the submatrix. |
JA |
Host |
In |
Column index of the first column of the submatrix. |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_tau |
Device |
In |
Pointer into the local memory to an array of dimension |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpOrgqr_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpOrgqr_bufferSize(). |
d_info |
Device |
Out |
Optional. |
Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpOrgqr_bufferSize#
cusolverStatus_t cusolverMpOrgqr_bufferSize(
cusolverMpHandle_t handle,
int64_t M,
int64_t N,
int64_t K,
const void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
const void *d_tau,
cudaDataType_t computeType,
size_t *workspaceInBytesOnDevice,
size_t *workspaceInBytesOnHost)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
M |
Host |
In |
Number of rows of the matrix Q. |
N |
Host |
In |
Number of columns of the matrix Q. |
K |
Host |
In |
Number of elementary reflectors. |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the submatrix. |
JA |
Host |
In |
Column index of the first column of the submatrix. |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_tau |
Device |
In |
Pointer into the local memory to an array of dimension |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
On output, contains the size in bytes of the local device workspace needed by cusolverMpOrgqr(). |
workspaceInBytesOnHost |
Host |
Out |
On output, contains the size in bytes of the local host workspace needed by cusolverMpOrgqr(). |
Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpGels#
cusolverStatus_t cusolverMpGels(
cusolverMpHandle_t handle,
cublasOperation_t trans,
int64_t M,
int64_t N,
int64_t NRHS,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
void *d_B,
int64_t IB,
int64_t JB,
cusolverMpMatrixDescriptor_t descB,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *info)
A(IA:IA+M-1, JA:JA+N-1) or its transpose, using QR or LQ factorization of sub(A).M >= N) with a no-transpose option is only supported via QR factorization cusolverMpGeqrf().B(IB:IB+M-1, JB:JB+NRHS-1) and the solution multi-vector X is overwritten on the sub(B).Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
trans |
Host |
In |
Indicate that the linear system of sub(A) involves with no-transpose or (conj)transpose. |
M |
Host |
In |
Number of rows of sub(A). |
N |
Host |
In |
Number of columns of sub(A). |
NRHS |
Host |
In |
Number of right hand side vectors i.e., number of columns of sub(B) and X. |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). |
JA |
Host |
In |
Column index of the first column of the sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_B |
Device |
In/Out |
Pointer into the local memory to an array of dimension |
IB |
Host |
In |
Row index of the first row of the sub(B). |
JB |
Host |
In |
Column index of the first column of the sub(B). |
descB |
Host |
In |
Matrix descriptor associated to the global matrix B. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpGels_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpGels_bufferSize(). |
info |
Device |
Out |
Optional. |
(MB_A == NB_A) and alignment of sub(A) and sub(B) matrices, meaning (MB_A == MB_B) and (IA == IB).Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpGels_bufferSize#
cusolverStatus_t cusolverMpGels_bufferSize(
cusolverMpHandle_t handle,
cublasOperation_t trans,
int64_t M,
int64_t N,
int64_t NRHS,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
void *d_B,
int64_t IB,
int64_t JB,
cusolverMpMatrixDescriptor_t descB,
cudaDataType_t computeType,
size_t* workspaceInBytesOnDevice,
size_t* workspaceInBytesOnHost)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
trans |
Host |
In |
Indicate that the linear system of sub(A) involves with no-transpose or (conj)transpose. |
M |
Host |
In |
Number of rows of sub(A). |
N |
Host |
In |
Number of columns of sub(A). |
NRHS |
Host |
In |
Number of right hand side vectors i.e., number of columns of sub(B) and X. |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). |
JA |
Host |
In |
Column index of the first column of the sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_B |
Device |
In |
Pointer into the local memory to an array of dimension |
IB |
Host |
In |
Row index of the first row of the sub(B). |
JB |
Host |
In |
Column index of the first column of the sub(B). |
descB |
Host |
In |
Matrix descriptor associated to the global matrix B. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
The size in bytes of the local device workspace needed by cusolverMpGels(). |
workspaceInBytesOnHost |
Host |
Out |
The size in bytes of the local host workspace needed by cusolverMpGels(). |
Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpSytrd#
cusolverStatus_t cusolverMpSytrd(
cusolverMpHandle_t handle,
cublasFillMode_t uplo,
int64_t N,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
void *d_d,
void *d_e,
void *d_tau,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *info)
A(IA:IA+N-1, JA:JA+N-1) to a tridiagonal form.cusolverMpSytrd() API supports only CUBLAS_FILL_MODE_LOWER. The routine stores Householder reflectors in lower form; use CUBLAS_FILL_MODE_LOWER when applying these reflectors with cusolverMpOrmtr().ncclMemAlloc registered via cusolverMpBufferRegister(), before calling this routine.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
uplo |
Host |
In |
Indicate which triangular part of sub(A) is used. Currently, only |
N |
Host |
In |
Number of rows/columns of square matrix sub(A). |
d_A |
Device |
In/Out |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
1-based row index of the first row of sub(A). Current support requires |
JA |
Host |
In |
1-based column index of the first column of sub(A). Current support requires |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_d |
Device |
Out |
Pointer into the local memory to an array of dimension |
d_e |
Device |
Out |
Pointer into the local memory to an array of dimension |
d_tau |
Device |
Out |
Pointer into the local memory to an array of dimension |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpSytrd_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpSytrd_bufferSize(). |
info |
Device |
Out |
Optional. |
(MB_A == NB_A) and block-aligned starts, i.e. (IA - 1) is a multiple of MB_A and (JA - 1) is a multiple of NB_A.Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpSytrd_bufferSize#
cusolverStatus_t cusolverMpSytrd_bufferSize(
cusolverMpHandle_t handle,
cublasFillMode_t uplo,
int64_t N,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
void *d_d,
void *d_e,
void *d_tau,
cudaDataType_t computeType,
size_t *workspaceInBytesOnDevice,
size_t *workspaceInBytesOnHost)
CUBLAS_FILL_MODE_LOWER is supported, MB_A == NB_A, and starts must be block-aligned.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
uplo |
Host |
In |
Indicate which triangular part of sub(A) is used. Currently, only |
N |
Host |
In |
Number of rows/columns of square matrix sub(A). |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
1-based row index of the first row of sub(A). Current support requires |
JA |
Host |
In |
1-based column index of the first column of sub(A). Current support requires |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_d |
Device |
In |
Pointer into the local memory to an array of dimension |
d_e |
Device |
In |
Pointer into the local memory to an array of dimension |
d_tau |
Device |
In |
Pointer into the local memory to an array of dimension |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
The size in bytes of the local device workspace needed by cusolverMpSytrd(). |
workspaceInBytesOnHost |
Host |
Out |
The size in bytes of the local host workspace needed by cusolverMpSytrd(). |
Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpStedc#
cusolverStatus_t cusolverMpStedc(
cusolverMpHandle_t handle,
char *compz,
int64_t N,
void *d_D,
void *d_E,
void *d_Q,
int64_t IQ,
int64_t JQ,
cusolverMpMatrixDescriptor_t descQ,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *info)
compz=N) or all eigenvalues and eigenvectors (compz=I) of a symmetric tridiagonal matrix using the divide and conquer algorithm.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
compz |
Host |
In |
Option to compute eigenvalues only ( |
N |
Host |
In |
Number of rows/columns of square matrix sub(A). |
d_D |
Device |
In/Out |
Pointer to an array of dimension |
d_E |
Device |
In/Out |
Pointer to an array of dimension |
d_Q |
Device |
Out |
Pointer into the local memory to an array of dimension |
IQ |
Host |
In |
1-based row index of the first row of sub(Q). Current support requires |
JQ |
Host |
In |
1-based column index of the first column of sub(Q). Current support requires |
descQ |
Host |
In |
Matrix descriptor associated to the global matrix Q. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpStedc_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpStedc_bufferSize(). |
info |
Device |
Out |
Optional. |
compz = I or compz = N, square block size for Q (MB_Q == NB_Q), and a block-aligned Q column start, i.e. (JQ - 1) is a multiple of NB_Q.Data Type of Tridiagonal Matrix |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpStedc_bufferSize#
cusolverStatus_t cusolverMpStedc_bufferSize(
cusolverMpHandle_t handle,
char *compz,
int64_t N,
void *d_D,
void *d_E,
void *d_Q,
int64_t IQ,
int64_t JQ,
cusolverMpMatrixDescriptor_t descQ,
cudaDataType_t computeType,
size_t *workspaceInBytesOnDevice,
size_t *workspaceInBytesOnHost,
int *iwork)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
compz |
Host |
In |
Option to compute eigenvalues only ( |
N |
Host |
In |
Number of rows/columns of square matrix sub(A). |
d_D |
Device |
In |
Pointer to an array of dimension |
d_E |
Device |
In |
Pointer to an array of dimension |
d_Q |
Device |
In |
Pointer into the local memory to an array of dimension |
IQ |
Host |
In |
1-based row index of the first row of sub(Q). Current support requires |
JQ |
Host |
In |
1-based column index of the first column of sub(Q). Current support requires |
descQ |
Host |
In |
Matrix descriptor associated to the global matrix Q. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
The size in bytes of the local device workspace needed by the routine cusolverMpStedc(). |
workspaceInBytesOnHost |
Host |
Out |
The size in bytes of the local host workspace needed by cusolverMpStedc(). |
iwork |
Host |
In/Out |
Host scratch buffer used internally during buffer-size computation. Must be pre-allocated by the caller with at least |
Data Type of Tridiagonal Matrix |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpOrmtr#
cusolverStatus_t cusolverMpOrmtr(
cusolverMpHandle_t handle,
cublasSideMode_t side,
cublasFillMode_t uplo,
cublasOperation_t trans,
int64_t M,
int64_t N,
const void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
const void *d_tau,
void *d_C,
int64_t IC,
int64_t JC,
cusolverMpMatrixDescriptor_t descC,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *info)
C(IC:IC+M-1, JC:JC+N-1) by the orthogonal matrix Q can be given from cusolverMpSytrd().side = CUBLAS_SIDE_LEFT and uplo = CUBLAS_FILL_MODE_LOWER, where uplo describes the storage of the Householder reflectors in sub(A). It performs the following matrix product and overwrites the result on sub(C):op can be translated to \(Q\), \(Q^T\), \(Q^H\) based on the trans argument. Real types support CUBLAS_OP_N and CUBLAS_OP_T; complex types support CUBLAS_OP_N and CUBLAS_OP_C.CUBLAS_FILL_MODE_LOWER, Q is an orthogonal matrix formed as the following product of Householder reflectors:nq is m for the supported CUBLAS_SIDE_LEFT case.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
side |
Host |
In |
Indicate that Q is applied from left or right side. Currently, only |
uplo |
Host |
In |
Indicate whether upper or lower triangular of sub(A) contains Householder reflectors. Currently, only |
trans |
Host |
In |
Indicate that Q is applied with no-transpose or (conj)transpose. Real types support |
M |
Host |
In |
Number of rows of sub(C) and order of Q for the supported left-side case. |
N |
Host |
In |
Number of columns of sub(C). |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
1-based row index of the first row of sub(A). |
JA |
Host |
In |
1-based column index of the first column of sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_tau |
Device |
In |
Pointer into the local memory to an array of dimension |
d_C |
Device |
In/Out |
Pointer into the local memory to an array of dimension |
IC |
Host |
In |
1-based row index of the first row of sub(C). |
JC |
Host |
In |
1-based column index of the first column of sub(C). |
descC |
Host |
In |
Matrix descriptor associated to the global matrix C. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpOrmtr_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpOrmtr_bufferSize(). |
info |
Device |
Out |
Optional. |
(MB_A == NB_A) and alignment of sub(A) and sub(B) matrices, meaning (MB_A == MB_C) and (IA == IC).Data Type of A and C |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpOrmtr_bufferSize#
cusolverStatus_t cusolverMpOrmtr_bufferSize(
cusolverMpHandle_t handle,
cublasSideMode_t side,
cublasFillMode_t uplo,
cublasOperation_t trans,
int64_t M,
int64_t N,
const void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
const void *d_tau,
void *d_C,
int64_t IC,
int64_t JC,
cusolverMpMatrixDescriptor_t descC,
cudaDataType_t computeType,
size_t* workspaceInBytesOnDevice,
size_t* workspaceInBytesOnHost)
side = CUBLAS_SIDE_LEFT and uplo = CUBLAS_FILL_MODE_LOWER reflector storage are implemented; A and C require compatible communicator and process-grid properties; MB_A == MB_C is required; descriptor grids with multiple ranks require IA = JA = IC = 1; and a non-unit JC is supported only when the multi-column-rank tile-span and active-owner-containment requirements are both satisfied.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
side |
Host |
In |
Indicate that Q is applied from left or right side. Currently, only |
uplo |
Host |
In |
Indicate whether upper or lower triangular of sub(A) contains Householder reflectors. Currently, only |
trans |
Host |
In |
Indicate that Q is applied with no-transpose or (conj)transpose. Real types support |
M |
Host |
In |
Number of rows of sub(C) and order of Q for the supported left-side case. |
N |
Host |
In |
Number of columns of sub(C). |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
1-based row index of the first row of sub(A). |
JA |
Host |
In |
1-based column index of the first column of sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_tau |
Device |
In |
Pointer into the local memory to an array of dimension |
d_C |
Device |
In |
Pointer into the local memory to an array of dimension |
IC |
Host |
In |
1-based row index of the first row of sub(C). |
JC |
Host |
In |
1-based column index of the first column of sub(C). |
descC |
Host |
In |
Matrix descriptor associated to the global matrix C. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
The size in bytes of the local device workspace needed by cusolverMpOrmtr(). |
workspaceInBytesOnHost |
Host |
Out |
The size in bytes of the local host workspace needed by cusolverMpOrmtr(). |
Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpSyevd#
cusolverStatus_t cusolverMpSyevd(
cusolverMpHandle_t handle,
char *jobz,
cublasFillMode_t uplo,
int64_t N,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
void *d_D,
void *d_Q,
int64_t IQ,
int64_t JQ,
cusolverMpMatrixDescriptor_t descQ,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *d_info)
A(IA:IA+N-1, JA:JA+N-1) using the divide and conquer algorithm cusolverMpStedc().ncclMemAlloc registered via cusolverMpBufferRegister(), before calling this routine.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
jobz |
Host |
In |
If |
uplo |
Host |
In |
Indicate that upper or lower triangular of sub(A) is used to compute eigen solutions. |
N |
Host |
In |
Number of rows and columns of sub(A). |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). |
JA |
Host |
In |
Column index of the first column of the sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_D |
Device |
Out |
Pointer to a device buffer of length |
d_Q |
Device |
Out |
Pointer into the local memory to an array of dimension |
IQ |
Host |
In |
Row index of the first row of the sub(Q). |
JQ |
Host |
In |
Column index of the first column of the sub(Q). |
descQ |
Host |
In |
Matrix descriptor associated to the global matrix Q. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpSyevd_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpSyevd_bufferSize(). |
d_info |
Device |
Out |
Optional. |
(MB_A == NB_A) and alignment of sub(A) and sub(Q) matrices, meaning (MB_A == MB_Q) and (IA == IQ). The current implementation supports full-matrix submatrix starts (IA == JA == IQ == JQ == 1).uplo selects which triangular part of the input sub(A) is read, and both CUBLAS_FILL_MODE_UPPER and CUBLAS_FILL_MODE_LOWER are supported. The standalone cusolverMpSytrd() and cusolverMpOrmtr() routines remain lower-only.Data Type of A and Q |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpSyevd_bufferSize#
cusolverStatus_t cusolverMpSyevd_bufferSize(
cusolverMpHandle_t handle,
char *jobz,
cublasFillMode_t uplo,
int64_t N,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
void *d_D,
void *d_Q,
int64_t IQ,
int64_t JQ,
cusolverMpMatrixDescriptor_t descQ,
cudaDataType_t computeType,
size_t *workspaceInBytesOnDevice,
size_t *workspaceInBytesOnHost)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
jobz |
Host |
In |
If |
uplo |
Host |
In |
Indicate that upper or lower triangular of sub(A) is used to compute eigen solutions. |
N |
Host |
In |
Number of rows and columns of sub(A). |
d_A |
Device |
In |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). |
JA |
Host |
In |
Column index of the first column of the sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_D |
Device |
In |
Pointer to a device buffer of length |
d_Q |
Device |
In |
Pointer into the local memory to an array of dimension |
IQ |
Host |
In |
Row index of the first row of the sub(Q). |
JQ |
Host |
In |
Column index of the first column of the sub(Q). |
descQ |
Host |
In |
Matrix descriptor associated to the global matrix Q. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
The size in bytes of the local device workspace needed by cusolverMpSyevd(). |
workspaceInBytesOnHost |
Host |
Out |
The size in bytes of the local host workspace needed by cusolverMpSyevd(). |
uplo selects which triangular part of the input sub(A) is read, and both CUBLAS_FILL_MODE_UPPER and CUBLAS_FILL_MODE_LOWER are supported. The reported workspace sizes do not depend on uplo.Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpSygst#
cusolverStatus_t cusolverMpSygst(
cusolverMpHandle_t handle,
cusolverEigType_t ibtype,
cublasFillMode_t uplo,
int64_t N,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
const void *d_B,
int64_t IB,
int64_t JB,
cusolverMpMatrixDescriptor_t descB,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *info)
ibtype = CUSOLVER_EIG_TYPE_1: the problem is sub(A)*x = lambda*sub(B)*x, and sub(A) is overwritten by inv(L)*sub(A)*inv(L^H) or inv(U^H)*sub(A)*inv(U).
ibtype = CUSOLVER_EIG_TYPE_2 or 3: the problem is sub(A)*sub(B)*x = lambda*x or sub(B)*sub(A)*x = lambda*x, and sub(A) is overwritten by L^H*sub(A)*L or U*sub(A)*U^H.
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
ibtype |
Host |
In |
Indicate the eigen problem type sub(A)*x=(lambda)*sub(B)*x, sub(A)*sub(B)x=(lambda)*x, or sub(B)*sub(A)*x=(lambda)*x. |
uplo |
Host |
In |
Indicate that lower |
N |
Host |
In |
Number of rows and columns of sub(A) and sub(B). |
d_A |
Device |
In/Out |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). |
JA |
Host |
In |
Column index of the first column of the sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_B |
Device |
In |
Pointer into the local memory to an array of dimension |
IB |
Host |
In |
Row index of the first row of the sub(B). |
JB |
Host |
In |
Column index of the first column of the sub(B). |
descB |
Host |
In |
Matrix descriptor associated to the global matrix B. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by cusolverMpSygst(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by cusolverMpSygst(). |
info |
Device |
Out |
Optional. |
Same square blocksize is used
(MB == NB)for the matrix A and B.The beginning row and column of A and B are aligned each other i.e.,
(IA == IB)and(JA == JB).
ibtype = CUSOLVER_EIG_TYPE_1, uplo = CUBLAS_FILL_MODE_LOWER, and full-matrix submatrix starts (IA == JA == IB == JB == 1).Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpSygst_bufferSize#
cusolverStatus_t cusolverMpSygst_bufferSize(
cusolverMpHandle_t handle,
cusolverEigType_t ibtype,
cublasFillMode_t uplo,
int64_t N,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
int64_t IB,
int64_t JB,
cusolverMpMatrixDescriptor_t descB,
cudaDataType_t computeType,
size_t *workspaceInBytesOnDevice,
size_t *workspaceInBytesOnHost)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
ibtype |
Host |
In |
Indicate the eigen problem type sub(A)*x=(lambda)*sub(B)*x, sub(A)*sub(B)x=(lambda)*x, or sub(B)*sub(A)*x=(lambda)*x. |
uplo |
Host |
In |
Indicate that lower |
N |
Host |
In |
Number of rows and columns of sub(A) and sub(B). |
IA |
Host |
In |
Row index of the first row of the sub(A). |
JA |
Host |
In |
Column index of the first column of the sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
IB |
Host |
In |
Row index of the first row of the sub(B). |
JB |
Host |
In |
Column index of the first column of the sub(B). |
descB |
Host |
In |
Matrix descriptor associated to the global matrix B. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
The size in bytes of the local device workspace needed by cusolverMpSygst(). |
workspaceInBytesOnHost |
Host |
Out |
The size in bytes of the local host workspace needed by cusolverMpSygst(). |
Same square blocksize is used
(MB == NB)for the matrix A and B.The beginning row and column of A and B are aligned each other i.e.,
(IA == IB)and(JA == JB).
ibtype = CUSOLVER_EIG_TYPE_1, uplo = CUBLAS_FILL_MODE_LOWER, and full-matrix submatrix starts (IA == JA == IB == JB == 1).Data Type of A |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpSygvd#
cusolverStatus_t cusolverMpSygvd(
cusolverMpHandle_t handle,
cusolverEigType_t ibtype,
cusolverEigMode_t jobz,
cublasFillMode_t uplo,
int64_t N,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
void *d_B,
int64_t IB,
int64_t JB,
cusolverMpMatrixDescriptor_t descB,
void *d_W,
void *d_Z,
int64_t IZ,
int64_t JZ,
cusolverMpMatrixDescriptor_t descZ,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *info)
ibtype = CUSOLVER_EIG_TYPE_1: the problem is sub(A)*x = lambda*sub(B)*x.
ibtype = CUSOLVER_EIG_TYPE_2: the problem is sub(A)*sub(B)*x = lambda*x.
ibtype = CUSOLVER_EIG_TYPE_3: the problem is sub(B)*sub(A)*x = lambda*x.
ncclMemAlloc registered via cusolverMpBufferRegister(), before calling this routine.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
ibtype |
Host |
In |
Indicate the eigen problem type sub(A)*x=(lambda)*sub(B)*x, sub(A)*sub(B)x=(lambda)*x, or sub(B)*sub(A)*x=(lambda)*x. |
jobz |
Host |
In |
Indicate whether the routine computes eigenvalues only |
uplo |
Host |
In |
Indicate that lower |
N |
Host |
In |
Number of rows and columns of sub(A) and sub(B). |
d_A |
Device |
In/Out |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the sub(A). |
JA |
Host |
In |
Column index of the first column of the sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_B |
Device |
In/Out |
Pointer into the local memory to an array of dimension |
IB |
Host |
In |
Row index of the first row of the sub(B). |
JB |
Host |
In |
Column index of the first column of the sub(B). |
descB |
Host |
In |
Matrix descriptor associated to the global matrix B. |
d_W |
Device |
Out |
Pointer into the memory to an array of global size |
d_Z |
Device |
Out |
Pointer into the local memory to an array of dimension |
IZ |
Host |
In |
Row index of the first row of the sub(Z). |
JZ |
Host |
In |
Column index of the first column of the sub(Z). |
descZ |
Host |
In |
Matrix descriptor associated to the global matrix Z. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpSygvd_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpSygvd_bufferSize(). |
info |
Device |
Out |
Optional. |
Same square blocksize is used
(MB == NB)for the matrix A, B, and Z.The beginning row and column of A, B and Z are aligned each other i.e.,
(IA == IB == IZ)and(JA == JB == JZ).
ibtype = CUSOLVER_EIG_TYPE_1) is supported. The current implementation also requires jobz = CUSOLVER_EIG_MODE_VECTOR, uplo = CUBLAS_FILL_MODE_LOWER, and full-matrix submatrix starts (IA == JA == IB == JB == IZ == JZ == 1).Data Type of A, B, and Z |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpSygvd_bufferSize#
cusolverStatus_t cusolverMpSygvd_bufferSize(
cusolverMpHandle_t handle,
cusolverEigType_t ibtype,
cusolverEigMode_t jobz,
cublasFillMode_t uplo,
int64_t N,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
int64_t IB,
int64_t JB,
cusolverMpMatrixDescriptor_t descB,
int64_t IZ,
int64_t JZ,
cusolverMpMatrixDescriptor_t descZ,
cudaDataType_t computeType,
size_t *workspaceInBytesOnDevice,
size_t *workspaceInBytesOnHost)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
ibtype |
Host |
In |
Indicate the eigen problem type sub(A)*x=(lambda)*sub(B)*x, sub(A)*sub(B)x=(lambda)*x, or sub(B)*sub(A)*x=(lambda)*x. |
jobz |
Host |
In |
Indicate whether the routine computes eigenvalues only |
uplo |
Host |
In |
Indicate that lower |
N |
Host |
In |
Number of rows and columns of sub(A) and sub(B). |
IA |
Host |
In |
Row index of the first row of the sub(A). |
JA |
Host |
In |
Column index of the first column of the sub(A). |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
IB |
Host |
In |
Row index of the first row of the sub(B). |
JB |
Host |
In |
Column index of the first column of the sub(B). |
descB |
Host |
In |
Matrix descriptor associated to the global matrix B. |
IZ |
Host |
In |
Row index of the first row of the sub(Z). |
JZ |
Host |
In |
Column index of the first column of the sub(Z). |
descZ |
Host |
In |
Matrix descriptor associated to the global matrix Z. |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
workspaceInBytesOnDevice |
Host |
Out |
The size in bytes of the local device workspace needed by cusolverMpSygvd(). |
workspaceInBytesOnHost |
Host |
Out |
The size in bytes of the local host workspace needed by cusolverMpSygvd(). |
Same square blocksize is used
(MB == NB)for the matrix A, B, and Z.The beginning row and column of A, B and Z are aligned each other i.e.,
(IA == IB == IZ)and(JA == JB == JZ).
ibtype = CUSOLVER_EIG_TYPE_1) is supported. The current implementation also requires jobz = CUSOLVER_EIG_MODE_VECTOR, uplo = CUBLAS_FILL_MODE_LOWER, and full-matrix submatrix starts (IA == JA == IB == JB == IZ == JZ == 1).Data Type of A, B, and Z |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpLaset#
cusolverStatus_t cusolverMpLaset(
cusolverMpHandle_t handle,
cublasFillMode_t uplo,
int64_t M,
int64_t N,
const void *alpha,
const void *beta,
void *d_A,
int64_t IA,
int64_t JA,
cusolverMpMatrixDescriptor_t descA,
int *d_info)
M-by-N distributed submatrix A(IA:IA+M-1, JA:JA+N-1) with alpha and the diagonal elements with beta. This is the distributed equivalent of LAPACK’s xLASET.uplo parameter controls which part of the submatrix is initialized:
CUBLAS_FILL_MODE_LOWER: only the lower triangular part (below and including the first subdiagonal) is set toalpha, and diagonal elements are set tobeta.
CUBLAS_FILL_MODE_UPPER: only the upper triangular part (above and including the first superdiagonal) is set toalpha, and diagonal elements are set tobeta.
CUBLAS_FILL_MODE_FULL: all off-diagonal elements are set toalpha, and diagonal elements are set tobeta.
alpha and beta scalars may reside in either host or device memory. The pointer type is detected automatically at runtime.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
uplo |
Host |
In |
Specifies the part of the submatrix to initialize: |
M |
Host |
In |
Number of rows of the submatrix. |
N |
Host |
In |
Number of columns of the submatrix. |
alpha |
Host/Device |
In |
Scalar value for off-diagonal elements. Must match the data type of matrix A. |
beta |
Host/Device |
In |
Scalar value for diagonal elements. Must match the data type of matrix A. |
d_A |
Device |
Out |
Pointer into the local memory to an array of dimension |
IA |
Host |
In |
Row index of the first row of the submatrix. |
JA |
Host |
In |
Column index of the first column of the submatrix. |
descA |
Host |
In |
Matrix descriptor associated to the global matrix A. |
d_info |
Device |
Out |
Optional. |
Data Type of A |
|---|
CUDA_R_32F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_64F |
cusolverMpNewtonSchulz#
cusolverStatus_t cusolverMpNewtonSchulz(
cusolverMpHandle_t handle,
cusolverMpNewtonSchulzDescriptor_t nsDesc,
int64_t M,
int64_t N,
void *d_X,
int64_t IX,
int64_t JX,
const cusolverMpMatrixDescriptor_t descX,
int64_t numberOfNewtonSchulzIterations,
const void *h_coeffs,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *d_info)
M >= N) distributed matrix X(IX:IX+M-1, JX:JX+N-1) in-place on supported Px1 process grids. The routine approximates the orthogonal polar factor U from the polar decomposition X = U * H, where U has orthonormal columns.M >= N), each iteration i applies a polynomial update using three user-supplied coefficients (alpha_i, beta_i, gamma_i):X := X / ||X||_F. This step can be disabled via the descriptor when the input is already normalized.h_coeffs must be provided as a host array of float triplets, with 3 * numberOfNewtonSchulzIterations elements stored as [alpha_0, beta_0, gamma_0, alpha_1, beta_1, gamma_1, ...]. See the Newton-Schulz sample for example coefficients optimized for quintic convergence in 5 iterations. The classical Newton-Schulz iteration can be recovered by setting (alpha, beta, gamma) = (1.5, -0.5, 0.0) for each iteration, though more iterations will be needed to converge.ncclMemAlloc registered via cusolverMpBufferRegister(), before calling this routine.
Only Px1 process grids (1D row distribution with
numColDevices = 1) are supported. 2D block-cyclic grids are not yet implemented.Only tall or square matrices (
M >= N) are supported. Wide rectangular matrices (M < N) are not yet supported.Only
IX = JX = 1is supported (no submatrix offsets).Only
CUDA_R_16BF(bfloat16) andCUDA_R_32F(float32) value types are supported.The only supported compute type is
CUDA_R_32F.
RSRC = CSRC = 0is required.
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
nsDesc |
Host |
In |
cusolverMpNewtonSchulzDescriptor_t descriptor (may be |
M |
Host |
In |
Number of rows of the submatrix X. |
N |
Host |
In |
Number of columns of the submatrix X. |
d_X |
Device |
In/Out |
Pointer into the local memory to an array of dimension |
IX |
Host |
In |
Row index of the first row of the submatrix. |
JX |
Host |
In |
Column index of the first column of the submatrix. |
descX |
Host |
In |
Matrix descriptor associated to the global matrix X. |
numberOfNewtonSchulzIterations |
Host |
In |
Number of Newton-Schulz iterations to perform. |
h_coeffs |
Host |
In |
Host array of triplets with |
computeType |
Host |
In |
Data type used for computations. See table below for supported combinations. |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpNewtonSchulz_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpNewtonSchulz_bufferSize(). |
d_info |
Device |
Out |
Optional. |
Data Type of X |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_16BF |
CUDA_R_32F |
CUDA_R_16BF |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
cusolverMpNewtonSchulz_bufferSize#
cusolverStatus_t cusolverMpNewtonSchulz_bufferSize(
cusolverMpHandle_t handle,
cusolverMpNewtonSchulzDescriptor_t nsDesc,
int64_t M,
int64_t N,
void *d_X,
int64_t IX,
int64_t JX,
const cusolverMpMatrixDescriptor_t descX,
int64_t numberOfNewtonSchulzIterations,
const void *h_coeffs,
cudaDataType_t computeType,
size_t *workspaceInBytesOnDevice,
size_t *workspaceInBytesOnHost)
Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
nsDesc |
Host |
In |
Newton-Schulz descriptor (may be |
M |
Host |
In |
Number of rows of the submatrix X. |
N |
Host |
In |
Number of columns of the submatrix X. |
d_X |
Device |
In |
Pointer into the local memory to an array of dimension |
IX |
Host |
In |
Row index of the first row of the submatrix. |
JX |
Host |
In |
Column index of the first column of the submatrix. |
descX |
Host |
In |
Matrix descriptor associated to the global matrix X. |
numberOfNewtonSchulzIterations |
Host |
In |
Number of Newton-Schulz iterations to perform. |
h_coeffs |
Host |
In |
Host array of triplets with |
computeType |
Host |
In |
Data type used for computations. |
workspaceInBytesOnDevice |
Host |
Out |
On output, contains the size in bytes of the local device workspace needed by cusolverMpNewtonSchulz(). |
workspaceInBytesOnHost |
Host |
Out |
On output, contains the size in bytes of the local host workspace needed by cusolverMpNewtonSchulz(). |
cusolverMpPolar#
cusolverStatus_t cusolverMpPolar(
cusolverMpHandle_t handle,
cusolverMpPolarDescriptor_t polarDesc,
cublasFillMode_t uplo,
int64_t m,
int64_t n,
void *d_A,
int64_t ia,
int64_t ja,
cusolverMpMatrixDescriptor_t descA,
void *d_H,
int64_t ih,
int64_t jh,
cusolverMpMatrixDescriptor_t descH,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *d_info)
A(ia:ia+m-1, ja:ja+n-1), where m >= n. In the default unperturbed case, the decomposition satisfies A = Up * H. On exit, d_A is overwritten with the m-by-n polar factor Up. For full-rank inputs, Up has orthonormal columns (Up^H * Up = I). For rank-deficient or numerically singular inputs without a positive perturbation, Up is not guaranteed to have orthonormal columns; the inaccurate directions of Up lie in the null space of A, so they are filtered out in products with A. In the unperturbed case, the computed factors still satisfy ||A - Up * H||_F <= O(n) * eps * ||A||_F and ||H^2 - A^H * A||_F <= O(n) * eps * ||A||_F^2, where H = Up^H * A and eps is the machine precision of the compute type. If H is requested, d_H receives the n-by-n Hermitian positive semidefinite factor H after the always-on projection 0.5 * (H + H^H).polarDesc is NULL or CUSOLVERMP_POLAR_DESCRIPTOR_ATTRIBUTE_REQUESTED_KSI is less than or equal to 0, no diagonal perturbation is applied. If REQUESTED_KSI is positive, that value is applied as a diagonal perturbation in original unscaled units. For perturbed square problems, the reported relation is A + ksi * I = Up * H. For perturbed rectangular problems, the reported relation is Up * H = Q_init * (R + ksi * I), where R is the n-by-n triangular factor and Q_init the orthonormal factor from the initial QR step. Equivalently, Up * H = A + ksi * Q_init, so the A-space residual ||A - Up * H||_F equals |ksi| * sqrt(n) in exact arithmetic.uplo may be CUBLAS_FILL_MODE_FULL for a general tall or square input. uplo may be CUBLAS_FILL_MODE_UPPER only when m == n and the input is already upper triangular; in that case only the upper triangle is read.MB == NB). The first tile of the submatrix must be owned by the same process row and process column as the first tile of the full matrix descriptor. In-bounds starts that are not tile-aligned or phase-preserving return CUSOLVER_STATUS_NOT_SUPPORTED.descH selects whether H is computed and must be rank-uniform. To skip H computation, pass NULL for descH on all ranks; d_H, ih, and jh are then ignored. To compute H, pass a valid descH on all ranks and pass a non-NULL d_H on every rank that owns local storage for the requested submatrix H(ih:ih+n-1, jh:jh+n-1).H is requested, descH must be non-NULL on all ranks and compatible with descA: same communicator, process grid, grid layout, tile sizes, and data type. The n-by-n submatrix H(ih:ih+n-1, jh:jh+n-1) must be in bounds, and local d_H storage must not alias d_A or d_work.polarDesc is non-NULL, CUSOLVERMP_POLAR_DESCRIPTOR_ATTRIBUTE_A_NORM_FROBENIUS receives the Frobenius norm of the logical, unperturbed input before overwrite. With uplo = CUBLAS_FILL_MODE_UPPER, the ignored strict lower triangle is treated as zero. CUSOLVERMP_POLAR_DESCRIPTOR_ATTRIBUTE_RCOND_ESTIMATE receives an estimate of 1 / (||R||_1 ||R^{-1}||_1) for the initial unperturbed triangular factor R used to initialize the iteration. For a FULL input, R is the upper-triangular factor from QR of the logical input; for uplo = CUBLAS_FILL_MODE_UPPER, R is the input itself. Descriptor output attributes are reset to NaN at routine entry and populated only if the corresponding quantity is computed.uplo, m, n, REQUESTED_KSI, and whether H is requested, must be identical on all ranks.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
polarDesc |
Host |
In |
cusolverMpPolarDescriptor_t descriptor for perturbation control and scalar diagnostics. May be |
uplo |
Host |
In |
Specifies the form of the input: |
m |
Host |
In |
Number of rows of the submatrix A. |
n |
Host |
In |
Number of columns of the submatrix A and order of H. |
d_A |
Device |
In/Out |
Pointer into the local memory for matrix A. On entry, contains the input submatrix. On exit, overwritten with the polar factor Up. |
ia |
Host |
In |
Row index of the first row of sub(A). Must satisfy the submatrix offset alignment rule above. |
ja |
Host |
In |
Column index of the first column of sub(A). Must satisfy the submatrix offset alignment rule above. |
descA |
Host |
In |
Matrix descriptor associated with A. |
d_H |
Device |
Out |
Pointer into the local memory for H when |
ih |
Host |
In |
Row index of the first row of sub(H). Used only when H is requested. |
jh |
Host |
In |
Column index of the first column of sub(H). Used only when H is requested. |
descH |
Host |
In |
Matrix descriptor associated with H and the rank-uniform H-output selector. Pass |
computeType |
Host |
In |
Data type used for computation. Must match the data type of |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpPolar_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpPolar_bufferSize(). |
d_info |
Device |
Out |
Optional. |
Data Type of A and H |
computeType |
Output Data Type |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_C_64F |
cusolverMpPolar_bufferSize#
cusolverStatus_t cusolverMpPolar_bufferSize(
cusolverMpHandle_t handle,
cusolverMpPolarDescriptor_t polarDesc,
cublasFillMode_t uplo,
int64_t m,
int64_t n,
const void *d_A,
int64_t ia,
int64_t ja,
cusolverMpMatrixDescriptor_t descA,
const void *d_H,
int64_t ih,
int64_t jh,
cusolverMpMatrixDescriptor_t descH,
cudaDataType_t computeType,
size_t *workspaceInBytesOnDevice,
size_t *workspaceInBytesOnHost)
H computation from the workspace query, pass NULL for descH. To include H computation, pass a valid descH. The query does not inspect d_H.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
polarDesc |
Host |
In |
Polar descriptor. May be |
uplo |
Host |
In |
Specifies the form of the input: |
m |
Host |
In |
Number of rows of the submatrix A. |
n |
Host |
In |
Number of columns of the submatrix A and order of H. |
d_A |
Device |
In |
Pointer into the local memory for matrix A. |
ia |
Host |
In |
Row index of the first row of sub(A). Must satisfy the submatrix offset alignment rule for cusolverMpPolar(). |
ja |
Host |
In |
Column index of the first column of sub(A). Must satisfy the submatrix offset alignment rule for cusolverMpPolar(). |
descA |
Host |
In |
Matrix descriptor associated with A. |
d_H |
Device |
In |
Not inspected by the workspace query. |
ih |
Host |
In |
Row index of the first row of sub(H). Used only when H computation is included. |
jh |
Host |
In |
Column index of the first column of sub(H). Used only when H computation is included. |
descH |
Host |
In |
Matrix descriptor associated with H and the H-workspace selector. Pass |
computeType |
Host |
In |
Data type used for computation. Must match the data type of |
workspaceInBytesOnDevice |
Host |
Out |
On output, contains the size in bytes of the local device workspace needed by cusolverMpPolar(). |
workspaceInBytesOnHost |
Host |
Out |
On output, contains the size in bytes of the local host workspace needed by cusolverMpPolar(). |
cusolverMpGesvd#
cusolverStatus_t cusolverMpGesvd(
cusolverMpHandle_t handle,
cusolverMpGesvdDescriptor_t gesvdDesc,
cusolverEigMode_t jobu,
cusolverEigMode_t jobvt,
int64_t m,
int64_t n,
void *d_A,
int64_t ia,
int64_t ja,
cusolverMpMatrixDescriptor_t descA,
void *d_S,
void *d_U,
int64_t iu,
int64_t ju,
cusolverMpMatrixDescriptor_t descU,
void *d_VT,
int64_t ivt,
int64_t jvt,
cusolverMpMatrixDescriptor_t descVT,
cudaDataType_t computeType,
void *d_work,
size_t workspaceInBytesOnDevice,
void *h_work,
size_t workspaceInBytesOnHost,
int *d_info)
A = U * Sigma * V^H of the distributed submatrix A(ia:ia+m-1, ja:ja+n-1). The routine supports both tall (m >= n) and wide (m < n) inputs.k = min(m,n). The singular values are written in descending order to the replicated device vector d_S of length k on every rank. By default, requested singular-vector outputs use THIN shape: if jobu is CUSOLVER_EIG_MODE_VECTOR, d_U receives the m-by-k thin left singular vectors; if jobvt is CUSOLVER_EIG_MODE_VECTOR, d_VT receives the k-by-n thin right singular vectors as V^H. Set CUSOLVERMP_GESVD_DESCRIPTOR_ATTRIBUTE_SHAPE to CUSOLVERMP_GESVD_OUTPUT_SHAPE_FULL to request full factors: U is m by m and V^H is n by n. If either mode is CUSOLVER_EIG_MODE_NOVECTOR, the corresponding pointer, descriptor, and offsets are ignored.A(ia:ia+m-1, ja:ja+n-1) must fit inside descA. Its offsets must be at least 1 and must start on a source-owned tile boundary in descA; bounds failures return CUSOLVER_STATUS_INVALID_VALUE and tile/phase failures return CUSOLVER_STATUS_NOT_SUPPORTED. When a factor is requested, its submatrix offsets must be at least 1 and must start on a source-owned tile boundary in the parent descriptor. The requested output window must also fit inside the parent descriptor. THIN requires descU.M >= iu - 1 + m and descU.N >= ju - 1 + k for U, and descVT.M >= ivt - 1 + k and descVT.N >= jvt - 1 + n for V^H. FULL requires descU.M >= iu - 1 + m and descU.N >= ju - 1 + m for U, and descVT.M >= ivt - 1 + n and descVT.N >= jvt - 1 + n for V^H. Larger parent descriptors are allowed as storage capacity; the descriptor shape attribute, not descriptor capacity alone, selects FULL.descA and the descriptors for requested U and V^H outputs must use square tiles (MB == NB) and RSRC = CSRC = 0. Requested output descriptors must also be structurally compatible with descA: same communicator, grid layout, process grid dimensions, tile sizes, and source ranks. Unsupported structural mismatches return CUSOLVER_STATUS_NOT_SUPPORTED. computeType must be one of the supported GESVD compute types and must match the data type of descA and every requested output descriptor; datatype mismatches return CUSOLVER_STATUS_INVALID_VALUE. Local-memory overlap between d_S and requested A, U, or V^H storage, or between the requested A, U, and V^H windows, is not supported and returns CUSOLVER_STATUS_NOT_SUPPORTED. The input storage d_A may be overwritten.d_A and replicated d_S must be non-NULL on every rank. When jobu or jobvt requests vectors, d_U or d_VT must also be non-NULL on every rank, including ranks that own no local elements of the corresponding distributed window; empty-owner ranks may pass a minimal dummy allocation. In CUSOLVER_EIG_MODE_NOVECTOR mode, the corresponding pointer, descriptor, and offsets are ignored.gesvdDesc is non-NULL, descriptor attributes control optional behavior, and diagnostic output attributes are reset at routine entry and populated as the algorithm progresses. Query output attributes after the call returns using cusolverMpGesvdDescriptorGetAttribute(). Passing NULL uses default attributes and discards diagnostics.jobu, jobvt, m, n, computeType, ia, ja, the requested-output offsets iu, ju, ivt, and jvt when the corresponding vector factor is requested, and GESVD descriptor input attributes (COMPUTE_RESIDUAL, SHAPE, and future attributes that affect behavior), must be identical on all ranks.ncclMemAlloc registered via cusolverMpBufferRegister(), before calling this routine.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
gesvdDesc |
Host |
In |
cusolverMpGesvdDescriptor_t descriptor for attributes and scalar diagnostics. May be |
jobu |
Host |
In |
Specifies whether to compute left singular vectors: |
jobvt |
Host |
In |
Specifies whether to compute right singular vectors as |
m |
Host |
In |
Number of rows of the submatrix A. |
n |
Host |
In |
Number of columns of the submatrix A. |
d_A |
Device |
In/Out |
Pointer into the local memory for matrix A. Required on every rank for non-empty problems; empty-owner ranks may pass a minimal dummy allocation. The input contents may be overwritten. |
ia |
Host |
In |
Row index of the first row of sub(A). Must be at least 1, fit the requested A window inside |
ja |
Host |
In |
Column index of the first column of sub(A). Must be at least 1, fit the requested A window inside |
descA |
Host |
In |
Matrix descriptor associated with A. |
d_S |
Device |
Out |
Replicated device vector of length |
d_U |
Device |
Out |
Pointer into the local memory for U. When |
iu |
Host |
In |
Row index of the first row of sub(U). Used only when |
ju |
Host |
In |
Column index of the first column of sub(U). Used only when |
descU |
Host |
In |
Matrix descriptor associated with U. Used only when |
d_VT |
Device |
Out |
Pointer into the local memory for V^H. When |
ivt |
Host |
In |
Row index of the first row of sub(VT). Used only when |
jvt |
Host |
In |
Column index of the first column of sub(VT). Used only when |
descVT |
Host |
In |
Matrix descriptor associated with V^H. Used only when |
computeType |
Host |
In |
Data type used for computation. Must match the data type of |
d_work |
Device |
Out |
Device workspace of size |
workspaceInBytesOnDevice |
Host |
In |
The size in bytes of the local device workspace needed by the routine as provided by cusolverMpGesvd_bufferSize(). |
h_work |
Host |
Out |
Host workspace of size |
workspaceInBytesOnHost |
Host |
In |
The size in bytes of the local host workspace needed by the routine as provided by cusolverMpGesvd_bufferSize(). |
d_info |
Device |
Out |
Optional. |
Data Type of A, U, and VT |
computeType |
Output Data Type of S |
|---|---|---|
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_32F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_R_64F |
CUDA_C_32F |
CUDA_C_32F |
CUDA_R_32F |
CUDA_C_64F |
CUDA_C_64F |
CUDA_R_64F |
cusolverMpGesvd_bufferSize#
cusolverStatus_t cusolverMpGesvd_bufferSize(
cusolverMpHandle_t handle,
cusolverMpGesvdDescriptor_t gesvdDesc,
cusolverEigMode_t jobu,
cusolverEigMode_t jobvt,
int64_t m,
int64_t n,
const void *d_A,
int64_t ia,
int64_t ja,
cusolverMpMatrixDescriptor_t descA,
const void *d_S,
const void *d_U,
int64_t iu,
int64_t ju,
cusolverMpMatrixDescriptor_t descU,
const void *d_VT,
int64_t ivt,
int64_t jvt,
cusolverMpMatrixDescriptor_t descVT,
cudaDataType_t computeType,
size_t *workspaceInBytesOnDevice,
size_t *workspaceInBytesOnHost)
U and V^H are requested, and GESVD descriptor attributes; query workspace with the same control-flow inputs and descriptor attributes used for execution. The device pointer values d_A, d_S, d_U, and d_VT are not dereferenced by the workspace query.Parameter |
Memory |
In/Out |
Description |
|---|---|---|---|
handle |
Host |
In |
cuSOLVERMp library handle. |
gesvdDesc |
Host |
In |
Singular value decomposition descriptor. May be |
jobu |
Host |
In |
Specifies whether to compute left singular vectors: |
jobvt |
Host |
In |
Specifies whether to compute right singular vectors as |
m |
Host |
In |
Number of rows of the submatrix A. |
n |
Host |
In |
Number of columns of the submatrix A. |
d_A |
Device |
In |
Pointer argument corresponding to matrix A. Not dereferenced by this workspace query. |
ia |
Host |
In |
Row index of the first row of sub(A). |
ja |
Host |
In |
Column index of the first column of sub(A). |
descA |
Host |
In |
Matrix descriptor associated with A. |
d_S |
Device |
In |
Pointer argument corresponding to the replicated singular-value vector. Not dereferenced by this workspace query. |
d_U |
Device |
In |
Pointer argument corresponding to U. Not dereferenced by this workspace query. |
iu |
Host |
In |
Row index of the first row of sub(U). Used only when |
ju |
Host |
In |
Column index of the first column of sub(U). Used only when |
descU |
Host |
In |
Matrix descriptor associated with U. Used only when |
d_VT |
Device |
In |
Pointer argument corresponding to V^H. Not dereferenced by this workspace query. |
ivt |
Host |
In |
Row index of the first row of sub(VT). Used only when |
jvt |
Host |
In |
Column index of the first column of sub(VT). Used only when |
descVT |
Host |
In |
Matrix descriptor associated with V^H. Used only when |
computeType |
Host |
In |
Data type used for computation. Must match the data type of |
workspaceInBytesOnDevice |
Host |
Out |
On output, contains the size in bytes of the local device workspace needed by cusolverMpGesvd(). |
workspaceInBytesOnHost |
Host |
Out |
On output, contains the size in bytes of the local host workspace needed by cusolverMpGesvd(). |