Data
sgemm-gpu-kernel

sgemm-gpu-kernel

active ARFF Publicly available Visibility: public Uploaded 28-05-2021 by Meilina Reksoprodjo
0 likes downloaded by 0 people , 0 total downloads 0 issues 0 downvotes
Issue #Downvotes for this reason By


Loading wiki
Help us complete this description Edit
Author: Enrique G. Paredes, Rafael Ballester-Ripoll Source: [UCI](https://archive.ics.uci.edu/ml/datasets/SGEMM+GPU+kernel+performance) - 2018 Please cite: [Paper](https://arxiv.org/abs/1712.00233) SGEMM GPU kernel performance dataset This data set measures the running time of a matrix-matrix product A x B = C, where all matrices have size 2048 x 2048, using a parameterizable SGEMM GPU kernel with 241600 possible parameter combinations. For each tested combination, 4 runs were performed and their results are reported as the 4 last columns. All times are measured in milliseconds*. There are 14 parameter, the first 10 are ordinal and can only take up to 4 different powers of two values, and the 4 last variables are binary. Out of 1327104 total parameter combinations, only 241600 are feasible (due to various kernel constraints). This data set contains the results for all these feasible combinations. The experiment was run on a desktop workstation running Ubuntu 16.04 Linux with an Intel Core i5 (3.5GHz), 16GB RAM, and a NVidia Geforce GTX 680 4GB GF580 GTX-1.5GB GPU. We use the 'gemm_fast' kernel from the automatic OpenCL kernel tuning library 'CLTune' ([Web Link]). * Note: for this kind of data sets it is usually better to work with the logarithm of the running times (see e.g. Falch and Elster, 'Machine learning-based auto-tuning for enhanced performance portability of OpenCL applications', 2015). ### Attribute information Independent variables: 1-2. MWG, NWG: per-matrix 2D tiling at workgroup level: {16, 32, 64, 128} (integer) 3. KWG: inner dimension of 2D tiling at workgroup level: {16, 32} (integer) 4-5. MDIMC, NDIMC: local workgroup size: {8, 16, 32} (integer) 6-7. MDIMA, NDIMB: local memory shape: {8, 16, 32} (integer) 8. KWI: kernel loop unrolling factor: {2, 8} (integer) 9-10. VWM, VWN: per-matrix vector widths for loading and storing: {1, 2, 4, 8} (integer) 11-12. STRM, STRN: enable stride for accessing off-chip memory within a single thread: {0, 1} (categorical) 13-14. SA, SB: per-matrix manual caching of the 2D workgroup tile: {0, 1} (categorical) Output: 15-18. Run1, Run2, Run3, Run4: performance times in milliseconds for 4 independent runs using the same parameters. They range between 13.25 and 3397.08.

18 features

MWGnumeric4 unique values
0 missing
NWGnumeric4 unique values
0 missing
KWGnumeric2 unique values
0 missing
MDIMCnumeric3 unique values
0 missing
NDIMCnumeric3 unique values
0 missing
MDIMAnumeric3 unique values
0 missing
NDIMBnumeric3 unique values
0 missing
KWInumeric2 unique values
0 missing
VWMnumeric4 unique values
0 missing
VWNnumeric4 unique values
0 missing
STRMnumeric2 unique values
0 missing
STRNnumeric2 unique values
0 missing
SAnumeric2 unique values
0 missing
SBnumeric2 unique values
0 missing
Run1 (ms)numeric58161 unique values
0 missing
Run2 (ms)numeric58269 unique values
0 missing
Run3 (ms)numeric58264 unique values
0 missing
Run4 (ms)numeric58154 unique values
0 missing

19 properties

241600
Number of instances (rows) of the dataset.
18
Number of attributes (columns) of the dataset.
Number of distinct values of the target attribute (if it is nominal).
0
Number of missing values in the dataset.
0
Number of instances with at least one value missing.
18
Number of numeric attributes.
0
Number of nominal attributes.
Average class difference between consecutive instances.
0
Percentage of missing values.
0
Number of attributes divided by the number of instances.
100
Percentage of numeric attributes.
Percentage of instances belonging to the most frequent class.
0
Percentage of nominal attributes.
Number of instances belonging to the most frequent class.
Percentage of instances belonging to the least frequent class.
Number of instances belonging to the least frequent class.
0
Number of binary attributes.
0
Percentage of binary attributes.
0
Percentage of instances having missing values.

0 tasks

Define a new task