High performance: close to roofline fp16 TensorCore (NVIDIA GPU) / MatrixCore (AMD GPU) performance on major models, including ResNet, MaskRCNN, BERT, VisionTransformer, Stable Diffusion, etc. Unified ...
These implementations are for demonstration purposes. They are less efficient than the implementations in the Python standard library.
Some results have been hidden because they may be inaccessible to you
Show inaccessible results