pengxu
ba9569e382
Loongarch64: fixed dot_lasx
2025-04-30 16:41:48 +08:00
pengxu
dc5fa29851
Loongarch64: fixed cscal_lasx
2025-04-30 16:41:39 +08:00
pengxu
a98dd6d911
Loongarch64: fixed copy_lasx
2025-04-30 16:41:28 +08:00
pengxu
d49319c2d2
Loongarch64: fixed cnrm2_lasx
2025-04-30 16:41:18 +08:00
pengxu
74c97ef814
Loongarch64: fixed cdot_lasx
2025-04-30 16:41:05 +08:00
pengxu
be525521ad
Loongarch64: fixed asum_lasx
2025-04-30 16:40:55 +08:00
pengxu
0cd5ca5527
Loongarch64: fixed amax_lasx
2025-04-30 16:40:44 +08:00
guoyuanplct
11ffc8680e
Format the code
2025-04-25 00:27:27 +08:00
guoyuanplct
7616c42095
Optimized RVV_ZVL256B Implementation of zgemv_n
...
The implementation of zgemv_n using RVV_ZVL256B has been optimized.
Compared to the previous implementation, it has achieved a 1.5x
performance improvement.
2025-04-25 00:05:15 +08:00
abhishek-fujitsu
9c02cdb073
optimise dot using thread throttling for NEOVERSE V1
2025-04-23 22:35:05 +05:30
Martin Kroeker
d0e8fd6d40
Merge pull request #5239 from annop-w/gemv_n_sve
...
Use SVE kernel for S/DGEMVN for SVE machines
2025-04-22 10:19:49 -07:00
Iha, Taisei
08b5c18d70
fixed a potential out-of-bounds on gemv.
2025-04-22 19:56:44 +09:00
Annop Wongwathanarat
e11744a411
Use SVE kernel for S/DGEMVN for SVE machines
2025-04-22 09:40:13 +00:00
Martin Kroeker
db0abfa907
Merge pull request #5238 from martin-frbg/revert5125
...
remove non-vectorized SGEMV transpose reduce path for POWER8, restoring optimizations frpm PR4880
2025-04-22 02:12:19 -07:00
Martin Kroeker
7389b6c483
Merge pull request #5237 from martin-frbg/revert5219
...
Fix and reinstate the Cooper Lake/Sapphire Rapids microkernel for non-transpose SBGEMV
2025-04-21 23:36:23 -07:00
Martin Kroeker
4ec62d7f73
remove non-vectorized code path for power8, restoring PR4880
2025-04-21 23:14:10 +02:00
Martin Kroeker
1df8738f27
Merge pull request #5235 from quickwritereader/issue_unaligned_ppc64le
...
Explicit unaligned vector load/stores in PPC64LE GEMV kernels
2025-04-21 14:03:56 -07:00
Martin Kroeker
99d9f1ff38
Fix conditional
2025-04-21 22:55:45 +02:00
Martin Kroeker
96d80801bc
Reinstate the CooperLake microkernel
2025-04-21 22:53:26 +02:00
Martin Kroeker
2e4309315c
Merge pull request #5219 from martin-frbg/sbgemvn_cooper
...
Temporarily disable the Cooper Lake/Sapphire Rapids microkernel for non-transpose SBGEMV
2025-04-20 07:29:20 -07:00
Ubuntu
0cc2485594
Explicit unaligned vector load/stores in PPC64LE GEMV kernels
2025-04-20 08:00:29 +00:00
Martin Kroeker
dd38b4e811
Merge pull request #5225 from annop-w/gemv_n
...
Improve performance for SGEMVN on NEONVERSEN1
2025-04-17 01:54:10 -07:00
Martin Kroeker
0241d516f6
Merge pull request #5220 from iha-taisei/sdgemv_n_unroll
...
Further performance improvements to non-transposed [SD]GEMV kernels for A64FX and Neoverse V1.
2025-04-16 12:55:55 -07:00
Annop Wongwathanarat
d535728803
Improve performance for SGEMVN on NEONVERSEN1
2025-04-16 09:54:30 +00:00
Usui, Tetsuzo
d711906e3e
Add symv kernels for arm64
2025-04-11 20:39:52 +09:00
Iha, Taisei
f1e628b889
Further performance improvements to [SD]GEMV.
2025-04-11 20:00:33 +09:00
Martin Kroeker
211dfd0754
disable the CooperLake microkernel as it produces wrong results
2025-04-10 22:21:57 +02:00
Martin Kroeker
b30dc9701f
Merge pull request #5215 from annop-w/gemv_t
...
Use SVE kernel for S/DGEMVT for SVE machines
2025-04-10 13:06:07 -07:00
Martin Kroeker
2893d0add4
Merge pull request #5211 from guoyuanplct/develop
...
Optimizing the Implementation of GEMV on the RISC-V V Extension
2025-04-10 09:43:03 -07:00
Annop Wongwathanarat
ec146157d3
Use SVE kernel for S/DGEMVT for SVE machines
2025-04-09 20:38:14 +00:00
Martin Kroeker
70865a894e
Merge pull request #5180 from ywwry66/openmp_use_cmake
...
CMake: Pass `OpenMP` compiler and linker flags through CMake targets
2025-04-08 13:16:07 -07:00
lglglglgy
1ff303f36e
Optimizing the Implementation of GEMV on the RISC-V V Extension
...
Specialized some scenarios, performed loop unrolling, and reduced the
number of multiplications.
2025-04-08 21:18:00 +08:00
ColumbusAI
7bf848454d
Update zsum.c -- fixed spelling error to successfully compile
...
spelling error where zsum_kernel is used and it should be zasum_kernel. Will not compile without fix.
2025-04-05 09:57:53 -07:00
Vaisakh K V
04915be829
Add vector registers to clobber list to prevent compiler optimization.
...
SME based SGEMMDIRECT kernel uses the vector registers (z) and adding
clobber list informs compiler not to optimize these registers.
2025-04-03 12:18:43 +05:30
Egbert Eich
ea6515c4b3
On zarch don't produce objects from assembler with a writable stack section
...
On z-series, the current version of the GNU toolchain produces warnings
such as:
```
/usr/lib64/gcc/[...]/s390x-suse-linux/bin/ld: warning: ztrmm_kernel_RC_Z14.o: missing .note.GNU-stack section implies
executable stack
/usr/lib64/[...]/s390x-suse-linux/bin/ld: NOTE: This behaviour is deprecated and will be removed in a future version of the linker
```
To prevent this message and make sure we are future proof, add
```
.section .note.GNU-stack,"",@progbits
```
Also add the `.size` bit to give the asm defined functions a proper size
in the symbol table.
Signed-off-by: Egbert Eich <eich@suse.com>
2025-03-28 18:47:48 +01:00
Ruiyang Wu
02fd1df10b
CMake: Pass `OpenMP` compiler and linker flags through CMake targets
...
Using `OpenMP::OpenMP_LANG` targets for CMake is less error-prone than
passing the compiler and linker flags manually. Furthermore, it allows
the user to customize those flags by setting `OpenMP_LANG_FLAGS`,
`OpenMP_LANG_LIB_NAMES`, and `OpenMP_omp_LIBRARY`.
2025-03-26 23:09:54 -04:00
Ye Tao
f27ba5efd1
fix bugs in aarch64 sbgemv_n kernel
2025-03-14 17:55:40 +00:00
Annop Wongwathanarat
edef2e4441
Fix bug in ARM64 sbgemv_t
2025-03-13 20:55:31 +00:00
Martin Kroeker
b55ca71d5b
Merge pull request #5182 from annop-w/sgemm_ncopy
...
Optimize aarch64 sgemm_ncopy
2025-03-13 16:04:39 +01:00
Martin Kroeker
2f778554b8
Merge pull request #5181 from taoye9/change_sbgemn_cast_bf16
...
replace customize bf16_to_fp32 with arm neon vcvtah_f32_bf16
2025-03-13 13:50:26 +01:00
Annop Wongwathanarat
9807f56580
Optimize aarch64 sgemm_ncopy
2025-03-13 10:17:43 +00:00
Martin Kroeker
a3e7b16072
Merge pull request #5157 from manaalmj/feature
...
Optimize gemv_n_sve kernel
2025-03-12 21:08:23 +01:00
Ye Tao
4c00099ed6
replace customize bf16_to_fp32 with arm neon vcvtah_f32_bf16
2025-03-12 16:20:15 +00:00
Annop Wongwathanarat
a085b6c9ec
Fix aarch64 sbgemv_t compilation error for GCC < 13
2025-03-12 14:52:42 +00:00
manjam01
5c4e38ab17
Optimize gemv_n_sve kernel
2025-03-10 16:39:20 +00:00
Martin Kroeker
1d5ed5c46b
Merge pull request #5168 from taoye9/add_sbgemvn_on_neonversen2
...
Add dispatch of SBGEMVNKERNEL for NEOVERSEN2 and NEOVERSEV2
2025-03-04 16:39:22 +01:00
Ye Tao
6b8b35cdf2
fix minior issues of redeclaration of float x0,x1 in sbgemv_n_neon.c
2025-03-03 11:55:27 +00:00
Ye Tao
38ee7c9301
Add dispatch of SBGEMVNKERNEL for NEOVERSEN2 and NEOVERSEV2
2025-03-03 11:32:05 +00:00
Martin Kroeker
2b941c44b5
Merge branch 'develop' into sbgemv_n_neon
2025-03-02 22:39:32 +01:00
Ye Tao
35bdbca153
Add sbgemv_n_neon kernel for arm64.
2025-02-28 14:37:06 +00:00
Annop Wongwathanarat
edaf51dd99
Add sbgemv_t_bfdot kernel for ARM64
...
This improves performance for sbgemv_t by up to 100x on NEOVERSEV1.
The geometric mean speedup is ~61x for M=N=[2,512].
2025-02-28 12:31:50 +00:00
Martin Kroeker
77fba0f400
Fix "dummy2" flag handling
2025-02-22 20:09:21 +01:00
Martin Kroeker
eb84aac7ad
Merge pull request #5084 from quic/topic/sgemm_direct_sme1
...
Support for SGEMM_DIRECT Kernel based on SME1
2025-02-19 10:56:49 +01:00
Martin Kroeker
b9ae246f20
define USE_TRMM for RISCV64 targets as well
2025-02-16 23:18:04 +01:00
Vaisakh K V
f66ca05b31
Merge branch 'develop' into topic/sgemm_direct_sme1
2025-02-13 14:54:37 +05:30
Vaisakh K V
d23eb3b93e
Support for SME1 based sgemm_direct kernel for cblas_sgemm level 3 API
...
* Added ARMV9SME target
* Added SGEMM_DIRECT kernel based on SME1
2025-02-13 14:51:21 +05:30
Martin Kroeker
8d487ef6eb
Merge pull request #5124 from XiWeiGu/LoongArch64-LA264-lapack-fixed
...
LoongArch64: Fixed lapack test for LA264
2025-02-12 14:58:30 +01:00
Martin Kroeker
81eed868b6
Restore the non-vectorized code from before PR4880 for POWER8
2025-02-12 09:07:20 +01:00
Martin Kroeker
98b5ef929c
Restore the non-vectorized code from before PR4880 for POWER8
2025-02-12 09:04:22 +01:00
gxw
2c4a5cc6e6
LoongArch64: Fixed snrm2_lsx.S and cnrm2_lsx.S
...
When the data type is single-precision real or single-precision complex,
converting it to double precision does not prevent overflow (as exposed in LAPACK tests).
The only solution is to follow C's approach: find the maximum value in the
array and divide each element by that maximum to avoid this issue
2025-02-12 15:48:01 +08:00
gxw
9e75d6b3d1
LoongArch64: Fixed swap_lsx.S
...
Fixed the error when the stride is zero
2025-02-12 14:57:35 +08:00
gxw
e8c740368c
LoongArch64: Fixed rot_lsx.S ane crot_lsx.S
...
Do not check whether the input parameters c and s are zero,
as this may cause errors with special values (same as scal).
Although OpenBLAS's own test suite doesn't catch this, it will
cause LAPACK test cases to fail.
2025-02-12 14:52:49 +08:00
Hao Chen
c2212d0abd
LoongArch64: Fixed copy_lsx.S
...
Fixed incorrect store operation
Signed-off-by: gxw <guxiwei-hf@loongson.cn>
2025-02-12 14:52:20 +08:00
Hao Chen
7f1ebc7ae6
LoongArch64: Fixed iamax_lsx.S
...
Fixed index retrieval issue when there are
identical maximum absolute values
Signed-off-by: Hao Chen <chenhao@loongson.cn>
Signed-off-by: gxw <guxiwei-hf@loongson.cn>
2025-02-12 14:44:44 +08:00
Hao Chen
31d326f895
LoongArch64: Fixed dot_lsx.S
...
Fixed incorrect register usage in instructions
Signed-off-by: gxw <guxiwei-hf@loongson.cn>
2025-02-12 14:44:11 +08:00
Hao Chen
5d6356bc16
LoongArch64: Fixed amax_lsx.S
...
Fixed register zeroing operation
Signed-off-by: Hao Chen <chenhao@loongson.cn>
Signed-off-by: gxw <guxiwei-hf@loongson.cn>
2025-02-12 14:39:29 +08:00
Ye Tao
c748e6a338
optimized sbgemm kernel for neoverse-v1 (sve-256)
...
Signed-off-by: Ye Tao <ye.tao@arm.com>
2025-02-05 10:06:37 +00:00
Aditya Tewari
4379a6fbe3
* checkpoint sbgemm for SVE-256
2025-02-03 12:49:49 +00:00
Martin Kroeker
d7036cfd74
Remove trailing blanks that break the cmake parser
2025-01-27 09:32:17 +01:00
Martin Kroeker
6e393a5599
Merge branch 'develop' into gemv_t
2025-01-25 12:54:04 +01:00
Martin Kroeker
876ba58e28
Merge pull request #5091 from goplanid/develop
...
Small gemm kernel improvements for AArch64
2025-01-24 10:59:16 +01:00
Martin Kroeker
180ba5e7d0
Merge pull request #5069 from tingboliao/dev_rotm_20250107
...
Further rearranged the rotm kernel for the different architectures.
2025-01-23 10:16:43 +01:00
Deeksha Goplani
d1bfa979f7
small gemm kernel packing modifications
2025-01-23 09:41:45 +05:30
Martin Kroeker
1a6a9fb22f
add another generator line for rotm
2025-01-23 00:17:04 +01:00
Martin Kroeker
4924319c50
fix position of srotm, qrotm
2025-01-22 16:07:35 +01:00
Martin Kroeker
b58cba9eb6
fix qrotm build rules
2025-01-22 15:51:49 +01:00
gxw
2da86b80c9
LoongArch64: Fixed scalar version of cscal and zscal
2025-01-22 14:32:20 +08:00
gxw
5392f6df69
LoongArch64: Fixed LASX version of cscal and zscal
2025-01-22 14:15:27 +08:00
tingbo.liao
3c8df6358f
Further rearranged the rotm kernel for the different architectures.
...
Signed-off-by: tingbo.liao <tingbo.liao@starfivetech.com>
2025-01-22 11:41:12 +08:00
Annop Wongwathanarat
c0318cea6e
Simplify gemv_t_sve_v1x3 kernel
2025-01-21 13:40:17 +00:00
gxw
b2117bb2ca
LoongArch64: Fixed LSX version of cscal and zscal
2025-01-21 14:49:44 +08:00
gxw
e114880dc4
kernel/generic: Fixed cscal and zscal
2025-01-21 11:44:22 +08:00
Martin Kroeker
87083fdbf6
[WIP] Work around assembler limitations in current LLVM for Windows on Arm ( #5076 )
...
* Protect align directives in assembly files that are currently problematic with LLVM on WoA
* use the armv8 zdot on WoA to work around other LLVM issues
2025-01-18 16:45:56 +01:00
tingbo.liao
ef7f54b357
Optimized the gemm_tcopy_8_rvv to be compatible with the vlens 128 and 256.
...
Signed-off-by: tingbo.liao <tingbo.liao@starfivetech.com>
2025-01-15 11:31:28 +08:00
gxw
e0a8216554
LoongArch64: Update dsymv LSX version
2025-01-14 19:45:42 +08:00
gxw
a9070ba3f9
LoongArch64: Update ssymv LSX version
2025-01-14 09:06:59 +00:00
Xi Ruoyao
af10c132b8
LoongArch64: Fix dsymv and ssymv LASX version
...
"fmov.d $f2, $f4" leaves all the bits higher than the 63-th bit
unpredictable but it's obvious that the following code uses the value of
those high bits. We actually want to replicate the lower 64 bits here,
so we should use xvreplve0.d instead.
LA464 (Loongson 3[A-Z]-5000) happens to replicate them for us due to
some uarch internal details so the issue was not detected, but for LA664
(Loongson 3[A-Z]-6000) and future uarch we need to do things correctly
or we end up getting a lot of test failures.
Closes: https://bbs.aosc.io/t/topic/302
Signed-off-by: Xi Ruoyao <xry111@xry111.site>
2025-01-13 22:16:00 +08:00
Martin Kroeker
d74eb02954
Merge pull request #5057 from martin-frbg/issue5050
...
Replace while loop in generic C/ZGEMM_BETA to avoid going out of bounds
2025-01-11 11:33:56 -08:00
Martin Kroeker
30f7a4120b
Merge pull request #5056 from tingboliao/dev_omatcopy_20250108
...
Optimize the omatcopy_cn/zomatcopy_cn kernels with RVV 1.0 intrinsic.
2025-01-11 09:42:57 -08:00
gxw
20a8e48f25
LoongArch64: Update ssymv LASX version
2025-01-10 16:02:54 +08:00
gxw
e0748588b8
LoongArch64: Update dsymv LASX version
2025-01-10 14:52:57 +08:00
Martin Kroeker
d91d4fa6e9
convert the beta=0 branch to a for loop as well
2025-01-09 23:11:26 +01:00
Martin Kroeker
09e75f1588
fix absurd typo
2025-01-09 00:52:14 +01:00
Martin Kroeker
2891fd8d6d
Replace while loop with for
2025-01-08 23:17:45 +01:00
tingbo.liao
0a5dbf13d3
Optimize the omatcopy_cn and zomatcopy_cn kernels with RVV 1.0 intrinsic.
...
Signed-off-by: tingbo.liao <tingbo.liao@starfivetech.com>
2025-01-08 11:00:35 +08:00
Sergey Fedorov
229efa42ff
scal.S: use r11 on 32-bit Darwin on powerpc
2025-01-05 00:31:27 +08:00
Sergey Fedorov
81e1be8d90
Revert "temporarily disable the default S/DSCAL kernel"
...
This reverts commit 9b9c0aa5c9 .
2025-01-04 22:54:54 +08:00
Martin Kroeker
9b9c0aa5c9
temporarily disable the default S/DSCAL kernel
2025-01-03 21:36:46 +01:00
tingbo.liao
c37509c213
Optimize the nrm2_rvv function to further improve performance.
...
Signed-off-by: tingbo.liao <tingbo.liao@starfivetech.com>
2024-12-31 08:46:55 +08:00
tingbo.liao
0bea1cfd9d
Optimize the zgemm_tcopy_4_rvv function to be compatible with the situations where the vector lengths(vlens) are 128 and 256.
...
Signed-off-by: tingbo.liao <tingbo.liao@starfivetech.com>
2024-12-24 10:33:27 +08:00
tingbo.liao
d00cc400b1
Replaced the __riscv_vid_v_i32m2 and __riscv_vid_v_i64m2 with __riscv_vid_v_u32m2 and __riscv_vid_v_u64m2 for riscv64-unknown-linux-gnu-gcc compiling.
...
Signed-off-by: tingbo.liao <tingbo.liao@starfivetech.com>
2024-12-18 08:38:30 +08:00
Martin Kroeker
229d8a025e
Merge pull request #4959 from CDAC-Bengaluru/level-1-sve
...
SVE Implementation for Level-1 BLAS Routines
2024-12-13 05:20:51 -08:00
SushilPratap04
3368a4e697
Update swap_kernel_sve.c
2024-12-13 16:47:58 +05:30
CDAC-SSDG
dd71e4234a
Added Updated swap and rot sve kernels.
2024-12-13 11:15:29 +05:30
CDAC-SSDG
06ffd411a5
Update KERNEL.ARMV8SVE
2024-12-13 11:05:47 +05:30
CDAC-SSDG
765850194e
Delete kernel/arm64/swap_kernel_sve.c
2024-12-13 11:02:01 +05:30
CDAC-SSDG
c17c19fbcf
Delete kernel/arm64/swap_kernel_c.c
2024-12-13 11:01:46 +05:30
CDAC-SSDG
f6416c0e37
Delete kernel/arm64/swap.c
2024-12-13 11:01:32 +05:30
CDAC-SSDG
3b7b74664c
Delete kernel/arm64/scal_kernel_sve.c
2024-12-13 11:01:03 +05:30
CDAC-SSDG
95a97012e8
Delete kernel/arm64/scal_kernel_c.c
2024-12-13 11:00:45 +05:30
CDAC-SSDG
5540f2121e
Delete kernel/arm64/scal.c
2024-12-13 11:00:12 +05:30
CDAC-SSDG
f62519cc87
Delete kernel/arm64/rot_kernel_sve.c
2024-12-13 10:59:35 +05:30
CDAC-SSDG
10857c9df4
Delete kernel/arm64/rot_kernel_c.c
2024-12-13 10:58:51 +05:30
CDAC-SSDG
b9f51a5cf7
Delete kernel/arm64/rot.c
2024-12-13 10:58:06 +05:30
Martin Kroeker
81666de4ef
Merge pull request #5007 from martin-frbg/issue5006
...
Revert the NRM2 kernels for NeoverseN2 and ARMV8SVE targets to the generic NEON version
2024-12-05 14:43:03 -08:00
Martin Kroeker
3345007d8f
retire the thunderx2 NRM2 kernels due to reported inaccuracies and NAN
2024-12-05 21:12:06 +01:00
Martin Kroeker
5fe983db29
retire the thunderx2 nrm2 kernels for now due to NAN and inaccuracies
2024-12-05 21:09:53 +01:00
Iha, Taisei
4918beecbe
Loop-unrolled transposed [SD]GEMV kernels for A64FX and Neoverse V1
2024-12-02 18:46:00 +09:00
Juliya32
3b2421cba0
Add files via upload
2024-10-30 14:23:42 +05:30
Juliya32
012fe4da36
Delete kernel/arm64/rot_kernel_sve.c
2024-10-30 14:23:15 +05:30
Juliya32
d90ee00f85
Delete kernel/arm64/rot_kernel_c.c
2024-10-30 14:22:51 +05:30
Juliya32
668e28adc4
Delete kernel/arm64/rot.c
2024-10-30 14:22:31 +05:30
SushilPratap04
fa880ab1cf
Update KERNEL.ARMV8SVE
...
updated KERNEL.ARMV8SVE for level 1 sve (swap, rot and scal) kernels.
2024-10-30 14:09:37 +05:30
SushilPratap04
7822ae9617
Added sve kernels for rot routine.
2024-10-30 14:05:21 +05:30
SushilPratap04
b8bc2a752e
Added sve optimized kernels for swap routine
2024-10-30 14:02:57 +05:30
CDAC-SSDG
0667cf6c92
Added optimized scal routine files
2024-10-30 14:01:09 +05:30
gxw
73c6a28073
x86_64: opt somatcopy_ct with AVX
2024-10-29 07:06:15 +00:00
Ayappan Perumal
020cce1068
Fix build issues with gcc compiler as well
2024-10-23 04:24:06 -05:00
Ayappan Perumal
b6ec73e77c
Fix AIX build
2024-10-21 07:38:03 -05:00
Martin Kroeker
016bdb9b0b
Merge pull request #4946 from XiWeiGu/la64_omatcopy_lasx
...
LoongArch64: Opt somatcopy with LASX
2024-10-18 14:03:06 +02:00
Chip Kerchner
ab71a1edf2
Better VSX.
2024-10-17 08:25:02 -05:00
gxw
bb31bbef52
LoongArch64: Opt somatcopy_ct with LASX
2024-10-17 11:45:13 +00:00
gxw
b37129341b
LoongArch64: Opt somatcopy_cn with LASX
2024-10-17 11:27:55 +00:00
gxw
acf6cab304
LoongArch64: Opt somatcopy_rn with LASX
2024-10-17 09:50:02 +00:00
gxw
15edb441bf
LoongArch64: Opt somatcopy_rt with LASX
2024-10-17 09:15:42 +00:00
Chip Kerchner
36bd3eeddf
Vectorize BF16 GEMV (VSX & MMA). Use GEMM_GEMV_FORWARD_BF16 (for Power).
2024-10-13 13:46:11 -05:00
Martin Kroeker
e52d9b4cf1
Merge pull request #4928 from austinpagan/czgemm_in_c
...
CGEMM & ZGEMM using C code, Power only, P10 only.
2024-10-09 20:26:21 +02:00
Gordon Fossum
0b7fb5c791
CGEMM & ZGEMM using C code.
2024-10-09 09:42:23 -05:00
Martin Kroeker
9783dd07ab
Rename KERNEL.LOONGSONGENERIC to KERNEL.LA64_GENERIC
2024-10-06 22:43:11 +02:00
Martin Kroeker
c9e92348a6
Handle inf/nan if dummy2 flag is set
2024-10-06 19:57:17 +02:00
Martin Kroeker
d714013ab9
change sgemm kernel to 4x4 as the 16x4 altivec goes out of bounds
2024-10-03 22:04:20 +02:00
Martin Kroeker
de421b7764
Merge pull request #4904 from XiWeiGu/la64_cross_cmake
...
LoongArch64: Enable cmake cross-compilation
2024-10-03 15:53:57 +02:00
gxw
30af9278dc
LoongArch64: Enable cmake cross-compilation
2024-09-29 10:13:30 +08:00
gxw
48698b2b1d
LoongArch64: Rename core
...
Use microarchitecture name instead of meaningless strings to name the core,
the legacy core is still retained.
1. Rename LOONGSONGENERIC to LA64_GENERIC
2. Rename LOONGSON3R5 to LA464
3. Rename LOONGSON2K1000 to LA264
2024-09-29 09:35:21 +08:00
Deeksha Goplani
4894c54055
Improve TN case with further unrolling
2024-09-02 22:22:49 +05:30
Martin Kroeker
e05d98d00a
expressly use fld.d/fst.d for floating point registers instead of LD/ST macros
2024-08-15 22:14:29 +02:00
Chip Kerchner
a0aeba631d
Merge branch 'develop' into betterPowerGEMVTail
2024-08-15 08:00:00 -05:00
Chip Kerchner
083faf7556
Merge branch 'develop' into betterPowerGEMVTail
2024-08-14 15:56:03 -05:00
Chip Kerchner
75472b830a
Merge branch 'develop' into betterPowerGEMVTail
2024-08-14 10:52:46 -05:00
Henry Chen
ef94b96530
Use ldc1 and sdc1 for the prologue and epilogue on LOONGSON3A
...
This fix is similar to
2d8064174c .
2024-08-14 18:05:11 +08:00
Martin Kroeker
7ca835a82c
address clang array overflow warning
2024-08-10 13:44:56 +02:00
Martin Kroeker
46e331a917
remove the unworkable GEMM3M restriction from GENERIC again
2024-08-07 19:41:10 +02:00
Martin Kroeker
ccc23338d7
have the dummy GEMM3M kernel at least forward to regular GEMM
2024-08-07 19:39:02 +02:00
Martin Kroeker
f1c9803f9a
add proper return statement
2024-08-04 00:14:31 +02:00
Martin Kroeker
60abcc3991
add proper return statement
2024-08-04 00:13:31 +02:00
Chip Kerchner
1a7b8c650d
Merge branch 'develop' into betterPowerGEMVTail
2024-08-01 14:59:12 -05:00
Martin Kroeker
9afd0c8afd
Merge pull request #4814 from Mousius/gemv-proxy
...
Forward GEMM to GEMV when one argument is actually a vector
2024-07-31 23:18:01 +02:00
Martin Kroeker
edbf093c98
Update zarch SCAL kernels to handle INF and NAN arguments ( #4829 )
...
* handle INF and NAN in input (for S/D only if DUMMY2 argument is set)
2024-07-31 19:45:15 +02:00
Chris Sidebottom
ba2e989c67
Add accumulators to AArch64 GEMV Kernels
...
This helps to reduce values going missing as we accumulate.
2024-07-31 13:09:14 +01:00
Martin Kroeker
a875304eb0
fix inverted conditional for NAN handling
2024-07-26 09:50:20 +02:00
Martin Kroeker
24acdd6bbb
correct offset
2024-07-26 09:49:24 +02:00
Martin Kroeker
fb7c53c5e5
Merge pull request #4807 from martin-frbg/scalfixes
...
[WIP]Make NAN handling in the SCAL kernels depend on the dummy2 parameter
2024-07-25 23:42:50 +02:00
Martin Kroeker
15c53dd2e0
Merge pull request #4794 from XiWeiGu/Fixed_Numpy_CI_Test
...
Try to fixed numpy ci test failures
2024-07-25 23:42:13 +02:00
Martin Kroeker
a4e56e0452
Merge pull request #4806 from Mousius/small-gemm
...
Small GEMM for AArch64 with SVE
2024-07-25 21:50:04 +02:00
yamazaki-mitsufumi
88caf02f62
Fix ambiguous error on Mac OS
2024-07-25 22:43:13 +09:00
Martin Kroeker
b613754143
Update scal..c
2024-07-24 14:31:29 +02:00
Martin Kroeker
f5d04318e3
Merge branch 'OpenMathLib:develop' into scalfixes
2024-07-21 13:43:43 +02:00
Martin Kroeker
73f8866ffb
make NAN handling depend on DUMMY2 parameter
2024-07-21 13:42:47 +02:00
Martin Kroeker
dfbc2348a8
fix NAN handling
2024-07-20 18:27:15 +02:00
Martin Kroeker
c064319ecb
fix alpha=NAN case
2024-07-20 17:42:31 +02:00
Martin Kroeker
c2ffd90e8c
make NAN handling depend on dummy2 parameter
2024-07-20 17:31:00 +02:00
Chris Sidebottom
ea4ab3b310
Better header guard around bridge
2024-07-20 14:39:57 +01:00
Chris Sidebottom
7311d93016
Unroll TT further
2024-07-19 17:51:20 +01:00
Martin Kroeker
a815594fd1
Merge pull request #4801 from markdryan/markdryan/riscv-dynamic-arch
...
Add autodetection for riscv64
2024-07-19 17:12:07 +02:00
Martin Kroeker
dd6c33d34d
make NAN handling depend on dummy2 parameter
2024-07-19 16:14:55 +02:00
Hong Bo Peng
db98f8753f
Try to fix LAPACK testing failures on P7.
...
1. Remove the FADD insn from the GEMV Transpose code.
2. Remove the FADD insn from GEMM and ZGEMM code.
3. Reorder the compution of the Imaginary part in ZGEMM code.
2024-07-19 02:08:19 -04:00
Chris Sidebottom
a9edddb695
Unroll TN further
2024-07-18 20:04:15 +01:00
Chris Sidebottom
9984c5ce9d
Clean up k2 removal more and unroll SGEMM more
2024-07-18 18:35:43 +01:00
Chris Sidebottom
b1c9fafabb
Remove k2 loop from DGEMM TN and use a more conservative heuristic for SGEMM
2024-07-18 17:37:18 +01:00
Martin Kroeker
2020569705
fix NAN handling and make it depend on dummy2 parameter
2024-07-17 23:55:54 +02:00
Martin Kroeker
3870995f01
make NAN handling depend on dummy2 parameter
2024-07-17 23:54:24 +02:00
Martin Kroeker
7284c533b5
make NAN handling depend on dummy2 parameter
2024-07-17 23:50:40 +02:00
Martin Kroeker
73751218a4
make NAN handling depend on dummy2 parameter
2024-07-17 23:41:26 +02:00
Martin Kroeker
b9bfc8ce09
make NAN handling depend on dummy2 parameter
2024-07-17 23:29:50 +02:00
Martin Kroeker
eb4879e04c
make NAN handling depend on the dummy2 parameter
2024-07-17 23:24:19 +02:00
Martin Kroeker
ee87cb90d0
Merge pull request #4803 from iha-taisei/SVESupportSDGEMV
...
A64FX: Add support for SVE to SGEMV/DGEMV kernels.
2024-07-17 23:14:21 +02:00
gxw
34b80ce03f
mips64: Fixed numpy CI failure
2024-07-17 10:32:22 +08:00
gxw
f6d6c14a96
mips: Fixed numpy CI failure
2024-07-17 10:31:49 +08:00
Chip Kerchner
ba47c7f4f3
Vectorize reduction stage of sgemv_t.
2024-07-16 15:57:24 -05:00
iha fujitsu
0985fdc82b
A64FX: Add support for SVE to SGEMV/DGEMV kernels.
2024-07-16 17:31:33 +09:00
Mark Ryan
67bf4b6998
Fix axpby_rvv kernels for cases where inc_y = 0
...
The following openblas_utest tests fail when the RISCV64_ZVL128B is
enabled.
TEST 89/103 axpby:zaxpby_inc_0 [FAIL]
TEST 92/103 axpby:caxpby_inc_0 [FAIL]
TEST 95/103 axpby:daxpby_inc_0 [FAIL]
TEST 98/103 axpby:saxpby_inc_0 [FAIL]
The issue is that the vectorized kernels do not work when inc_y == 0.
This patch updates the kernels to fall back to the scalar algorithms
when inc_y == 0, fixing the failing tests.
Signed-off-by: Mark Ryan <markdryan@rivosinc.com>
2024-07-15 14:24:47 +00:00
Mark Ryan
3b715e6162
Add autodetection for riscv64
...
Implement DYNAMIC_ARCH support for riscv64. Three cpu types are
supported, riscv64_generic, riscv64_zvl256b, riscv64_zvl128b.
The two non-generic kernels require CPU support for RVV 1.0 to
function correctly. Detecting that a riscv64 device supports
RVV 1.0 is a little complicated as there are some boards on the
market that advertise support for V via hwcap but only support
RVV 0.7.1, which is not binary compatible with RVV 1.0. The
approach taken is to first try hwprobe. If hwprobe is not
available, we fall back to hwcap + an additional check to distinguish
between RVV 1.0 and RVV 0.7.1.
Tested on a VM with VLEN=256, a CanMV K230 with VLEN=128 (with only
the big core enabled), a Lichee Pi with RVV 0.7.1 and a VF2 with no
vector.
A compiler with RVV 1.0 support must be used to build OpenBLAS for
riscv64 when DYNAMIC_ARCH=1.
Signed-off-by: Mark Ryan <markdryan@rivosinc.com>
2024-07-15 14:24:22 +00:00
gxw
3f39c8f94f
LoongArch: Fixed numpy CI failure
2024-07-15 11:43:08 +08:00
gxw
f3cebb3ca3
x86: Fixed numpy CI failure when the target is ZEN.
2024-07-12 16:09:30 +08:00
Martin Kroeker
5d08ec7ff3
Merge pull request #4782 from martin-frbg/azurewincl
...
Fix NAN handling in ARM/generic SCAL; have AzureCI Windows show errors on failure
2024-07-11 23:55:15 +02:00
Chip Kerchner
cb154832f8
Vectorize SBGEMM incopy - 4x faster.
2024-07-09 13:10:03 -05:00
Martin Kroeker
a5c04e326a
Update scal.c
2024-07-04 22:28:01 +02:00
Martin Kroeker
536200bc9e
fix handling of INF or NAN
2024-07-04 17:47:19 +02:00
Martin Kroeker
3677b3886c
Merge pull request #4702 from bashimao/detect-nv-grace
...
Correctly detect ARM Neoverse V2 CPUs.
2024-06-30 22:48:48 +02:00
Martin Kroeker
f3c364c2cc
temporarily(?) disable the alpha=0 branch as it fails to handle INF,NAN
2024-06-27 22:18:27 +02:00