davidz-ampere
be68ef03b4
Add support for Ampere processors
2025-06-15 22:00:40 -04:00
Srangrang
9f13b2c6ac
style: modify HALF to BFLOAT16 in benchmark folder
2025-06-15 20:57:05 +08:00
Srangrang
ec14e1648c
fix: resolve non-RISCV host build failed issue
...
- adjust interface to disable "small matrix" pathway
- separate HFLOAT16 from BFLOAT16
- remove SHGEMM_UNROLL_M and SHGEMM_UNROLL_N equal conditions
Related to PR#5290
Co-authored-by Martin
2025-06-15 20:25:15 +08:00
Martin Kroeker
e338d34ce1
fix path
2025-06-13 13:37:15 +02:00
Martin Kroeker
d36093d084
temporarily change default C/ZSCAL to the non-asm implementation
2025-06-13 13:32:02 +02:00
Martin Kroeker
b3c90564d7
resync with the generic arm version for inf/nan handling
2025-06-13 00:54:27 -07:00
Martin Kroeker
6bdc7f9eb7
Merge pull request #5300 from martin-frbg/fixup5296
...
kernel/riscv64: Fix cscal/zscal for riscv64_generic
2025-06-12 15:02:22 -07:00
Martin Kroeker
73af02b89f
use dummy2 as Inf/NAN handling flag
2025-06-12 13:33:56 -07:00
Martin Kroeker
549a9f1dbb
Disable the default SSE kernels for CSCAL/ZSCAL for now
2025-06-12 18:54:33 +02:00
Martin Kroeker
58eeb9041c
fix handling of dummy2
2025-06-12 03:03:01 -07:00
Martin Kroeker
7c77537b25
Merge pull request #5297 from martin-frbg/zscal_x86_sparc
...
kernel/(x86|sparc): Fix cscal and zscal by reverting to the generic C kernels
2025-06-12 01:10:35 -07:00
Martin Kroeker
63287e1855
Merge pull request #5296 from martin-frbg/zscal_riscv
...
kernel/riscv64: Fix cscal and zscal
2025-06-12 01:10:15 -07:00
Martin Kroeker
d2855d3dab
Merge pull request #5285 from martin-frbg/zscal_zarch
...
kernel/zarch: Fix cscal and zscal
2025-06-12 01:09:52 -07:00
Martin Kroeker
1408be5fe0
Merge pull request #5282 from martin-frbg/zscal_power
...
kernel/power: Fixed cscal and zscal
2025-06-12 01:04:38 -07:00
Martin Kroeker
1589d0b21e
Merge pull request #5281 from martin-frbg/zscal_arm64
...
kernel/arm64: fixed cscal and zscal
2025-06-12 01:04:18 -07:00
Martin Kroeker
a86419fb66
Merge pull request #5280 from martin-frbg/zscal_x86_64
...
kernel/x86_64: fixed cscal and zscal
2025-06-12 01:03:55 -07:00
Martin Kroeker
11ff18bb0f
Merge pull request #5081 from XiWeiGu/kernel_generic_fixed_cscal_zscal
...
kernel/generic: Fixed cscal and zscal
2025-06-12 01:03:00 -07:00
Martin Kroeker
f4194fc65f
Merge branch 'develop' into la64_fixed_cscal_zscal
2025-06-11 14:28:41 -07:00
Martin Kroeker
e12132abd4
Use generic C/ZSCAL kernels to address inf/nan handling for now
2025-06-11 22:12:10 +02:00
Martin Kroeker
1cefbea7ea
Use generic SCAL kernels to address inf/nan handling for now
2025-06-11 22:10:46 +02:00
Sharif Inamdar
8279e68805
Optimize gemv_n_sve_v1x3 kernel
...
- Calculate predicate outside the loop
- Divide matrix in blocks of 3
2025-06-11 10:16:56 +00:00
Martin Kroeker
f18b7a46bf
add dummy2 flag handling for inf/nan agnostic zeroing
2025-06-11 01:47:43 -07:00
Martin Kroeker
fe220a0d7d
Merge pull request #5291 from guoyuanplct/develop
...
kernel/riscv64:fixed the performance problem in RISCV64_ZVL256 when OPENBLAS_K is small
2025-06-09 23:42:04 -07:00
Arne Juul
5442aff218
Accumulate results in output register explicitly
2025-06-09 19:03:22 +00:00
guoyuanplct
2ae019161a
fixed the performance problem in RISCV64_ZVL256 when OPENBLAS_K is small
2025-06-05 21:53:03 +08:00
Srangrang
fb89820f20
Merge branch 'develop' of https://github.com/Srangrang/OpenBLAS into develop
2025-06-04 20:27:05 +08:00
Srangrang
4e1a381e5b
fix: resolve the compilation failure without zfh instruction
...
- modify the macro conditions in Makefile.system
- Delete development test code
Related to issue#5279
2025-06-04 20:00:12 +08:00
gkdddd
670ec6f757
Added shgemm_kernel_8x8 for RISCV64_ZVL128B and shgemm_kernel_16x8 for RISCV64_ZVL256B
...
Added HFLOAT16 support for RISCV64
Added shgemm_kernel_8x8 for RISCV64_ZVL128B and shgemm_kernel_16x8 for RISCV64_ZVL256B based on HFLOAT16
The instruction sets used are ZVFH and ZFH, which need to be supported by RVV1.0
Related to issue #5279
Co-authored-by Linjin Li <linjin_li@163.com>
2025-06-03 20:14:30 +08:00
guoyuanplct
d2003dc886
del lines
2025-05-29 18:38:22 +08:00
guoyuanplct
45fd2d9b07
Optimized the axpby function.
2025-05-29 17:50:44 +08:00
Martin Kroeker
fb8dc8ff5c
Add dummy2 flag handling
2025-05-25 14:47:06 -07:00
Srangrang
2996c25c94
add shgemm for RISCV_ZVL128B
2025-05-24 23:55:49 +08:00
Martin Kroeker
cf06250d36
add handling of dummy2 flag
2025-05-24 06:06:24 -07:00
Martin Kroeker
28f8fdaf0f
support flag for NaN/Inf handling and fix scaling of NaN/Inf values
2025-05-23 14:59:59 +02:00
Martin Kroeker
669c847ceb
support extra flag for NaN handling
2025-05-23 05:52:48 -07:00
Martin Kroeker
0b0bb9951d
Merge pull request #5265 from guoyuanplct/develop
...
kernel/riscv64:Added support for omatcopy on RISCV64_ZVL256B
2025-05-17 05:08:47 -07:00
guoyuanplct
be9f7550b5
Format Code
2025-05-15 18:55:47 +08:00
guoyuanplct
4d213653d8
kernel/riscv64:Added support for omatcopy on riscv64.
2025-05-15 13:29:14 +08:00
Martin Kroeker
8afddc1a81
Merge pull request #5262 from guoyuanplct/develop
...
kernel/riscv64:Fixed the bug of openblas_utest_ext failing in c/zgemv and some c/zgbmv tests:
2025-05-14 02:40:32 -07:00
guoyuanplct
9a7e3f102b
kernel/riscv64:Fixed the bug of openblas_utest_ext failing in c/zgemv and some c/zgbmv tests:
2025-05-14 00:09:26 +08:00
pengxu
a978ad3180
Loongarch64: add C functions of zgemm_ncopy_16
2025-05-13 16:09:12 +08:00
pengxu
0ccb050583
Loongarch64: fixed cgemm_ncopy_16_lasx
2025-05-13 16:08:33 +08:00
Martin Kroeker
5141a90993
Fix ARMV9SME target in DYNAMIC_ARCH and add SME query code for MacOS ( #5222 )
...
* Fix ARMV9SME target and add support_sme1 code for MacOS
* make sgemm_direct unconditionally available on all arm64
* build a (dummy) sgemm_direct kernel on all arm64
* Update dynamic_arm64.c
2025-05-10 22:39:32 +02:00
Martin Kroeker
151b74284e
Merge pull request #5203 from quic/fix-sgemmdirect-sme1
...
Add vector registers to clobber list to prevent compiler optimization.
2025-05-09 05:39:47 -07:00
Martin Kroeker
cba32d001a
Merge pull request #5245 from guoyuanplct/develop
...
Optimized RVV_ZVL256B Implementation of zgemv_n
2025-05-01 03:04:38 -07:00
pengxu
f19e72c402
Loongarch64: fixed swap_lasx
2025-04-30 16:42:52 +08:00
pengxu
b471fa337b
Loongarch64: fixed snrm2_lasx
2025-04-30 16:42:36 +08:00
pengxu
57bb46bedf
Loongarch64: fixed rot_lasx
2025-04-30 16:42:22 +08:00
pengxu
6dc4ca2391
Loongarch64: fixed icamax_lasx
2025-04-30 16:42:12 +08:00
pengxu
b528b1b8ea
Loongarch64: fixed iamax_lasx
2025-04-30 16:41:58 +08:00
pengxu
ba9569e382
Loongarch64: fixed dot_lasx
2025-04-30 16:41:48 +08:00
pengxu
dc5fa29851
Loongarch64: fixed cscal_lasx
2025-04-30 16:41:39 +08:00
pengxu
a98dd6d911
Loongarch64: fixed copy_lasx
2025-04-30 16:41:28 +08:00
pengxu
d49319c2d2
Loongarch64: fixed cnrm2_lasx
2025-04-30 16:41:18 +08:00
pengxu
74c97ef814
Loongarch64: fixed cdot_lasx
2025-04-30 16:41:05 +08:00
pengxu
be525521ad
Loongarch64: fixed asum_lasx
2025-04-30 16:40:55 +08:00
pengxu
0cd5ca5527
Loongarch64: fixed amax_lasx
2025-04-30 16:40:44 +08:00
guoyuanplct
11ffc8680e
Format the code
2025-04-25 00:27:27 +08:00
guoyuanplct
7616c42095
Optimized RVV_ZVL256B Implementation of zgemv_n
...
The implementation of zgemv_n using RVV_ZVL256B has been optimized.
Compared to the previous implementation, it has achieved a 1.5x
performance improvement.
2025-04-25 00:05:15 +08:00
abhishek-fujitsu
9c02cdb073
optimise dot using thread throttling for NEOVERSE V1
2025-04-23 22:35:05 +05:30
Martin Kroeker
d0e8fd6d40
Merge pull request #5239 from annop-w/gemv_n_sve
...
Use SVE kernel for S/DGEMVN for SVE machines
2025-04-22 10:19:49 -07:00
Iha, Taisei
08b5c18d70
fixed a potential out-of-bounds on gemv.
2025-04-22 19:56:44 +09:00
Annop Wongwathanarat
e11744a411
Use SVE kernel for S/DGEMVN for SVE machines
2025-04-22 09:40:13 +00:00
Martin Kroeker
db0abfa907
Merge pull request #5238 from martin-frbg/revert5125
...
remove non-vectorized SGEMV transpose reduce path for POWER8, restoring optimizations frpm PR4880
2025-04-22 02:12:19 -07:00
Martin Kroeker
7389b6c483
Merge pull request #5237 from martin-frbg/revert5219
...
Fix and reinstate the Cooper Lake/Sapphire Rapids microkernel for non-transpose SBGEMV
2025-04-21 23:36:23 -07:00
Martin Kroeker
4ec62d7f73
remove non-vectorized code path for power8, restoring PR4880
2025-04-21 23:14:10 +02:00
Martin Kroeker
1df8738f27
Merge pull request #5235 from quickwritereader/issue_unaligned_ppc64le
...
Explicit unaligned vector load/stores in PPC64LE GEMV kernels
2025-04-21 14:03:56 -07:00
Martin Kroeker
99d9f1ff38
Fix conditional
2025-04-21 22:55:45 +02:00
Martin Kroeker
96d80801bc
Reinstate the CooperLake microkernel
2025-04-21 22:53:26 +02:00
Martin Kroeker
2e4309315c
Merge pull request #5219 from martin-frbg/sbgemvn_cooper
...
Temporarily disable the Cooper Lake/Sapphire Rapids microkernel for non-transpose SBGEMV
2025-04-20 07:29:20 -07:00
Ubuntu
0cc2485594
Explicit unaligned vector load/stores in PPC64LE GEMV kernels
2025-04-20 08:00:29 +00:00
Martin Kroeker
dd38b4e811
Merge pull request #5225 from annop-w/gemv_n
...
Improve performance for SGEMVN on NEONVERSEN1
2025-04-17 01:54:10 -07:00
Martin Kroeker
0241d516f6
Merge pull request #5220 from iha-taisei/sdgemv_n_unroll
...
Further performance improvements to non-transposed [SD]GEMV kernels for A64FX and Neoverse V1.
2025-04-16 12:55:55 -07:00
Annop Wongwathanarat
d535728803
Improve performance for SGEMVN on NEONVERSEN1
2025-04-16 09:54:30 +00:00
Usui, Tetsuzo
d711906e3e
Add symv kernels for arm64
2025-04-11 20:39:52 +09:00
Iha, Taisei
f1e628b889
Further performance improvements to [SD]GEMV.
2025-04-11 20:00:33 +09:00
Martin Kroeker
211dfd0754
disable the CooperLake microkernel as it produces wrong results
2025-04-10 22:21:57 +02:00
Martin Kroeker
b30dc9701f
Merge pull request #5215 from annop-w/gemv_t
...
Use SVE kernel for S/DGEMVT for SVE machines
2025-04-10 13:06:07 -07:00
Martin Kroeker
2893d0add4
Merge pull request #5211 from guoyuanplct/develop
...
Optimizing the Implementation of GEMV on the RISC-V V Extension
2025-04-10 09:43:03 -07:00
Annop Wongwathanarat
ec146157d3
Use SVE kernel for S/DGEMVT for SVE machines
2025-04-09 20:38:14 +00:00
Martin Kroeker
70865a894e
Merge pull request #5180 from ywwry66/openmp_use_cmake
...
CMake: Pass `OpenMP` compiler and linker flags through CMake targets
2025-04-08 13:16:07 -07:00
lglglglgy
1ff303f36e
Optimizing the Implementation of GEMV on the RISC-V V Extension
...
Specialized some scenarios, performed loop unrolling, and reduced the
number of multiplications.
2025-04-08 21:18:00 +08:00
ColumbusAI
7bf848454d
Update zsum.c -- fixed spelling error to successfully compile
...
spelling error where zsum_kernel is used and it should be zasum_kernel. Will not compile without fix.
2025-04-05 09:57:53 -07:00
Vaisakh K V
04915be829
Add vector registers to clobber list to prevent compiler optimization.
...
SME based SGEMMDIRECT kernel uses the vector registers (z) and adding
clobber list informs compiler not to optimize these registers.
2025-04-03 12:18:43 +05:30
Egbert Eich
ea6515c4b3
On zarch don't produce objects from assembler with a writable stack section
...
On z-series, the current version of the GNU toolchain produces warnings
such as:
```
/usr/lib64/gcc/[...]/s390x-suse-linux/bin/ld: warning: ztrmm_kernel_RC_Z14.o: missing .note.GNU-stack section implies
executable stack
/usr/lib64/[...]/s390x-suse-linux/bin/ld: NOTE: This behaviour is deprecated and will be removed in a future version of the linker
```
To prevent this message and make sure we are future proof, add
```
.section .note.GNU-stack,"",@progbits
```
Also add the `.size` bit to give the asm defined functions a proper size
in the symbol table.
Signed-off-by: Egbert Eich <eich@suse.com>
2025-03-28 18:47:48 +01:00
Ruiyang Wu
02fd1df10b
CMake: Pass `OpenMP` compiler and linker flags through CMake targets
...
Using `OpenMP::OpenMP_LANG` targets for CMake is less error-prone than
passing the compiler and linker flags manually. Furthermore, it allows
the user to customize those flags by setting `OpenMP_LANG_FLAGS`,
`OpenMP_LANG_LIB_NAMES`, and `OpenMP_omp_LIBRARY`.
2025-03-26 23:09:54 -04:00
Ye Tao
f27ba5efd1
fix bugs in aarch64 sbgemv_n kernel
2025-03-14 17:55:40 +00:00
Annop Wongwathanarat
edef2e4441
Fix bug in ARM64 sbgemv_t
2025-03-13 20:55:31 +00:00
Martin Kroeker
b55ca71d5b
Merge pull request #5182 from annop-w/sgemm_ncopy
...
Optimize aarch64 sgemm_ncopy
2025-03-13 16:04:39 +01:00
Martin Kroeker
2f778554b8
Merge pull request #5181 from taoye9/change_sbgemn_cast_bf16
...
replace customize bf16_to_fp32 with arm neon vcvtah_f32_bf16
2025-03-13 13:50:26 +01:00
Annop Wongwathanarat
9807f56580
Optimize aarch64 sgemm_ncopy
2025-03-13 10:17:43 +00:00
Martin Kroeker
a3e7b16072
Merge pull request #5157 from manaalmj/feature
...
Optimize gemv_n_sve kernel
2025-03-12 21:08:23 +01:00
Ye Tao
4c00099ed6
replace customize bf16_to_fp32 with arm neon vcvtah_f32_bf16
2025-03-12 16:20:15 +00:00
Annop Wongwathanarat
a085b6c9ec
Fix aarch64 sbgemv_t compilation error for GCC < 13
2025-03-12 14:52:42 +00:00
manjam01
5c4e38ab17
Optimize gemv_n_sve kernel
2025-03-10 16:39:20 +00:00
Martin Kroeker
1d5ed5c46b
Merge pull request #5168 from taoye9/add_sbgemvn_on_neonversen2
...
Add dispatch of SBGEMVNKERNEL for NEOVERSEN2 and NEOVERSEV2
2025-03-04 16:39:22 +01:00
Ye Tao
6b8b35cdf2
fix minior issues of redeclaration of float x0,x1 in sbgemv_n_neon.c
2025-03-03 11:55:27 +00:00
Ye Tao
38ee7c9301
Add dispatch of SBGEMVNKERNEL for NEOVERSEN2 and NEOVERSEV2
2025-03-03 11:32:05 +00:00
Martin Kroeker
2b941c44b5
Merge branch 'develop' into sbgemv_n_neon
2025-03-02 22:39:32 +01:00
Ye Tao
35bdbca153
Add sbgemv_n_neon kernel for arm64.
2025-02-28 14:37:06 +00:00
Annop Wongwathanarat
edaf51dd99
Add sbgemv_t_bfdot kernel for ARM64
...
This improves performance for sbgemv_t by up to 100x on NEOVERSEV1.
The geometric mean speedup is ~61x for M=N=[2,512].
2025-02-28 12:31:50 +00:00
Martin Kroeker
77fba0f400
Fix "dummy2" flag handling
2025-02-22 20:09:21 +01:00
Martin Kroeker
eb84aac7ad
Merge pull request #5084 from quic/topic/sgemm_direct_sme1
...
Support for SGEMM_DIRECT Kernel based on SME1
2025-02-19 10:56:49 +01:00
Martin Kroeker
b9ae246f20
define USE_TRMM for RISCV64 targets as well
2025-02-16 23:18:04 +01:00
Vaisakh K V
f66ca05b31
Merge branch 'develop' into topic/sgemm_direct_sme1
2025-02-13 14:54:37 +05:30
Vaisakh K V
d23eb3b93e
Support for SME1 based sgemm_direct kernel for cblas_sgemm level 3 API
...
* Added ARMV9SME target
* Added SGEMM_DIRECT kernel based on SME1
2025-02-13 14:51:21 +05:30
Martin Kroeker
8d487ef6eb
Merge pull request #5124 from XiWeiGu/LoongArch64-LA264-lapack-fixed
...
LoongArch64: Fixed lapack test for LA264
2025-02-12 14:58:30 +01:00
Martin Kroeker
81eed868b6
Restore the non-vectorized code from before PR4880 for POWER8
2025-02-12 09:07:20 +01:00
Martin Kroeker
98b5ef929c
Restore the non-vectorized code from before PR4880 for POWER8
2025-02-12 09:04:22 +01:00
gxw
2c4a5cc6e6
LoongArch64: Fixed snrm2_lsx.S and cnrm2_lsx.S
...
When the data type is single-precision real or single-precision complex,
converting it to double precision does not prevent overflow (as exposed in LAPACK tests).
The only solution is to follow C's approach: find the maximum value in the
array and divide each element by that maximum to avoid this issue
2025-02-12 15:48:01 +08:00
gxw
9e75d6b3d1
LoongArch64: Fixed swap_lsx.S
...
Fixed the error when the stride is zero
2025-02-12 14:57:35 +08:00
gxw
e8c740368c
LoongArch64: Fixed rot_lsx.S ane crot_lsx.S
...
Do not check whether the input parameters c and s are zero,
as this may cause errors with special values (same as scal).
Although OpenBLAS's own test suite doesn't catch this, it will
cause LAPACK test cases to fail.
2025-02-12 14:52:49 +08:00
Hao Chen
c2212d0abd
LoongArch64: Fixed copy_lsx.S
...
Fixed incorrect store operation
Signed-off-by: gxw <guxiwei-hf@loongson.cn>
2025-02-12 14:52:20 +08:00
Hao Chen
7f1ebc7ae6
LoongArch64: Fixed iamax_lsx.S
...
Fixed index retrieval issue when there are
identical maximum absolute values
Signed-off-by: Hao Chen <chenhao@loongson.cn>
Signed-off-by: gxw <guxiwei-hf@loongson.cn>
2025-02-12 14:44:44 +08:00
Hao Chen
31d326f895
LoongArch64: Fixed dot_lsx.S
...
Fixed incorrect register usage in instructions
Signed-off-by: gxw <guxiwei-hf@loongson.cn>
2025-02-12 14:44:11 +08:00
Hao Chen
5d6356bc16
LoongArch64: Fixed amax_lsx.S
...
Fixed register zeroing operation
Signed-off-by: Hao Chen <chenhao@loongson.cn>
Signed-off-by: gxw <guxiwei-hf@loongson.cn>
2025-02-12 14:39:29 +08:00
Ye Tao
c748e6a338
optimized sbgemm kernel for neoverse-v1 (sve-256)
...
Signed-off-by: Ye Tao <ye.tao@arm.com>
2025-02-05 10:06:37 +00:00
Aditya Tewari
4379a6fbe3
* checkpoint sbgemm for SVE-256
2025-02-03 12:49:49 +00:00
Martin Kroeker
d7036cfd74
Remove trailing blanks that break the cmake parser
2025-01-27 09:32:17 +01:00
Martin Kroeker
6e393a5599
Merge branch 'develop' into gemv_t
2025-01-25 12:54:04 +01:00
Martin Kroeker
876ba58e28
Merge pull request #5091 from goplanid/develop
...
Small gemm kernel improvements for AArch64
2025-01-24 10:59:16 +01:00
Martin Kroeker
180ba5e7d0
Merge pull request #5069 from tingboliao/dev_rotm_20250107
...
Further rearranged the rotm kernel for the different architectures.
2025-01-23 10:16:43 +01:00
Deeksha Goplani
d1bfa979f7
small gemm kernel packing modifications
2025-01-23 09:41:45 +05:30
Martin Kroeker
1a6a9fb22f
add another generator line for rotm
2025-01-23 00:17:04 +01:00
Martin Kroeker
4924319c50
fix position of srotm, qrotm
2025-01-22 16:07:35 +01:00
Martin Kroeker
b58cba9eb6
fix qrotm build rules
2025-01-22 15:51:49 +01:00
gxw
2da86b80c9
LoongArch64: Fixed scalar version of cscal and zscal
2025-01-22 14:32:20 +08:00
gxw
5392f6df69
LoongArch64: Fixed LASX version of cscal and zscal
2025-01-22 14:15:27 +08:00
tingbo.liao
3c8df6358f
Further rearranged the rotm kernel for the different architectures.
...
Signed-off-by: tingbo.liao <tingbo.liao@starfivetech.com>
2025-01-22 11:41:12 +08:00
Annop Wongwathanarat
c0318cea6e
Simplify gemv_t_sve_v1x3 kernel
2025-01-21 13:40:17 +00:00
gxw
b2117bb2ca
LoongArch64: Fixed LSX version of cscal and zscal
2025-01-21 14:49:44 +08:00
gxw
e114880dc4
kernel/generic: Fixed cscal and zscal
2025-01-21 11:44:22 +08:00
Martin Kroeker
87083fdbf6
[WIP] Work around assembler limitations in current LLVM for Windows on Arm ( #5076 )
...
* Protect align directives in assembly files that are currently problematic with LLVM on WoA
* use the armv8 zdot on WoA to work around other LLVM issues
2025-01-18 16:45:56 +01:00
tingbo.liao
ef7f54b357
Optimized the gemm_tcopy_8_rvv to be compatible with the vlens 128 and 256.
...
Signed-off-by: tingbo.liao <tingbo.liao@starfivetech.com>
2025-01-15 11:31:28 +08:00
gxw
e0a8216554
LoongArch64: Update dsymv LSX version
2025-01-14 19:45:42 +08:00
gxw
a9070ba3f9
LoongArch64: Update ssymv LSX version
2025-01-14 09:06:59 +00:00
Xi Ruoyao
af10c132b8
LoongArch64: Fix dsymv and ssymv LASX version
...
"fmov.d $f2, $f4" leaves all the bits higher than the 63-th bit
unpredictable but it's obvious that the following code uses the value of
those high bits. We actually want to replicate the lower 64 bits here,
so we should use xvreplve0.d instead.
LA464 (Loongson 3[A-Z]-5000) happens to replicate them for us due to
some uarch internal details so the issue was not detected, but for LA664
(Loongson 3[A-Z]-6000) and future uarch we need to do things correctly
or we end up getting a lot of test failures.
Closes: https://bbs.aosc.io/t/topic/302
Signed-off-by: Xi Ruoyao <xry111@xry111.site>
2025-01-13 22:16:00 +08:00
Martin Kroeker
d74eb02954
Merge pull request #5057 from martin-frbg/issue5050
...
Replace while loop in generic C/ZGEMM_BETA to avoid going out of bounds
2025-01-11 11:33:56 -08:00
Martin Kroeker
30f7a4120b
Merge pull request #5056 from tingboliao/dev_omatcopy_20250108
...
Optimize the omatcopy_cn/zomatcopy_cn kernels with RVV 1.0 intrinsic.
2025-01-11 09:42:57 -08:00
gxw
20a8e48f25
LoongArch64: Update ssymv LASX version
2025-01-10 16:02:54 +08:00
gxw
e0748588b8
LoongArch64: Update dsymv LASX version
2025-01-10 14:52:57 +08:00
Martin Kroeker
d91d4fa6e9
convert the beta=0 branch to a for loop as well
2025-01-09 23:11:26 +01:00
Martin Kroeker
09e75f1588
fix absurd typo
2025-01-09 00:52:14 +01:00
Martin Kroeker
2891fd8d6d
Replace while loop with for
2025-01-08 23:17:45 +01:00
tingbo.liao
0a5dbf13d3
Optimize the omatcopy_cn and zomatcopy_cn kernels with RVV 1.0 intrinsic.
...
Signed-off-by: tingbo.liao <tingbo.liao@starfivetech.com>
2025-01-08 11:00:35 +08:00
Sergey Fedorov
229efa42ff
scal.S: use r11 on 32-bit Darwin on powerpc
2025-01-05 00:31:27 +08:00
Sergey Fedorov
81e1be8d90
Revert "temporarily disable the default S/DSCAL kernel"
...
This reverts commit 9b9c0aa5c9 .
2025-01-04 22:54:54 +08:00
Martin Kroeker
9b9c0aa5c9
temporarily disable the default S/DSCAL kernel
2025-01-03 21:36:46 +01:00
tingbo.liao
c37509c213
Optimize the nrm2_rvv function to further improve performance.
...
Signed-off-by: tingbo.liao <tingbo.liao@starfivetech.com>
2024-12-31 08:46:55 +08:00
tingbo.liao
0bea1cfd9d
Optimize the zgemm_tcopy_4_rvv function to be compatible with the situations where the vector lengths(vlens) are 128 and 256.
...
Signed-off-by: tingbo.liao <tingbo.liao@starfivetech.com>
2024-12-24 10:33:27 +08:00