llvm-project/llvm/test/Analysis/CostModel/AArch64/store.ll

; RUN: opt < %s  -cost-model -analyze -mtriple=aarch64-apple-ios | FileCheck %s
; RUN: opt < %s  -cost-model -analyze -mtriple=aarch64-apple-ios -mattr=slow-misaligned-128store | FileCheck %s --check-prefix=SLOW_MISALIGNED_128_STORE

target datalayout = "e-p:32:32:32-i1:8:32-i8:8:32-i16:16:32-i32:32:32-i64:32:64-f32:32:32-f64:32:64-v64:32:64-v128:32:128-a0:0:32-n32-S32"
; CHECK-LABEL: getMemoryOpCost
; SLOW_MISALIGNED_128_STORE-LABEL: getMemoryOpCost
define void @getMemoryOpCost() {
    ; If FeatureSlowMisaligned128Store is set, we penalize <2 x i64> stores. On
    ; Cyclone, for example, such stores should be expensive because we don't
    ; split them and misaligned 16b stores have bad performance.
    ;
    ; CHECK: cost of 1 {{.*}} store
    ; SLOW_MISALIGNED_128_STORE: cost of 12 {{.*}} store
    store <2 x i64> undef, <2 x i64> * undef

    ; We scalarize the loads/stores because there is no vector register name for
    ; these types (they get extended to v.4h/v.2s).
    ; CHECK: cost of 16 {{.*}} store
    store <2 x i8> undef, <2 x i8> * undef
    ; CHECK: cost of 64 {{.*}} store
    store <4 x i8> undef, <4 x i8> * undef
    ; CHECK: cost of 16 {{.*}} load
    load <2 x i8> , <2 x i8> * undef
    ; CHECK: cost of 64 {{.*}} load
    load <4 x i8> , <4 x i8> * undef

    ret void
}
[AArch64] Guard Misaligned 128-bit store penalty by subtarget feature This patch checks that the SlowMisaligned128Store subtarget feature is set when penalizing such stores in getMemoryOpCost. Differential Revision: https://reviews.llvm.org/D27677 llvm-svn: 289845 2016-12-16 02:36:59 +08:00			`; RUN: opt < %s -cost-model -analyze -mtriple=aarch64-apple-ios \| FileCheck %s`
			`; RUN: opt < %s -cost-model -analyze -mtriple=aarch64-apple-ios -mattr=slow-misaligned-128store \| FileCheck %s --check-prefix=SLOW_MISALIGNED_128_STORE`

ARM64: initial backend import This adds a second implementation of the AArch64 architecture to LLVM, accessible in parallel via the "arm64" triple. The plan over the coming weeks & months is to merge the two into a single backend, during which time thorough code review should naturally occur. Everything will be easier with the target in-tree though, hence this commit. llvm-svn: 205090 2014-03-29 18:18:08 +08:00			`target datalayout = "e-p:32:32:32-i1:8:32-i8:8:32-i16:16:32-i32:32:32-i64:32:64-f32:32:32-f64:32:64-v64:32:64-v128:32:128-a0:0:32-n32-S32"`
[AArch64] Guard Misaligned 128-bit store penalty by subtarget feature This patch checks that the SlowMisaligned128Store subtarget feature is set when penalizing such stores in getMemoryOpCost. Differential Revision: https://reviews.llvm.org/D27677 llvm-svn: 289845 2016-12-16 02:36:59 +08:00			`; CHECK-LABEL: getMemoryOpCost`
			`; SLOW_MISALIGNED_128_STORE-LABEL: getMemoryOpCost`
			`define void @getMemoryOpCost() {`
			`; If FeatureSlowMisaligned128Store is set, we penalize <2 x i64> stores. On`
			`; Cyclone, for example, such stores should be expensive because we don't`
			`; split them and misaligned 16b stores have bad performance.`
			`;`
			`; CHECK: cost of 1 {{.*}} store`
			`; SLOW_MISALIGNED_128_STORE: cost of 12 {{.*}} store`
ARM64: initial backend import This adds a second implementation of the AArch64 architecture to LLVM, accessible in parallel via the "arm64" triple. The plan over the coming weeks & months is to merge the two into a single backend, during which time thorough code review should naturally occur. Everything will be easier with the target in-tree though, hence this commit. llvm-svn: 205090 2014-03-29 18:18:08 +08:00			`store <2 x i64> undef, <2 x i64> * undef`

			`; We scalarize the loads/stores because there is no vector register name for`
			`; these types (they get extended to v.4h/v.2s).`
			`; CHECK: cost of 16 {{.*}} store`
			`store <2 x i8> undef, <2 x i8> * undef`
			`; CHECK: cost of 64 {{.*}} store`
			`store <4 x i8> undef, <4 x i8> * undef`
			`; CHECK: cost of 16 {{.*}} load`
[opaque pointer type] Add textual IR support for explicit type parameter to load instruction Essentially the same as the GEP change in r230786. A similar migration script can be used to update test cases, though a few more test case improvements/changes were required this time around: (r229269-r229278) import fileinput import sys import re pat = re.compile(r"((?:=\|:\|^)\sload (?:atomic )?(?:volatile )?(.?))(\| addrspace\(\d+\) )\($\| (?:%\|@\|null\|undef\|blockaddress\|getelementptr\|addrspacecast\|bitcast\|inttoptr\|\[\[[a-zA-Z]\|\{\{).$)") for line in sys.stdin: sys.stdout.write(re.sub(pat, r"\1, \2\3*\4", line)) Reviewers: rafael, dexonsmith, grosser Differential Revision: http://reviews.llvm.org/D7649 llvm-svn: 230794 2015-02-28 05:17:42 +08:00			`load <2 x i8> , <2 x i8> * undef`
ARM64: initial backend import This adds a second implementation of the AArch64 architecture to LLVM, accessible in parallel via the "arm64" triple. The plan over the coming weeks & months is to merge the two into a single backend, during which time thorough code review should naturally occur. Everything will be easier with the target in-tree though, hence this commit. llvm-svn: 205090 2014-03-29 18:18:08 +08:00			`; CHECK: cost of 64 {{.*}} load`
[opaque pointer type] Add textual IR support for explicit type parameter to load instruction Essentially the same as the GEP change in r230786. A similar migration script can be used to update test cases, though a few more test case improvements/changes were required this time around: (r229269-r229278) import fileinput import sys import re pat = re.compile(r"((?:=\|:\|^)\sload (?:atomic )?(?:volatile )?(.?))(\| addrspace\(\d+\) )\($\| (?:%\|@\|null\|undef\|blockaddress\|getelementptr\|addrspacecast\|bitcast\|inttoptr\|\[\[[a-zA-Z]\|\{\{).$)") for line in sys.stdin: sys.stdout.write(re.sub(pat, r"\1, \2\3*\4", line)) Reviewers: rafael, dexonsmith, grosser Differential Revision: http://reviews.llvm.org/D7649 llvm-svn: 230794 2015-02-28 05:17:42 +08:00			`load <4 x i8> , <4 x i8> * undef`
ARM64: initial backend import This adds a second implementation of the AArch64 architecture to LLVM, accessible in parallel via the "arm64" triple. The plan over the coming weeks & months is to merge the two into a single backend, during which time thorough code review should naturally occur. Everything will be easier with the target in-tree though, hence this commit. llvm-svn: 205090 2014-03-29 18:18:08 +08:00
			`ret void`
			`}`