llvm-project/llvm/test/CodeGen/X86/avx-trunc.ll

; RUN: llc < %s -mtriple=x86_64-apple-darwin -mcpu=corei7-avx -mattr=+avx | FileCheck %s

define <4 x i32> @trunc_64_32(<4 x i64> %A) nounwind uwtable readnone ssp{
; CHECK-LABEL: trunc_64_32
; CHECK: pshufd
; CHECK: pshufd
; CHECK: pblendw
  %B = trunc <4 x i64> %A to <4 x i32>
  ret <4 x i32>%B
}
define <8 x i16> @trunc_32_16(<8 x i32> %A) nounwind uwtable readnone ssp{
; CHECK-LABEL: trunc_32_16
; CHECK: pshufb
  %B = trunc <8 x i32> %A to <8 x i16>
  ret <8 x i16>%B
}
define <16 x i8> @trunc_16_8(<16 x i16> %A) nounwind uwtable readnone ssp{
; CHECK-LABEL: trunc_16_8
; CHECK: pshufb
  %B = trunc <16 x i16> %A to <16 x i8>
  ret <16 x i8> %B
}
Unix line endings llvm-svn: 149615 2012-02-03 03:00:49 +08:00			`; RUN: llc < %s -mtriple=x86_64-apple-darwin -mcpu=corei7-avx -mattr=+avx \| FileCheck %s`

			`define <4 x i32> @trunc_64_32(<4 x i64> %A) nounwind uwtable readnone ssp{`
Lower AVX v4i64->v4i32 truncate to one shuffle. llvm-svn: 202996 2014-03-06 03:41:16 +08:00			`; CHECK-LABEL: trunc_64_32`
[x86] Teach the 128-bit vector shuffle lowering routines to take advantage of the existence of a reasonable blend instruction. The 256-bit vector shuffle lowering has leveraged the general technique of decomposed shuffles and blends for quite some time, but this never made it back into the 128-bit code, and there are a large number of patterns where this is substantially better. For example, this removes almost all domain crossing in vector shuffles that involve some blend and some permutation with SSE4.1 and later. See the massive reduction in 'shufps' for integer test cases in this commit. This isn't perfect yet for a few reasons: 1) The v8i16 shuffle lowering continues to plague me. We don't always form an unpack-based blend when that would be better. But the wins pretty drastically outstrip the losses here. 2) The v16i8 shuffle lowering is just a disaster here. I never went and implemented blend support here for some terrible reason. I'll do that next probably. I've not updated it for now. More variations on this technique are coming as well -- we don't shuffle-into-unpack or shuffle-into-palignr, both of which would also be profitable. Note that some test cases grow significantly in the number of instructions, but I expect to actually be faster. We use pshufd+pshufd+blendw instead of a single shufps, but the pshufd's are very likely to pipeline well (two ports on most modern intel chips) and the blend is a very fast instruction. The domain switch penalty will essentially always be more than a blend instruction, which is the only increase in tree height. llvm-svn: 229350 2015-02-16 09:52:02 +08:00			`; CHECK: pshufd`
			`; CHECK: pshufd`
			`; CHECK: pblendw`
Unix line endings llvm-svn: 149615 2012-02-03 03:00:49 +08:00			`%B = trunc <4 x i64> %A to <4 x i32>`
			`ret <4 x i32>%B`
			`}`
			`define <8 x i16> @trunc_32_16(<8 x i32> %A) nounwind uwtable readnone ssp{`
Lower AVX v4i64->v4i32 truncate to one shuffle. llvm-svn: 202996 2014-03-06 03:41:16 +08:00			`; CHECK-LABEL: trunc_32_16`
Unix line endings llvm-svn: 149615 2012-02-03 03:00:49 +08:00			`; CHECK: pshufb`
			`%B = trunc <8 x i32> %A to <8 x i16>`
			`ret <8 x i16>%B`
			`}`
X86: Custom lower sext v16i8 to v16i16, and the corresponding truncate. Also update the cost model. llvm-svn: 193270 2013-10-24 05:06:07 +08:00			`define <16 x i8> @trunc_16_8(<16 x i16> %A) nounwind uwtable readnone ssp{`
			`; CHECK-LABEL: trunc_16_8`
			`; CHECK: pshufb`
			`%B = trunc <16 x i16> %A to <16 x i8>`
			`ret <16 x i8> %B`
			`}`