llvm-project

Commit Graph

Author	SHA1	Message	Date
Hongtao Yu	9641b9be9d	[Inliner] Preserve !prof metadata when converting call to invoke. When a callee function is inlined via an invoke instruction, every function call inside the callee, if not an invoke, will be converted to an invoke after cloned to the caller body. I found that during the conversion the !prof metadata was dropped. This in turned caused a cloned indirect call not properly promoted in subsequent passes. The particular scenario I was investigating was with AutoFDO and thinLTO. In prelink, no ICP was triggered (neither by the sample loader nor PGO ICP), no indirect call was promoted. This is because 1) the particular indirect call did not have inlined samples; and 2) PGO ICP was intentionally disabled. After inlining, the prof metadata was dropped. Then in postlink, PGO ICP jumped in but didn't do anything. Thus the opportunity was missed. I'm making a simple fix to preserve !prof metadata when converting call to invoke. Reviewed By: davidxl Differential Revision: https://reviews.llvm.org/D125249	2022-05-09 15:08:09 -07:00
Alexey Bataev	4212ef8a0e	Revert "[SLP]Further improvement of the cost model for scalars used in buildvectors." This reverts commit `99f31acfce` and several others to fix detected crashes, reported in https://reviews.llvm.org/D115750	2022-05-09 13:46:06 -07:00
Alexey Bataev	cce80bd8b7	[SLP]Adjust assertion check for scalars in several insertelements. If the same scalar is inserted several times into the same buildvector, the mask index can be used already. In this case need to check, that this scalar is already part of the vectorized buildvector.	2022-05-09 13:07:59 -07:00
Alexey Bataev	9dc4ced204	[SLP]Try partial store vectorization if supported by target. We can try to vectorize number of stores less than MinVecRegSize / scalar_value_size, if it is allowed by target. Gives an extra opportunity for the vectorization. Fixes PR54985. Differential Revision: https://reviews.llvm.org/D124284	2022-05-09 09:48:15 -07:00
Alexey Bataev	9c3a75eabf	[SLP]Fix a crash when preparing a mask for external scalars. Need to use actual index instead of the tree entry position, since the insert index may be different than 0. It mean, that we vectorized part of the buildvector starting from not initial insertelement instruction beause of some reason.	2022-05-09 07:59:34 -07:00
Florian Hahn	41e142fdc7	Recommit "[SimpleLoopUnswitch] Collect either logical ANDs/ORs but not both." This reverts commit `7211d5ce07`. This version fixes a crash that caused buildbot failures with the first version.	2022-05-09 13:49:12 +01:00
Florian Hahn	4c569ceeaa	[SimpleLoopUnswitch] Add test case for crash with `db7a87ed4f`.	2022-05-09 13:48:56 +01:00
David Sherwood	45f2e92d97	[NFC][LoopVectorize] Add SVE test for tail-folding combined with interleaving Differential Revision: https://reviews.llvm.org/D125001	2022-05-09 13:08:25 +01:00
Florian Hahn	61bb2e4ea8	[ConstraintElimination] Add initial ssub.with.overflow tests.	2022-05-09 10:02:59 +01:00
Simon Pilgrim	9a12138b5f	[SLP][X86] Add test coverage for PR50392 / Issue #49736	2022-05-08 19:40:04 +01:00
Simon Pilgrim	751005a2ca	[SLP][X86] Add test coverage for PR42652 / Issue #41997	2022-05-08 12:09:14 +01:00
Simon Pilgrim	7d94597048	[SLP][X86] Add test coverage for PR41892 / Issue #41237	2022-05-08 11:40:53 +01:00
Simon Pilgrim	2233a61500	[SLP][X86] Add test coverage for PR49934 / Issue #49278 D124284 should help us vectorize the sub-128-bit vector cases	2022-05-08 11:33:01 +01:00
Simon Pilgrim	96d2d2508e	[SLP][X86] Add test coverage for PR47491 / Issue #46835 D124284 should help us vectorize the sub-128-bit vector cases	2022-05-08 11:24:46 +01:00
Simon Pilgrim	993d9462e1	[InstCombine] Add test coverage for PR43261 / Issue #42606	2022-05-08 11:10:49 +01:00
David Green	6f9e1ea0ef	[VectorCombine] Attempt to fold select shuffles from reductions Given a commutative reduction leading from a shuffle, the order of the lanes on the shuffle are not important for the result. This means we can reorder the shuffle to something simpler, which we try shuffling the first vector lanes first. This was D123494. The new shuffle may not be profitable though, and if it is not we can try the folding of select shuffles from D123911. This, with some adjustment as the output lane ordering is now unimportant, can allow the final shuffle to simplify given the inputs to the patterns from D123911. Where as each transformation on their own are not profitable, the combination is. We can only support a single shuffle when called from reductions, but we are able to sort the ReconstructMask, potentially allowing it to simplify to an identity or concat mask. Differential Revision: https://reviews.llvm.org/D125086	2022-05-08 10:32:41 +01:00
Andrew Litteken	e38f014c40	[IROutliner] Accomodate blocks containing PHINodes with one entry outside the region and others inside the region. When a PHINode has an incoming block from outside the region, it must be handled specially when assigning a global value number to each incoming value. A PHINode has multiple predecessors, and we must handle this case rather than only the single predecessor case. Reviewer: paquette Differential Revision: https://reviews.llvm.org/D124777	2022-05-07 17:11:21 -05:00
David Green	802e15c576	[SLP] Cluster ordering for loads Given a load without a better order, this patch partially sorts the elements to form clusters of adjacent elements in memory. These clusters can potentially be loaded in fewer loads, meaning less overall shuffling (for example loading v4i8 clusters of a v16i8 as a single f32 loads, as opposed to multiple independent bytes loads and inserts). Differential Revision: https://reviews.llvm.org/D122145	2022-05-07 14:38:11 +01:00
Sanjay Patel	8650f05c97	[InstCombine] fix miscompile when casting int->FP->int As shown in https://github.com/llvm/llvm-project/issues/55150 - the existing fold may be wrong when converting to a signed value. This is a quick fix to avoid the miscompile. I added tests/comments for all of the signed/unsigned combinations at either side of the boundary width, and tried to confirm with Alive2: https://alive2.llvm.org/ce/z/3p9DSu There are already some TODO items in the test file that suggest possible refinements, so the regression with ui->FP->si is probably ok. It seems unlikely that we'd see these kind of edge cases with non-byte-width integer types in real code. The potential miscompile went undetected for several years. This and `747c6a0c73` fixes #55150. Differential Revision: https://reviews.llvm.org/D124692	2022-05-07 08:46:25 -04:00
Serge Pavlov	eb28da89a6	[InstCombine] Remove side effect of replaced constrained intrinsics If a constrained intrinsic call was replaced by some value, it was not removed in some cases. The dangling instruction resulted in useless instructions executed in runtime. It happened because constrained intrinsics usually have side effect, it is used to model the interaction with floating-point environment. In some cases side effect is actually absent or can be ignored. This change adds specific treatment of constrained intrinsics so that their side effect can be removed if it actually absents. Differential Revision: https://reviews.llvm.org/D118426	2022-05-07 19:04:11 +07:00
David Green	2db46db54d	[SLP] Add tests for awkward laod orders from SLP. NFC	2022-05-07 10:27:32 +01:00
Chenbing Zheng	394c683d40	[InstCombine] sub(add(X,Y),umin(Y,Z)) --> add(X,usub.sat(Y,Z)) Alive2: https://alive2.llvm.org/ce/z/2UNVbp Reviewed By: RKSimon, spatel Differential Revision: https://reviews.llvm.org/D124503	2022-05-07 17:17:48 +08:00
Chenbing Zheng	1fd7929ae5	[InstCombine] precommit some tests for reassociate add	2022-05-07 15:52:28 +08:00
Chenbing Zheng	8eaa1ef0d8	[InstCombine] add casts from splat-a-bit pattern if necessary Splatting a bit of constant-index across a value: sext (ashr (trunc iN X to iM), M-1) to iN --> ashr (shl X, N-M), N-1 If the dest type is different, use a cast (adjust use check). https://alive2.llvm.org/ce/z/acAan3 Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D124590	2022-05-07 15:34:57 +08:00
Florian Hahn	7211d5ce07	Revert "[SimpleLoopUnswitch] Collect either logical ANDs/ORs but not both." This reverts commit `db7a87ed4f`. This seems to cause a PPC buildbot failure: https://lab.llvm.org/buildbot#builders/93/builds/8787	2022-05-06 22:38:15 +01:00
Sanjay Patel	b331a7ebc1	[InstCombine] canonicalize fneg after shuffle For the unary shuffle pattern, this is opposite to what we try to do with binops, but it seems better to keep it consistent with the motivating binary shuffle pattern. On that, it is clearly better on the usual no-extra uses case. There is a chance that this will pull an fneg away from some other binop and cause a regression in codegen, but that should be invertible in the backend. The transform is birectional: https://alive2.llvm.org/ce/z/kKaKCU https://alive2.llvm.org/ce/z/3Desfw Fixes #45631	2022-05-06 16:30:26 -04:00
Sanjay Patel	ef9d39de2f	[InstCombine] add tests for shuffle with fneg operand(s); NFC issue #45631	2022-05-06 16:30:26 -04:00
Florian Hahn	1d042312f8	[InstCombine] Add tests for combining AArch64 neon min/max intrinsics.	2022-05-06 17:59:23 +01:00
Martin Sebor	cc2ce81bd8	[SimplifyLibcalls] Tests for libcall folding of subobjects [NFC] Add tests exercising the future enancement of folding library function calls with arguments involving subobjects such as elements of arrays or struct members.	2022-05-06 10:43:02 -06:00
Nikita Popov	82190f917a	[InstCombine] Fold icmp of select with implied condition When threading the icmp over the select, check whether the condition can be folded when taking into account the select condition.	2022-05-06 17:13:32 +02:00
Nikita Popov	a94589d52f	[InstCombine] Add icmp of select with implied condition tests (NFC)	2022-05-06 17:13:32 +02:00
Nikita Popov	0863abe3ac	[InstCombine] Fold icmp of select with non-constant operand Try to push an icmp into a select even if the icmp operand isn't constant - perform a generic SimplifyICmpInst instead. This doesn't appear to impact compile-time much, and forming logical and/or is generally profitable, as we have very good support for them.	2022-05-06 16:04:39 +02:00
Nikita Popov	d7b6fd47b2	[InstCombine] Add additional icmp of select tests (NFC)	2022-05-06 16:04:39 +02:00
Simon Pilgrim	cbd300f62d	[SLP][X86] Add test coverage for Issue #51088	2022-05-06 14:49:47 +01:00
Max Kazantsev	5a08e81779	[RS4GC] Add support for 'freeze' instruction to findBaseDefiningValue Because this instruction is a noop, we can simply go through it in search of the base.	2022-05-06 20:46:29 +07:00
Fraser Cormack	bafab9c09f	[InstCombine] Fix scalable-vector bitwise select matching D113035 enhanced the matching of bitwise selects from vector types. This change unfortunately introduced crashes as it tries to cast scalable vector types to integers. Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D124997	2022-05-06 12:59:39 +01:00
Simon Pilgrim	cbfa857346	[CostModel][X86] Adjust 128-bit select costs to account for slow BLENDV op Based off the script from D103695 - Jaguar, Bulldozer, Silvermont (et al) and Haswell all have slow BLENDV ops, so adjust the worse case cost values	2022-05-06 13:07:34 +01:00
Simon Pilgrim	d21bf51494	[CostModel][X86] Adjust pre-SSE41 fp scalar select costs to account for vector ops Based off the script from D103695, we now mainly use BLENDV or OR(AND,ANDN) to select scalar float/double ops	2022-05-06 11:41:55 +01:00
Simon Pilgrim	f0e8c1d6d9	[CostModel][X86] Adjust 256-bit select costs to account for slow BLENDV op Based off the script from D103695, on AVX1, Jaguar/Bulldozer both have low throughput for ymm select patterns (BLENDV + OR(AND,ANDN))), and even on AVX2 Haswell still struggles with BLENDV ops	2022-05-06 11:27:37 +01:00
Simon Pilgrim	f22e993e4e	[SLP][X86] Regenerate ssat tests to remove defunct AVX1/AVX2 checks	2022-05-06 11:21:52 +01:00
Florian Hahn	db7a87ed4f	[SimpleLoopUnswitch] Collect either logical ANDs/ORs but not both. After D97756, collectHomogenousInstGraphLoopInvariants may collect conditions for both logical ANDs and logical ORs in case the root is a select that matches both logical AND & OR. This means the function won't return invariant values of either AND/OR chains, but both. This can result in incorrect transformations. See llvm/test/Transforms/SimpleLoopUnswitch/trivial-unswitch-logical-and-or.ll. Without the patch, Alive2 rejects the modified tests with: Source and target don't have the same return domain. Note that this also applies to the test case added in D97756 (@test_partial_condition_unswitch_or_select). We can't unswitch on %cond6, because the graph leading to it contains and AND and an OR. This only fixes trivial unswitching for now, but a similar problem likely exists with non-trivial unswitching. Reviewed By: nikic Differential Revision: https://reviews.llvm.org/D124526	2022-05-06 09:50:03 +01:00
David Green	100cb9a2ba	[VectorCombine] Fold shuffle select pattern This patch adds a combine to attempt to reduce the costs of certain select-shuffle patterns. The form of code it attempts to detect is: %x = shuffle ... %y = shuffle ... %a = binop %x, %y %b = binop %x, %y shuffle %a, %b, selectmask A classic select-mask will pick items from each lane of a or b. These do not always have a great lowering on many architectures. This patch attempts to pack a and b into the lower elements, creating a differently ordered shuffle for reconstructing the orignal which may be better than the select mask. This can be better for performance, especially if less elements of a and b need to be computed and the input shuffles are cheaper. Because select-masks are just one form of shuffle, we generalize to any mask. So long as the backend has decent costmodel for the shuffles, this can generally improve things when they come up. For more basic cost models the folds do not appear to be profitable, not getting past the cost checks. Differential Revision: https://reviews.llvm.org/D123911	2022-05-06 08:13:18 +01:00
Chenbing Zheng	53069de6aa	[InstCombine] precommit tests for D124590	2022-05-06 10:53:12 +08:00
Chenbing Zheng	4c8c101b49	[InstCombine] try to narrow more shifted bswap-of-zext Try to narrow more bswap, if the shift amount is less than the zext (bswap (zext X)) >> C --> (zext (bswap X)) << C' https://alive2.llvm.org/ce/z/i7ddjn Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D124598	2022-05-06 10:45:10 +08:00
Serge Pavlov	e1554ac63a	Revert "[InstCombine] Remove side effect of replaced constrained intrinsics" This reverts commit `83914ee96f`. The change caused discussion: https://lists.llvm.org/pipermail/llvm-commits/Week-of-Mon-20220502/1034841.html	2022-05-06 01:09:16 +07:00
Sanjay Patel	21c028ac94	[InstCombine] fix typo in test name; NFC	2022-05-05 12:47:11 -04:00
Sanjay Patel	7bad1d281c	[InstCombine] add scalable vector test for logical select; NFC D124997 shows that the code is not ready to handle scalable vectors, so add some more coverage for a potential crashing case.	2022-05-05 12:46:28 -04:00
Benjamin Kramer	08b20f20d2	[ConstantFold] Use getFltSemantics instead of manually checking the type Simplifies the code and makes fpext/fptrunc constant folding not crash when the result is bf16.	2022-05-05 15:52:19 +02:00
Alexey Bataev	99f31acfce	[SLP]Further improvement of the cost model for scalars used in buildvectors. Further improvement of the cost model for the scalars used in buildvectors sequences. The main functionality is outlined into a separate function. The cost is calculated in the following way: 1. If the Base vector is not undef vector, resizing the very first mask to have common VF and perform action for 2 input vectors (including non-undef Base). Other shuffle masks are combined with the resulting after the 1 stage and processed as a shuffle of 2 elements. 2. If the Base is undef vector and have only 1 shuffle mask, perform the action only for 1 vector with the given mask, if it is not the identity mask. 3. If > 2 masks are used, perform serie of shuffle actions for 2 vectors, combing the masks properly between the steps. The original implementation misses the very first analysis for the Base vector, so the cost might too optimistic in some cases. But it improves the cost for the insertelements which are part of the current SLP graph. Part of D107966. Differential Revision: https://reviews.llvm.org/D115750	2022-05-05 06:04:25 -07:00
Florian Hahn	3497a4f396	[LICM] Add test to exercise assertion from D123473. Add a test case that triggers an assertion with earlier versions of D123473.	2022-05-05 10:49:52 +01:00

1 2 3 4 5 ...

21823 Commits