llvm-project

Commit Graph

Author	SHA1	Message	Date
Craig Topper	4b89cc1b96	[X86] Autogenerate complete checks. NFC llvm-svn: 324984	2018-02-13 04:19:23 +00:00
Craig Topper	ef619918f2	[X86] Remove duplicate CHECK-LABEL line the update script didn't delete when I converted the test. llvm-svn: 324979	2018-02-13 01:36:27 +00:00
Craig Topper	52b5558ca0	[X86] Auto generate complete checks. NFC llvm-svn: 324964	2018-02-12 23:43:10 +00:00
Adam Nemet	031a00c660	Revert "[LSR] Avoid UB overflow when examining reuse opportunities" This reverts commit r324943. Breaking bots, reverting for Gerolf. llvm-svn: 324958	2018-02-12 22:42:13 +00:00
Craig Topper	8d19c6fba2	[X86] Reverse the operand order of the autoupgrade of the kunpack builtins. The second operand needs to be in the lower bits of the concatenation. This matches llvm 5.0, gcc, and icc behavior. Fixes PR36360. llvm-svn: 324953	2018-02-12 22:38:34 +00:00
Sanjay Patel	0537deb5cd	[x86] add select test to show there's no single right answer (PR28968); NFC llvm-svn: 324947	2018-02-12 22:19:24 +00:00
Gerolf Hoflehner	edcd564820	[LSR] Avoid UB overflow when examining reuse opportunities llvm-svn: 324943	2018-02-12 21:49:32 +00:00
Sanjay Patel	014c000f6a	[DAG] make binops with undef operands consistent with IR This started by noticing that scalar and vector types were producing different results with div ops in PR36305: https://bugs.llvm.org/show_bug.cgi?id=36305 ...but the problem is bigger. I couldn't keep it straight without a table, so I'm attaching that as a PDF to the review. The x86 tests in undef-ops.ll correspond to that table. Green means that instsimplify and the DAG agree on the result for all types. Red means the DAG was returning undef when IR was not. Yellow means the DAG was returning a non-undef result when IR returned undef. This patch assumes that we're currently doing the right thing in IR. Note: I couldn't find any problems with lowering vector constants as the code comments were warning, but those comments were written long ago in rL36413 . Differential Revision: https://reviews.llvm.org/D43141 llvm-svn: 324941	2018-02-12 21:37:27 +00:00
Simon Pilgrim	d0693a6501	[X86][MMX] Add missing scheduling class tag for EMMS/FEMMS We only tagged it with the itinerary class, so completeness checks were erroneously passed (PR35639). AMD targets can perform these a lot quicker than WriteMicrocoded so will need an override in the models. llvm-svn: 324897	2018-02-12 15:52:59 +00:00
Hans Wennborg	7e19dfc45f	Revert r324835 "[X86] Reduce Store Forward Block issues in HW" It asserts building Chromium; see PR36346. (This also reverts the follow-up r324836.) > If a load follows a store and reloads data that the store has written to memory, Intel microarchitectures can in many cases forward the data directly from the store to the load, This "store forwarding" saves cycles by enabling the load to directly obtain the data instead of accessing the data from cache or memory. > A "store forward block" occurs in cases that a store cannot be forwarded to the load. The most typical case of store forward block on Intel Core microarchiticutre that a small store cannot be forwarded to a large load. > The estimated penalty for a store forward block is ~13 cycles. > > This pass tries to recognize and handle cases where "store forward block" is created by the compiler when lowering memcpy calls to a sequence > of a load and a store. > > The pass currently only handles cases where memcpy is lowered to XMM/YMM registers, it tries to break the memcpy into smaller copies. > breaking the memcpy should be possible since there is no atomicity guarantee for loads and stores to XMM/YMM. llvm-svn: 324887	2018-02-12 12:43:39 +00:00
Craig Topper	98ae8f833f	[X86] Change some compare patterns to use loadi8/loadi16/loadi32/loadi64 helper fragments. This enables CMP8mi to fold zextloadi8i1 which in all tests allows us to avoid creating a TEST8rr that peephole can't fold. llvm-svn: 324863	2018-02-12 02:48:42 +00:00
Craig Topper	27d5b6e4a6	[X86] Autogenerate complete checks. NFC llvm-svn: 324862	2018-02-12 02:03:36 +00:00
Craig Topper	dfc322ddf4	[X86] Allow zextload/extload i1->i8 to be folded into instructions during isel Previously we just emitted this as a MOV8rm which would likely get folded during the peephole pass anyway. This just makes it explicit earlier. The gpr-to-mask.ll test changed because the kaddb instruction has no memory form. llvm-svn: 324860	2018-02-12 01:33:36 +00:00
Craig Topper	3a354152dd	[X86] Update some required-vector-width.ll test cases to not pass 512-bit vectors in arguments or return. ABI for these would require 512 bits support so we don't want to test that. llvm-svn: 324845	2018-02-11 18:52:16 +00:00
Simon Pilgrim	0d8c4bfc2a	[X86][SSE] Use SplitBinaryOpsAndApply to recognise PSUBUS patterns before they're split on AVX1 This needs to be generalised further to support AVX512BW cases but I want to add non-uniform constants first. llvm-svn: 324844	2018-02-11 17:29:42 +00:00
Craig Topper	ca5a340171	[X86] Use min/max for vector ult/ugt compares if avoids a sign flip. Summary: Currently we only use min/max to help with ule/uge compares because it removes an invert of the result that would otherwise be needed. But we can also use it for ult/ugt compares if it will prevent the need for a sign bit flip needed to use pcmpgt at the cost of requiring an invert after the compare. I also refactored the code so that the max/min code is self contained and does its own return instead of setting up a flag to manipulate the rest of the function's behavior. Most of the test cases look ok with this. I did notice that we added instructions when one of the operands being sign flipped is a constant vector that we were able to constant fold the flip into. I also noticed that sometimes the SSE min/max clobbers a register that is needed after the compare. This resulted in an extra move being inserted before the min/max to preserve the register. We could try to detect this and switch from min to max and change the compare operands to use the operand that gets reused in the compare. Reviewers: spatel, RKSimon Reviewed By: RKSimon Subscribers: llvm-commits Differential Revision: https://reviews.llvm.org/D42935 llvm-svn: 324842	2018-02-11 17:11:40 +00:00
Sanjay Patel	eb8c408e50	[TargetLowering] try to create -1 constant operand for math ops via demanded bits This reverses instcombine's demanded bits' transform which always tries to clear bits in constants. As noted in PR35792 and shown in the test diffs: https://bugs.llvm.org/show_bug.cgi?id=35792 ...we can do better in codegen by trying to form -1. The x86 sub test shows a missed opportunity. I did investigate changing instcombine's behavior, but it would be more work to change canonicalization in IR. Clearing bits / shrinking constants can allow killing instructions, so we'd have to figure out how to not regress those cases. Differential Revision: https://reviews.llvm.org/D42986 llvm-svn: 324839	2018-02-11 14:38:23 +00:00
Simon Pilgrim	7630150222	[X86] Add PR33747 test case llvm-svn: 324838	2018-02-11 13:12:50 +00:00
Simon Pilgrim	0be5567a89	[X86][SSE] Enable SMIN/SMAX/UMIN/UMAX custom lowering for all legal types This allows us to recognise more saturation patterns and also simplify some MINMAX codegen that was failing to combine CMPGE comparisons to a legal CMPGT. Differential Revision: https://reviews.llvm.org/D43014 llvm-svn: 324837	2018-02-11 10:52:37 +00:00
Lama Saba	91e2b9d081	fix test/CodeGen/X86/fixup-sfb.ll test failure after commit https://reviews.llvm.org/rL324835 Change-Id: I2526c2f342654e85ce054237de03ae9db9ab4994 llvm-svn: 324836	2018-02-11 10:33:06 +00:00
Lama Saba	c2ba6c387e	[X86] Reduce Store Forward Block issues in HW If a load follows a store and reloads data that the store has written to memory, Intel microarchitectures can in many cases forward the data directly from the store to the load, This "store forwarding" saves cycles by enabling the load to directly obtain the data instead of accessing the data from cache or memory. A "store forward block" occurs in cases that a store cannot be forwarded to the load. The most typical case of store forward block on Intel Core microarchiticutre that a small store cannot be forwarded to a large load. The estimated penalty for a store forward block is ~13 cycles. This pass tries to recognize and handle cases where "store forward block" is created by the compiler when lowering memcpy calls to a sequence of a load and a store. The pass currently only handles cases where memcpy is lowered to XMM/YMM registers, it tries to break the memcpy into smaller copies. breaking the memcpy should be possible since there is no atomicity guarantee for loads and stores to XMM/YMM. Change-Id: I620b6dc91583ad9a1444591e3ddc00dd25d81748 llvm-svn: 324835	2018-02-11 09:34:12 +00:00
Craig Topper	24d3b28d93	[X86] Don't make 512-bit vectors legal when preferred vector width is 256 bits and 512 bits aren't required This patch adds a new function attribute "required-vector-width" that can be set by the frontend to indicate the maximum vector width present in the original source code. The idea is that this would be set based on ABI requirements, intrinsics or explicit vector types being used, maybe simd pragmas, etc. The backend will then use this information to determine if its save to make 512-bit vectors illegal when the preference is for 256-bit vectors. For code that has no vectors in it originally and only get vectors through the loop and slp vectorizers this allows us to generate code largely similar to our AVX2 only output while still enabling AVX512 features like mask registers and gather/scatter. The loop vectorizer doesn't always obey TTI and will create oversized vectors with the expectation the backend will legalize it. In order to avoid changing the vectorizer and potentially harm our AVX2 codegen this patch tries to make the legalizer behavior similar. This is restricted to CPUs that support AVX512F and AVX512VL so that we have good fallback options to use 128 and 256-bit vectors and still get masking. I've qualified every place I could find in X86ISelLowering.cpp and added tests cases for many of them with 2 different values for the attribute to see the codegen differences. We still need to do frontend work for the attribute and teach the inliner how to merge it, etc. But this gets the codegen layer ready for it. Differential Revision: https://reviews.llvm.org/D42724 llvm-svn: 324834	2018-02-11 08:06:27 +00:00
Simon Pilgrim	d229bfd20d	[X86][SSE] Add SMIN/SMAX combine test As discussed on D43014, we need the ability to flip SMIN/SMAX to (legal) UMIN/UMAX llvm-svn: 324829	2018-02-10 23:38:50 +00:00
Craig Topper	4dccffc84a	[X86] Change signatures of avx512 packed fp compare intrinsics to return a vXi1 mask type to be closer to an fcmp. Summary: This patch changes the signature of the avx512 packed fp compare intrinsics to return a vXi1 vector and no longer take a mask as input. The casts to scalar type will now need to be explicit in the IR. The masking node will now be an explicit and in the IR. This makes the intrinsic look much more similar to an fcmp instruction that we wish we could use for these but can't. We already use icmp instructions for integer compares. Previously the lowering step of isel would turn the intrinsic into an X86 specific ISD node and a emit the masking nodes as well as some bitcasts. This means DAG combines can't see the vXi1 type until somewhat late, making it more difficult to combine out gpr<->mask transition sequences. By exposing the vXi1 type explicitly in the IR and initial SelectionDAG we give earlier DAG combines and even InstCombine the chance to see it and optimize it. This should make any issues with gpr<->mask sequences the same between integer and fp. Meaning we only have to fix them once. Reviewers: spatel, delena, RKSimon, zvi Reviewed By: RKSimon Subscribers: llvm-commits Differential Revision: https://reviews.llvm.org/D43137 llvm-svn: 324827	2018-02-10 23:33:55 +00:00
Simon Pilgrim	b781b2e24c	[X86][SSE] Add UMIN/UMAX combine test As discussed on D43014, we need the ability to flip UMIN/UMAX to (legal) SMIN/SMAX llvm-svn: 324826	2018-02-10 22:27:35 +00:00
Craig Topper	9121eb575e	[X86] Custom legalize (v2i32 (setcc (v2f32))) so that we don't end up with a (v4i1 (setcc (v4f32))) Undef VLX, getSetCCResultType returns v2i1/v4i1 for v2f32/v4f32 so default type legalization will end up changing the setcc result type back to vXi1 if it had been extended. The resulting extend gets messed up further by type legalization and is difficult to recombine back to (v4i32 (setcc (v4f32))) after legalization. I went ahead and enabled this for SSE2 and later since its always the result we want and this helps type legalization get there in less steps. llvm-svn: 324822	2018-02-10 19:12:58 +00:00
Craig Topper	28d3a73c81	[X86] Extend inputs with elements smaller than i32 to sint_to_fp/uint_to_fp before type legalization. This prevents extends of masks being introduced during lowering where it become difficult to combine them out. There are a few oddities in here. We sometimes concatenate two k-registers produced by two compares, sign_extend the combined pair, then extract two halves. This worked better previously because the sign_extend wasn't created until after the fp_to_sint was split which led to a split sign_extend being created. We probably also need to custom type legalize (v2i32 (sext v2i1)) via widening. llvm-svn: 324820	2018-02-10 17:58:58 +00:00
Craig Topper	0fbdaa6d81	[X86] Remove some check-prefixes from avx512-cvt.ll to prepare for an upcoming patch. The update script sometimes has trouble when there are check-prefixes representing every possible combination of feature flags. I have a patch where the update script was generating something that didn't pass lit. This patch just removes some check-prefixes and expands out some of the checks to workaround this. llvm-svn: 324819	2018-02-10 17:58:56 +00:00
Sanjay Patel	3a7e447cd6	[x86] preserve test intent by removing undef D43141 proposes to correct undef folding in the DAG, and this test would not survive that change. llvm-svn: 324817	2018-02-10 15:36:23 +00:00
Sanjay Patel	5d176da09c	[x86] preserve test intent by removing undef D43141 proposes to correct undef folding in the DAG, and this test would not survive that change. llvm-svn: 324816	2018-02-10 15:28:08 +00:00
Simon Pilgrim	17b00ec4f5	[X86][SSE] Regenerate old sitofp v2i32 test llvm-svn: 324812	2018-02-10 14:45:58 +00:00
Craig Topper	b8d7b1620b	[X86] Custom legalize (v2i1 (fp_to_uint/fp_to_sint v2f64)) without AVX512VL. Strangely the code was already present, just the setOperationAction wasn't being called without VLX. llvm-svn: 324806	2018-02-10 08:39:31 +00:00
Craig Topper	c3aab4bbe1	[X86] Legalize zero extends from vXi1 to vXi16/vXi32/vXi64 using a sign extend and a shift. This avoids a constant pool load to create 1. The int->float are showing converts to mask and back. We probably need to widen inputs to sint_to_fp/uint_to_fp before type legalization. llvm-svn: 324805	2018-02-10 08:06:52 +00:00
Craig Topper	d34af6f636	[X86] Teach combineExtSetcc to handle ZERO_EXTEND by widening the setcc and then masking. A later DAG combine will convert to a shift. This helps to avoid a constant pool load needed to zero extend from the mask. llvm-svn: 324804	2018-02-10 08:06:49 +00:00
Craig Topper	fa6113b3d7	[X86] Teach combineInsertSubvector how to combine some k-register insert_subvectors and extract_subvector sequences to remove extra zeroing.wq llvm-svn: 324791	2018-02-10 01:00:41 +00:00
Craig Topper	99db883d55	[X86] Teach lower1BitVectorShuffle to recognize shuffles that are just filling upper elements with zero. Replace with insert_subvector. There's still some extra kshifts in one of the modified test cases here, but hopefully that's only a DAG combine away. llvm-svn: 324782	2018-02-09 23:32:27 +00:00
Sanjay Patel	4031ce15b8	[x86] remove duplicate undef tests; NFC These are incomplete and were made redundant with the consolidation in: https://reviews.llvm.org/rL324678 llvm-svn: 324754	2018-02-09 17:46:38 +00:00
Rafael Espindola	c052fa0bd3	Emit smaller exception tables for non-SJLJ mode. * Use uleb128 for code offsets in the LSDA call site table. * Omit the TTBase offset if the type table is empty. This change can reduce the size of the DWARF/Itanium LSDA by about half. Patch by Ryan Prichard! llvm-svn: 324750	2018-02-09 17:13:37 +00:00
Rafael Espindola	d09b416943	Use assembler expressions to lay out the EH LSDA. Rely on the assembler to finalize the layout of the DWARF/Itanium exception-handling LSDA. Rather than calculate the exact size of each thing in the LSDA, use assembler directives: To emit the offset to the TTBase label: .uleb128 .Lttbase0-.Lttbaseref0 .Lttbaseref0: To emit the size of the call site table: .uleb128 .Lcst_end0-.Lcst_begin0 .Lcst_begin0: ... call site table entries ... .Lcst_end0: To align the type info table: ... action table ... .balign 4 .long _ZTIi .long _ZTIl .Lttbase0: Using assembler directives simplifies the compiler and allows switching the encoding of offsets in the call site table from udata4 to uleb128 for a large code size savings. (This commit does not change the encoding.) The combination of the uleb128 followed by a balign creates an unfortunate dependency cycle that the assembler must sometimes resolve either by padding an LEB or by inserting zero padding before the type table. See PR35809 or GNU as bug 4029. Patch by Ryan Prichard! llvm-svn: 324749	2018-02-09 17:00:25 +00:00
Craig Topper	ca5841b4e4	[X86] Simplify some code in lowerV4X128VectorShuffle and lowerV2X128VectorShuffle Previously we extracted two subvectors and concatenate. But the concatenate will be lowered to two insert subvectors. Then DAG combine will merge once of the inserts and one of the extracts back into the original vector. We might as well just directly use one extract and one insert. llvm-svn: 324710	2018-02-09 05:54:36 +00:00
Craig Topper	28166a877d	[X86] Teach shuffle lowering to recognize 128/256 bit insertions into a zero vector. This regresses a couple cases in the shuffle combining test. But those cases use intrinsics that InstCombine knows how to turn into a generic shuffle earlier. This should give opportunities to fold this earlier in InstCombine or DAG combine. llvm-svn: 324709	2018-02-09 05:54:34 +00:00
Craig Topper	090e41d0cc	[X86] Add 512-bit shuffle test cases for concatenating 128/256-bits with zeros in the upper portion. We should recognize this and just use a mov that will zero the upper bits. llvm-svn: 324708	2018-02-09 05:54:31 +00:00
Craig Topper	79c3255fe4	[x86] Add test cases to demonstrate some dumb mask->gpr->mask transition sequences. llvm-svn: 324693	2018-02-09 01:14:17 +00:00
Francis Visoiu Mistrih	39ec2e95ae	[CodeGen] Unify the syntax of MBB successors in MIR and -debug output Instead of: Successors according to CFG: %bb.6(0x12492492 / 0x80000000 = 14.29%) print: successors: %bb.6(0x12492492); %bb.6(14.29%) llvm-svn: 324685	2018-02-09 00:10:31 +00:00
Sanjay Patel	b7e13938a9	[x86] consolidate and add tests for undef binop folds; NFC As was already shown in the div/rem tests and noted in PR36305, the behavior is inconsistent, but it's not limited to div/rem only. llvm-svn: 324678	2018-02-08 23:21:44 +00:00
Alexander Ivchenko	da9e81c462	[GlobalISel][X86] Fixing failures after https://reviews.llvm.org/D37775 The patch essentially makes sure that X86CallLowering adds proper G_COPY/G_TRUNC and G_ANYEXT/G_COPY when we are doing lowering of arguments/returns for floating point values passed on registers. Tests are updated accordingly Reviewed By: qcolombet Differential Revision: https://reviews.llvm.org/D42287 llvm-svn: 324665	2018-02-08 22:41:47 +00:00
Alexander Ivchenko	a85c4fc029	[GlobalIsel][X86] Making {G_IMPLICIT_DEF, s128} legal The patch is a split from D42287 and is related to fixing failures after https://reviews.llvm.org/D37775 Reviewed By: qcolombet Differential Revision: https://reviews.llvm.org/D42287 llvm-svn: 324664	2018-02-08 22:40:31 +00:00
Craig Topper	9e030c9e00	[X86] Improve combineCastedMaskArithmetic to fold (bitcast (vXi1 (and/or/xor X, C)))->(vXi1 (and/or/xor (bitcast X), (bitcast C)) where C is a constant build_vector. Most vxi1 constant build vectors have to be implemented in the scalar domain anyway so we'll probably end up with a cast there later. But by then its too late to do the combine to get rid of it. llvm-svn: 324662	2018-02-08 22:26:39 +00:00
Craig Topper	1b5b4ccb77	[X86] Add DAG combine to constant fold a bitcast of a vXi1 constant build_vector into a scalar integer. llvm-svn: 324661	2018-02-08 22:26:36 +00:00
Craig Topper	dccf72b583	[X86] Remove kortest intrinsics and replace with native IR. llvm-svn: 324646	2018-02-08 20:16:06 +00:00

1 2 3 4 5 ...

11279 Commits