llvm-project

Commit Graph

Author	SHA1	Message	Date
Peter Smith	e63455d5e0	[MC] Use local MCSubtargetInfo in writeNops On some architectures such as Arm and X86 the encoding for a nop may change depending on the subtarget in operation at the time of encoding. This change replaces the per module MCSubtargetInfo retained by the targets AsmBackend in favour of passing through the local MCSubtargetInfo in operation at the time. On Arm using the architectural NOP instruction can have a performance benefit on some implementations. For Arm I've deleted the copy of the AsmBackend's MCSubtargetInfo to limit the chances of this causing problems in the future. I've not done this for other targets such as X86 as there is more frequent use of the MCSubtargetInfo and it looks to be for stable properties that we would not expect to vary per function. This change required threading STI through MCNopsFragment and MCBoundaryAlignFragment. I've attempted to take into account the in tree experimental backends. Differential Revision: https://reviews.llvm.org/D45962	2021-09-07 15:46:19 +01:00
Peter Smith	5e71839f77	[MC] Add MCSubtargetInfo to MCAlignFragment In preparation for passing the MCSubtargetInfo (STI) through to writeNops so that it can use the STI in operation at the time, we need to record the STI in operation when a MCAlignFragment may write nops as padding. The STI is currently unused, a further patch will pass it through to writeNops. There are many places that can create an MCAlignFragment, in most cases we can find out the STI in operation at the time. In a few places this isn't possible as we are in initialisation or finalisation, or are emitting constant pools. When possible I've tried to find the most appropriate existing fragment to obtain the STI from, when none is available use the per module STI. For constant pools we don't actually need to use EmitCodeAlign as the constant pools are data anyway so falling through into it via an executable NOP is no better than falling through into data padding. This is a prerequisite for D45962 which uses the STI to emit the appropriate NOP for the STI. Which can differ per fragment. Note that involves an interface change to InitSections. It is now called initSections and requires a SubtargetInfo as a parameter. Differential Revision: https://reviews.llvm.org/D45961	2021-09-07 15:46:19 +01:00
Michael Liao	640beb38e7	[amdgpu] Enable selection of `s_cselect_b64`. Differential Revision: https://reviews.llvm.org/D109159	2021-09-07 10:45:07 -04:00
Mirko Brkusanin	6c4b634da6	[AMDGPU][GlobalISel] Legalize G_MUL for non-standard types Legalizing G_MUL for non-standard types (like i33) generated an error. Putting minScalar and maxScalar instead of clampScalar. Also using new rule, instead of widening to the next power of 2, widen to the next multiple of the passed argument (32 in this case), so instead of widening i65 to i128, we widen it to i96. Patch by: Mateja Marjanovic Differential Revision: https://reviews.llvm.org/D109228	2021-09-07 16:33:24 +02:00
Mirko Brkusanin	5263bf583a	[AMDGPU][GlobalISel] Legalization of G_ROTL and G_ROTR Add implementation for the legalization of G_ROTL and G_ROTR machine instructions. They are very similar to funnel shift instructions, the only difference is funnel shifts have 3 operands, whereas rotate instructions have two operands, the first being the register that is being rotated and the second being the number of shifts. The legalization of G_ROTL/G_ROTR is just lowering them into funnel shift instructions if they are legal. Patch by: Mateja Marjanovic Differential Revision: https://reviews.llvm.org/D105347	2021-09-07 16:33:24 +02:00
Simon Pilgrim	0d48ee2774	[X86] X86InstrSSE.td - remove unused template parameters. NFC. Identified in D109359	2021-09-07 15:13:05 +01:00
Simon Pilgrim	b50a60c234	[X86] X86InstrVecCompiler.td - remove unused template parameters. NFC. Identified in D109359	2021-09-07 14:46:08 +01:00
Simon Pilgrim	fb38795062	[X86] X86InstrFMA.td - remove unused template parameters. NFC. Identified in D109359	2021-09-07 14:46:07 +01:00
Sander de Smalen	448d47f743	[AArch64][SVE] Implement all-inactive predicate with PFALSE. Instead of using a WHILE XZR, XZR instruction, just emit a PFALSE. Reviewed By: david-arm Differential Revision: https://reviews.llvm.org/D109311	2021-09-07 14:29:02 +01:00
Mirko Brkusanin	36527cbe02	[AMDGPU][GlobalISel] Legalize memcpy family of intrinsics Legalize G_MEMCPY, G_MEMMOVE, G_MEMSET and G_MEMCPY_INLINE. Corresponding intrinsics are replaced by a loop that uses loads/stores in AMDGPULowerIntrinsics pass unless their length is a constant lower then MemIntrinsicExpandSizeThresholdOpt (default 1024). Any G_MEM* instruction that reaches legalizer should have a const length argument and should be expanded into appropriate number of loads + stores. Differential Revision: https://reviews.llvm.org/D108357	2021-09-07 12:24:07 +02:00
Fraser Cormack	a823bdf3ab	[RISCV][VP] Custom lower VP_STORE and VP_LOAD This patch adds support for the vector-predicated `VP_STORE` and `VP_LOAD` nodes. We do this in the same way we lower `MSTORE` and `MLOAD`: to regular load/store instructions via intrinsics. One necessary change was made to `SelectionDAGLegalize` so that `VP_STORE` nodes' operation actions are taken from the stored "value" operands, in the same vein as `STORE` or `MSTORE`. Reviewed By: craig.topper, rogfer01 Differential Revision: https://reviews.llvm.org/D108999	2021-09-07 10:53:25 +01:00
Fraser Cormack	f4dee8cb82	[RISCV][VP] Custom lower VP_SCATTER and VP_GATHER This patch adds support for the `VP_SCATTER` and `VP_GATHER` nodes by lowering them to RVV's `vsox`/`vlux` instructions, respectively. This process is almost identical to the existing `MSCATTER`/`MGATHER` support. One extra change was made to `SelectionDAGLegalize` so that `VP_SCATTER`'s operation action is derived from its stored "value" operand rather than its return type (which is always the chain). Reviewed By: craig.topper, rogfer01 Differential Revision: https://reviews.llvm.org/D108987	2021-09-07 10:43:07 +01:00
Andrew Wei	da9ed3dc71	[AArch64] Avoid adding duplicate implicit operands when expanding pseudo insts. When expanding pseudo insts, in order to create a new machine instr, we use BuildMI, which will add implicit operands by default. And transferImpOps will also copy implicit operands from old ones. Finally, duplicate implicit operands are added to the same inst. Sometimes this can cause correctness issues. Like below inst, renamable $w18 = nsw SUBSWrr renamable $w30, renamable $w14, implicit-def dead $nzcv After expanding, it will become $w18 = SUBSWrs renamable $w13, renamable $w14, 0, implicit-def $nzcv, implicit-def dead $nzcv A redundant implicit-def $nzcv is added, but the dead flag is missing. Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D109069	2021-09-07 17:11:58 +08:00
Ben Shi	63ca9371c7	[ARM] Implement target hook function to decide folding (mul (add x, c1), c2) Prevent the folding in DAGCombine if it leads to worse code. Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D109124	2021-09-07 15:42:43 +08:00
Craig Topper	da3ef8b756	[X86] Handle inverted inputs when matching VPTERNLOG from 2 binary ops. This is a more general version of D109273. Though it doesn't peek through bitcasts or rearange broadcasts. Reviewed By: LuoYuanke Differential Revision: https://reviews.llvm.org/D109295	2021-09-06 17:44:52 -07:00
Fangrui Song	76529b4468	[X86] Simplify condition guarding emitCalleeSavedFrameMoves. NFC	2021-09-06 15:54:02 -07:00
Fangrui Song	4f1e410a1b	[X86] Simplify two hasFP(F). NFC	2021-09-06 15:47:40 -07:00
Victor Campos	79f9c79aaf	[AArch64][MC] Merge FeaturePMU into FeaturePerfMon FeaturePMU was created in AArch64 to accommodate one missing system register, PMMIR_EL1, in commit `ffcd7698ae`. However, the Performance Monitors extension already had a target feature, which is called FeaturePerfMon. Therefore, FeaturePMU is redundant. This patch removes FeaturePMU and merges its contents into FeaturePerfMon. Reviewed By: dnsampaio Differential Revision: https://reviews.llvm.org/D109246	2021-09-06 14:56:49 +01:00
David Truby	b297531ece	[AArch64][sve] Prevent incorrect function call on fixed width vector The isEssentiallyExtractHighSubvector function currently calls getVectorNumElements on a type that in specific cases might be scalable. Since this function only has correct behaviour at the moment on scalable types anyway, the function can just return false when given a fixed type. Differential Revision: https://reviews.llvm.org/D109163	2021-09-06 14:25:03 +01:00
Tianqing Wang	12fa608af4	[X86] Add CRC32 feature. `d8faf03807` implemented general-regs-only for X86 by disabling all features with vector instructions. But the CRC32 instruction in SSE4.2 ISA, which uses only GPRs, also becomes unavailable. This patch adds a CRC32 feature for this instruction and allows it to be used with general-regs-only. Reviewed By: pengfei Differential Revision: https://reviews.llvm.org/D105462	2021-09-06 17:24:30 +08:00
Fangrui Song	0e03450ae4	[AArch64] Remove an uneeded !NeedsWinCFI check. NFC	2021-09-05 21:02:56 -07:00
guopeilin	5f48c144c5	[AArch64][GlobalISel] Use ZExtValue for zext(xor) when invert tb(n)z Currently, we use SExtValue to decide whether to invert tbz or tbnz. However, for the case zext (xor x, c), we should use ZExt rather than SExt otherwise we will generate totally opposite branches. Reviewed By: paquette Differential Revision: https://reviews.llvm.org/D108755	2021-09-06 11:12:07 +08:00
Simon Pilgrim	f114ef3731	[CostModel][X86] Add generic costs for vXi32 MUL -> v2Xi16 PMADDDW folds Based off the improved fold in D108522 This should eventually allow us to replace the SLM only cost patterns with generic versions.	2021-09-05 16:08:11 +01:00
Shivam Gupta	5449d2da65	[NFC] Run clang-format on llvm/lib/Trget/AVR/ The current inconsistency confuse contributors which coding guidlines to follow. It would be better to have it consistent using clang-format tool. Reviewed By: mhjacobson Differential Revision: https://reviews.llvm.org/D109270	2021-09-04 20:05:15 +05:30
Simon Pilgrim	2005ae15a6	[X86][SLM] WriteVecIMul instructions only take 1uop (REAPPLIED) The xmm variant have half the throughput (and +1cy latency) of the mmx variants, but are still 1uop. I still need to do more thorough testing of SLM on test-suite before fixing the obvious bad numbers for WritePMULLD. But this helps the D103695 helper script get to more accurate numbers for vXi32 multiplies of extended operands (i.e. we can use PMADDWD, PMULLW/PMULHW etc). Matches what Intel AoM / Agner / llvm-exegesis reports.	2021-09-04 15:03:56 +01:00
Simon Pilgrim	ac51d69208	Revert rG994da657076900f5ad7fe593c3b5e5f89ab3d53d "[X86][SLM] WriteVecIMul instructions only take 1uop" This changed some codegen tests that I forgot about in my rebase, I'll recommit shortly with a fix.	2021-09-04 13:39:10 +01:00
Simon Pilgrim	994da65707	[X86][SLM] WriteVecIMul instructions only take 1uop The xmm variant have half the throughput (and +1cy latency) of the mmx variants, but are still 1uop. I still need to do more thorough testing of SLM on test-suite before fixing the obvious bad numbers for WritePMULLD. But this helps the D103695 helper script get to more accurate numbers for vXi32 multiplies of extended operands (i.e. we can use PMADDWD, PMULLW/PMULHW etc). Matches what Intel AoM / Agner / llvm-exegesis reports.	2021-09-04 13:21:34 +01:00
Simon Pilgrim	c6371020a8	[X86][SLM] RMW instructions don't require an extra uop For RMW instructions, the load and store hold the MEC for an extra cycle, but within the same single uop. This is alluded to in the Intel AOM: "The MEC also owns the MEC RSV, which is responsible for scheduling of all loads and stores. Load and store instructions go through addresses generation phase in program order to avoid on-the-fly memory ordering later in the pipeline. Therefore, an unknown address will stall younger memory instructions." Noticed while trying to get a cheap SLM test box up and running with llvm-exegesis - RMW arithmetic is always 1uop - and matches what Agner / InstLatX64 report as well.	2021-09-04 13:21:34 +01:00
Simon Pilgrim	da965a77d5	[X86][SLM] Fix MUL uops, latency and throughput These were all set to the same best case mul i32 values (which seems to be the only version of MUL that SLM actually performs well with). Noticed while trying to improve multiplication costs for vectorization via the D103695 helper script. Confirmed with Intel AoM / Agner / InstLatX64.	2021-09-04 13:21:34 +01:00
Simon Pilgrim	7d062d2c47	[X86][Atom] MUL/DIV instructions require both ports, not either. Noticed while trying to improve multiplication costs for vectorization via the D103695 helper script. Confirmed with Intel AoM.	2021-09-04 11:58:09 +01:00
Simon Pilgrim	0d0f39b0f3	[X86][Atom] Add missing UOps override to AtomWriteResPair multiclass Make it easier to describe microcoded instructions.	2021-09-04 11:58:09 +01:00
Nikita Popov	66a54af967	[WebAssembly] Support opaque pointers in AddMissingPrototypes The change here is basically the same as in D108880: Rather than looking at bitcasts, look at calls and their function type. We still need to look through bitcasts to find those calls. The change in llvm/test/CodeGen/WebAssembly/add-prototypes-conflict.ll is due to different visitation order. add-prototypes-opaque-ptrs.ll is a copy of add-prototypes.ll with -force-opaque-pointers. Differential Revision: https://reviews.llvm.org/D109256	2021-09-04 11:25:42 +02:00
Kevin Athey	c7f50a445e	Revert "[AArch64] Implement target hook function to decide folding (mul (add x, c1), c2)" This reverts commit `095bea23d0`. Broke buildbot: https://lab.llvm.org/buildbot/#/builders/5/builds/11411	2021-09-03 18:08:58 -07:00
Ben Shi	095bea23d0	[AArch64] Implement target hook function to decide folding (mul (add x, c1), c2) Prevent the folding if it leads to worse code. Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D108871	2021-09-04 07:24:23 +08:00
Stanislav Mekhanoshin	d0c064715c	[AMDGPU] Small cleanup in optimizeCompareInstr. NFC.	2021-09-03 11:31:40 -07:00
David Green	adfd12e6d1	[ARM] Add patterns for store(fptosisat(..)) As an extension to D107866, this adds store(fptosisat(..)) patterns, similar to the existing fptosi patterns, to prevent unnecessarily moving into gpr regs where we can use fp stores directly. Differential Revision: https://reviews.llvm.org/D108378	2021-09-03 19:22:11 +01:00
David Green	f37e132263	[ARM] Add VFP lowering for fptosi.sat This extends D107865 to the VFP insructions, lowering llvm.fptosi.sat and llvm.fptoui.sat to VCVT instructions that inherently perform the saturate. Differential Revision: https://reviews.llvm.org/D107866	2021-09-03 18:11:08 +01:00
Craig Topper	75620fadf5	[RISCV] Change how we encode AVL operands in vector pseudoinstructions to use GPRNoX0. This patch changes the register class to avoid accidentally setting the AVL operand to X0 through MachineIR optimizations. There are cases where we really want to use X0, but we can't get that past the MachineVerifier with the register class as GPRNoX0. So I've use a 64-bit -1 as a sentinel for X0. All other immediate values should be uimm5. I convert it to X0 at the earliest possible point in the VSETVLI insertion pass to avoid touching the rest of the algorithm. In SelectionDAG lowering I'm using a -1 TargetConstant to hide it from instruction selection and treat it differently than if the user used -1. A user -1 should be selected to a register since it doesn't fit in uimm5. This is the rest of the changes started in D109110. As mentioned there, I don't have a failing test from MachineIR optimizations anymore. Reviewed By: frasercrmck Differential Revision: https://reviews.llvm.org/D109116	2021-09-03 09:19:25 -07:00
Simon Pilgrim	6ba0b9f68a	[X86][SLM] Fix PBLENDVB uops and throughput SLM PBLENDVB is just as bad as BLENDVPD/PS - so model it as such, fixing the rr vs rm uops diff as well. The Intel AoM appears to have a copy+paste typo with PBLENDW, it doesn't match Agner or InstLatX64. Noticed while investigating some of the weird discrepancies reported by the D103695 helper script (SLM had much better vector shift throughputs than it should).	2021-09-03 11:31:29 +01:00
Florian Mayer	abf8ed8a82	[hwasan] Support more complicated lifetimes. This is important as with exceptions enabled, non-POD allocas often have two lifetime ends: the exception handler, and the normal one. Reviewed By: eugenis Differential Revision: https://reviews.llvm.org/D108365	2021-09-03 10:29:50 +01:00
Cullen Rhodes	dc5dd77ac7	[AArch64][SME] Support NEON vector to GPR integer moves in streaming mode A small subset of the NEON instruction set is legal in streaming mode. This patch adds support for the following vector to integer move instructions: 0x00 1110 0000 0001 0010 11xx xxxx xxxx # SMOV W\|Xd,Vn.B[0] 0x00 1110 0000 0010 0010 11xx xxxx xxxx # SMOV W\|Xd,Vn.H[0] 0100 1110 0000 0100 0010 11xx xxxx xxxx # SMOV Xd,Vn.S[0] 0000 1110 0000 0001 0011 11xx xxxx xxxx # UMOV Wd,Vn.B[0] 0000 1110 0000 0010 0011 11xx xxxx xxxx # UMOV Wd,Vn.H[0] 0000 1110 0000 0100 0011 11xx xxxx xxxx # UMOV Wd,Vn.S[0] 0100 1110 0000 1000 0011 11xx xxxx xxxx # UMOV Xd,Vn.D[0] Only the zero index variants are legal, all others indexes are illegal. To support this, new instructions are defined specifically for zero index which is hardcoded, along an implicit 'VectorIndex0' operand. Since the index operand is implicit and takes no bits in the encoding, custom decoding is required to add the operand. I'm not sure if this is the best approach but the predicate constraint on a subset of an operand is unusual. Would be interested to hear some alternatives. The instructions are predicated on 'HasNEONorStreamingSVE', i.e. they're enabled by either +neon or +streaming-sve. This follows on from the work in D106272 to support the subset of SVE(2) instructions that are legal in streaming mode. Depends on D107902. Reviewed By: sdesmalen Differential Revision: https://reviews.llvm.org/D107903	2021-09-03 07:59:17 +00:00
Cullen Rhodes	1dcd900d1d	[AArch64][ISel] NFC: DAG.getMachineFunction() -> MF Reviewed By: sdesmalen Differential Revision: https://reviews.llvm.org/D109135	2021-09-03 07:59:17 +00:00
Amara Emerson	6d9505b8e0	[AArch64][GlobalISel] Support for folding G_ROTR as shifted operands. This allows selection like: eor w0, w1, w2, ror #8 Saves 500 bytes on ClamAV -Os, which is 0.1%. Differential Revision: https://reviews.llvm.org/D109206	2021-09-02 21:37:24 -07:00
Qiu Chaofan	d0f9553ef5	[PowerPC] Enable fast-isel on AIX 64 subtarget This patch basically enables fast-isel for AIX 64-bit subtarget (previously enabled only for ELF 64). The initial motivation is to introduce branch folding to AIX generated code for correct debug behavior. I also saw some compiling time improvement in a few LLVM test-suite benchmarks. (toast, dbms, cjpeg, burg, etc.) Reviewed By: jsji Differential Revision: https://reviews.llvm.org/D98844	2021-09-03 11:33:45 +08:00
Matt Arsenault	79bcd4a7db	AMDGPU: Remove FeatureLocalMemorySize0 There's no reason to make this an explicit feature, since it's implied by the lack of a feature with a size.	2021-09-02 22:43:01 -04:00
Alexander Pivovarov	6cd4b508a8	[RISCV] Add SiFive core S51 Add SiFive core s51 as rv64imac RocketModel Reviewed-By: MaskRay, evandro Differential Revision: https://reviews.llvm.org/D108886	2021-09-02 18:45:25 -07:00
Alexander Pivovarov	1104e3258b	Fix typo in RISCVMatInt.cpp comments	2021-09-02 18:11:09 -07:00
Stanislav Mekhanoshin	78fbd1aa3d	[AMDGPU] Process any power of 2 in optimizeCompareInstr Differential Revision: https://reviews.llvm.org/D109201	2021-09-02 17:39:17 -07:00
Stanislav Mekhanoshin	2cfda6a691	[AMDGPU] Fold immediates in the optimizeCompareInstr Peephole works before the first SIFoldOperands so most of the immediates are in registers. Differential Revision: https://reviews.llvm.org/D109186	2021-09-02 17:23:26 -07:00
Sam Clegg	c32884c482	[WebAssembly] Rename WrapperPIC -> WrapperREL. NFC This ISD node/wrapper represents am address which is relative to a base address and therefore lowers to `i32.const` rather than `global.get`. Use this wrapper type for TLS-relative addresses, paving the way for the non-REL wrapper to be used to external TLS address once those are supported. Differential Revision: https://reviews.llvm.org/D109179	2021-09-02 20:04:34 -04:00
Kirill Stoimenov	cf53c6c971	[asan] Fixed link error by setting jump symbol to R_X86_64_PLT32. Fixing this link error: ld: error: relocation R_X86_64_PC32 cannot be used against symbol __asan_report_load...; recompile with -fPIC Reviewed By: vitalybuka Differential Revision: https://reviews.llvm.org/D109183	2021-09-02 21:50:56 +00:00
Sam Clegg	4664590d53	[WebAssemlby] Remove redundant SDTypeProfile. NFC I added this back in https://reviews.llvm.org/D54647 but it wasn't actually needed. Differential Revision: https://reviews.llvm.org/D109176	2021-09-02 15:21:22 -04:00
Sam Clegg	ad2f94f398	[WebAssembly] Fix names of WebAssemblyWrapper SDNodes. NFC Other platforms all use CamelCase as normal for these wrapper nodes. Differential Revision: https://reviews.llvm.org/D109172	2021-09-02 13:54:44 -04:00
Heejin Ahn	28780e59f6	[WebAssembly] Add Wasm SjLj support This add support for SjLj using Wasm exception handling instructions: https://github.com/WebAssembly/exception-handling/blob/master/proposals/exception-handling/Exceptions.md This does not yet support the mixed use of EH and SjLj within a function. It will be added in a follow-up CL. This currently passes all SjLj Emscripten tests for wasm0/1/2/3/s, except for the below: - `test_longjmp_standalone`: Uses Node - `test_dlfcn_longjmp`: Uses NodeRAWFS - `test_longjmp_throw`: Mixes EH and SjLj - `test_exceptions_longjmp1`: Mixes EH and SjLj - `test_exceptions_longjmp2`: Mixes EH and SjLj - `test_exceptions_longjmp3`: Mixes EH and SjLj Reviewed By: dschuff, tlively Differential Revision: https://reviews.llvm.org/D108960	2021-09-02 10:51:02 -07:00
Nick Desaulniers	6860b136b9	[MipsISelLowering] avoid emitting libcalls to __multi3 Similar to D108842 and D108844. __has_builtin(builtin_mul_overflow) returns true for 32b MIPS targets, but Clang is deferring to compiler RT when encountering long long types. This breaks MIPS malta_defconfig builds of the Linux kernel that are using __builtin_mul_overflow with these types for these targets. If the semantics of __has_builtin mean "the compiler resolves these, always" then we shouldn't conditionally emit a libcall. This will still need to be worked around in the Linux kernel in order to continue to support malta_defconfig builds of the Linux kernel for this target with older releases of clang. Link: https://bugs.llvm.org/show_bug.cgi?id=28629 Link: https://github.com/ClangBuiltLinux/linux/issues/1438 Reviewed By: rengolin Differential Revision: https://reviews.llvm.org/D108926	2021-09-02 10:41:37 -07:00
Craig Topper	3e89cc5cda	[X86] Remove isel predicates for xgetbv/xsetbv instructions so they can work on Windows. https://reviews.llvm.org/D56686 was supposed to allow these to work on Windows without needing to enable the xsave feature to match MSVC. It seems this didn't work because the backend isel patterns would still block it. This patch removes the predicates from the isel patterns. Fixes PR51706. Reviewed By: pengfei Differential Revision: https://reviews.llvm.org/D109097	2021-09-02 10:25:02 -07:00
Stanislav Mekhanoshin	832c87b4fb	[AMDGPU] Use S_BITCMP0_* to replace AND in optimizeCompareInstr These can be used for reversed conditions if result of the AND is unused except in the compare: s_cmp_eq_u32 (s_and_b32 $src, 1), 0 => s_bitcmp0_b32 $src, 0 s_cmp_eq_i32 (s_and_b32 $src, 1), 0 => s_bitcmp0_b32 $src, 0 s_cmp_eq_u64 (s_and_b64 $src, 1), 0 => s_bitcmp0_b64 $src, 0 s_cmp_lg_u32 (s_and_b32 $src, 1), 1 => s_bitcmp0_b32 $src, 0 s_cmp_lg_i32 (s_and_b32 $src, 1), 1 => s_bitcmp0_b32 $src, 0 s_cmp_lg_u64 (s_and_b64 $src, 1), 1 => s_bitcmp0_b64 $src, 0 Differential Revision: https://reviews.llvm.org/D109099	2021-09-02 09:38:01 -07:00
Simon Pilgrim	d66d520fe1	[X86][SSE] combineMulToPMADDWD - improve recognition of sign/zero extended upper bits PMADDWD(v8i16 x, v8i16 y) == (v4i32) { (int)x[0]y[0] + (int)x[1]y[1], ..., (int)x[6]y[6] + (int)x[7]y[7] } Currently combineMulToPMADDWD only folds cases where the upper 17 bits of both vXi32 inputs are known zero (i.e. the first half is positive and the second half of the pair is zero in each 2xi16 pair), this can be relaxed to only require one zero-extended input if the other input has at least 17 sign bits. That way the sign of the result is still preserved, and the second half is still zero. Noticed while investigating PR47437. Differential Revision: https://reviews.llvm.org/D108522	2021-09-02 17:36:22 +01:00
Kazu Hirata	e1bb54b593	[clangd, llvm] Remove redundant calls to c_str() (NFC) Identified with readability-redundant-string-cstr.	2021-09-02 09:07:13 -07:00
Bradley Smith	14e1a4a6ee	[AArch64][SVE] Workaround incorrect types when lowering fixed length gather/scatter When lowering a fixed length gather/scatter the index type is assumed to be the same as the memory type, this is incorrect in cases where the extension of the index has been folded into the addressing mode. For now add a temporary workaround to fix the codegen faults caused by this by preventing the removal of this extension. At a later date the lowering for SVE gather/scatters will be redesigned to improve the way addressing modes are handled. As a short term side effect of this change, the addressing modes generated for fixed length gather/scatters will not be optimal. Differential Revision: https://reviews.llvm.org/D109145	2021-09-02 15:07:24 +00:00
Craig Topper	b5fd6b46f5	[RISCV] Teach instruction selection to elide sext.w in some cases. If a sext_inreg is up for isel, and all its users are W instructions, we can skip emitting the sext_inreg. This helpful if the producing instruction can't become a W instruction. Reviewed By: asb Differential Revision: https://reviews.llvm.org/D108966	2021-09-02 07:54:34 -07:00
Evandro Menezes	5ebdb07e7e	[RISCV] Enable shrink wrap by default Differential Revision: https://reviews.llvm.org/D109037	2021-09-02 09:47:58 -05:00
Craig Topper	e4e69ba4d1	[RISCV] Split PseudoVSETVLI into 2 instructions to allow different register classes for rs1. X0 has special meaning for vsetvli, we need to make sure we never create it a vsetvli that uses it by accident. This could happen if the register coalescer coalesces a copy from X0 into this instruction. This patch splits the instruction so that we can have GPRNoX0 register class to use for the cases where we don't want the source to be X0. The verifier won't let us explicitly use X0 on a GPRNoX0 operand so we need a separate pseudo for those cases. I don't currently have a failing example for this. There was a failure in D107957, but the coalescable copy from that example should have been optimized away much earlier so I've fixed that. This is not a complete fix. We still need to prevent the same possible issue on the AVL operand of all of the vector instruction pseudos. I don't want to make two versions of all of those so we need to find a different solution for those. I have an idea I'm going to try. Differential Revision: https://reviews.llvm.org/D109110	2021-09-02 07:45:31 -07:00
Piotr Sobczak	30d6c39bca	[AMDGPU] Add merging into S_BUFFER_LOAD_DWORDX8_IMM Extend SILoadStoreOptimizer to merge into DWORDX8 variant of S_BUFFER_LOAD. Merging into DWORDX2 and DWORDX4 variants is handled already. Differential Revision: https://reviews.llvm.org/D108909	2021-09-02 16:26:25 +02:00
David Green	9cb8f4d1ad	[ARM] Add a tail-predication loop predicate register The semantics of tail predication loops means that the value of LR as an instruction is executed determines the predicate. In other words: mov r3, #3 DLSTP lr, r3 // Start tail predication, lr==3 VADD.s32 q0, q1, q2 // Lanes 0,1 and 2 are updated in q0. mov lr, #1 VADD.s32 q0, q1, q2 // Only first lane is updated. This means that the value of lr cannot be spilled and re-used in tail predication regions without potentially altering the behaviour of the program. More lanes than required could be stored, for example, and in the case of a gather those lanes might not have been setup, leading to alignment exceptions. This patch adds a new lr predicate operand to MVE instructions in order to keep a reference to the lr that they use as a tail predicate. It will usually hold the zeroreg meaning not predicated, being set to the LR phi value in the MVETPAndVPTOptimisationsPass. This will prevent it from being spilled anywhere that it needs to be used. A lot of tests needed updating. Differential Revision: https://reviews.llvm.org/D107638	2021-09-02 13:42:58 +01:00
Roman Lebedev	3f1f08f0ed	Revert @llvm.isnan intrinsic patchset. Please refer to https://lists.llvm.org/pipermail/llvm-dev/2021-September/152440.html (and that whole thread.) TLDR: the original patch had no prior RFC, yet it had some changes that really need a proper RFC discussion. It won't be productive to discuss such an RFC, once it's actually posted, while said patch is already committed, because that introduces bias towards already-committed stuff, and the tree is potentially in broken state meanwhile. While the end result of discussion may lead back to the current design, it may also not lead to the current design. Therefore i take it upon myself to revert the tree back to last known good state. This reverts commit `4c4093e6e3`. This reverts commit `0a2b1ba33a`. This reverts commit `d9873711cb`. This reverts commit `791006fb8c`. This reverts commit `c22b64ef66`. This reverts commit `72ebcd3198`. This reverts commit `5fa6039a5f`. This reverts commit `9efda541bf`. This reverts commit `94d3ff09cf`.	2021-09-02 13:53:56 +03:00
Simon Pilgrim	b0acd6c369	[X86] Fold PMADD(x,0) or PMADD(0,x) -> 0 Pulled out of D108522 - handle zero-operand cases for PMADDWD/VPMADDUBSW ops	2021-09-02 10:48:50 +01:00
David Sherwood	d581d94385	[SVE] Fix the FP arithmetic instruction costs for SVE Several FP instructions (fadd, fsub, etc.) were incorrectly assigned a higher cost for SVE because they have custom lowering, however we know they are legal. This patch explicitly assigns a cost of 2 to these opcodes. Tests added here: Analysis/CostModel/AArch64/arith-fp-sve.ll Differential Revision: https://reviews.llvm.org/D108993	2021-09-02 09:55:13 +01:00
Jinsong Ji	8671191d26	[NFC][PowerPC] Small code refactor in LoopInstrFormPrep Avoid some duplicate code. Reviewed By: #powerpc, shchenz Differential Revision: https://reviews.llvm.org/D109083	2021-09-02 03:16:01 +00:00
Chen Zheng	2596120199	[PowerPC] small code format refactor ; NFC address the code review comments in patch https://reviews.llvm.org/D105872	2021-09-02 01:39:32 +00:00
Jon Roelofs	9237eda304	Revert "[AArch64][GlobalISel] Legalize bswap <2 x i16>" This reverts commit `5cd63e9ec2`. https://bugs.llvm.org/show_bug.cgi?id=51707 The sequence feeding in/out of the rev32/ushr isn't quite right: _swap: ldr h0, [x0] ldr h1, [x0, #2] - mov v0.h[1], v1.h[0] + mov v0.s[1], v1.s[0] rev32 v0.8b, v0.8b ushr v0.2s, v0.2s, #16 - mov h1, v0.h[1] + mov s1, v0.s[1] str h0, [x0] str h1, [x0, #2] ret	2021-09-01 16:49:20 -07:00
Stanislav Mekhanoshin	f3645c792a	[AMDGPU] Use S_BITCMP1_* to replace AND in optimizeCompareInstr Differential Revision: https://reviews.llvm.org/D109082	2021-09-01 15:59:12 -07:00
Stanislav Mekhanoshin	bf77b11277	[AMDGPU] Introduce optimizeCompareInstr The following patterns are currently handled: s_cmp_eq_u32 (s_and_b32 $src, 1), 1 => s_and_b32 $src, 1 s_cmp_eq_i32 (s_and_b32 $src, 1), 1 => s_and_b32 $src, 1 s_cmp_eq_u64 (s_and_b64 $src, 1), 1 => s_and_b64 $src, 1 s_cmp_ge_u32 (s_and_b32 $src, 1), 1 => s_and_b32 $src, 1 s_cmp_ge_i32 (s_and_b32 $src, 1), 1 => s_and_b32 $src, 1 s_cmp_lg_u32 (s_and_b32 $src, 1), 0 => s_and_b32 $src, 1 s_cmp_lg_i32 (s_and_b32 $src, 1), 0 => s_and_b32 $src, 1 s_cmp_lg_u64 (s_and_b64 $src, 1), 0 => s_and_b64 $src, 1 s_cmp_gt_u32 (s_and_b32 $src, 1), 0 => s_and_b32 $src, 1 s_cmp_gt_i32 (s_and_b32 $src, 1), 0 => s_and_b32 $src, 1 Differential Revision: https://reviews.llvm.org/D109031	2021-09-01 15:57:05 -07:00
Alexander Pivovarov	4b04d54206	[RISCV] Fix typo in RISCVSchedSiFive7.td Fix typo in "microarchitecure". Differential Revision: https://reviews.llvm.org/D109006	2021-09-01 16:39:48 -05:00
David Green	49476a4d66	[ARM] Add MVE lowering for fptosi.sat This adds lowering of the llvm.fptosi.sat and llvm.fptoui.sat intinsics, selecting a VCVT instruction which under MVE will inherently perform the saturate. Differential Revision: https://reviews.llvm.org/D107865	2021-09-01 22:38:47 +01:00
alex-t	e3cbf1d437	[AMDGPU] enable scalar compare in truncate selection Currently, the truncate selection dag node is expanded as a bitwise AND plus compare to 1. This change enables scalar comparison in the pattern if the truncate node is uniform. Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D108925	2021-09-01 23:35:11 +03:00
Nikita Popov	7f058ce8c2	[WebAssembly] Support opaque pointers in FixFunctionBitcasts With opaque pointers, no actual bitcasts will be present. Instead, there will be a mismatch between the call FunctionType and the function ValueType. Change the code to collect CallBases specifically (rather than general Uses) and compare these types. RAUW is no longer performed, as there would no longer be any bitcasts that can be RAUWd. Differential Revision: https://reviews.llvm.org/D108880	2021-09-01 22:17:24 +02:00
Craig Topper	ccbb4c8b4f	[RISCV] Fold (RISCVISD::SELECT_CC X, Y, CC, Z, Z) -> Z. If the true and false values are the same, we don't need a SELECT_CC. This would normally be folded before a select is legalized to select_cc. The test case exploits the late legalization of vscale to trigger a case where they become identical after legalization. This works around an issue found on a test case in D107957. In that case the true/false values were both eventually 0 and the select was used by a vector AVL operand. The select_cc got expanded to control flow and a phi, but the phi inputs were both copies from X0. MachineIR optimizations simplified this to a single copy from X0 going into the vector instruction. This became the input of a vsetvli after vsetvli insertion. Then register coalescing folded the copy into the vsetvli. X0 as the source of a vsetvli is a special encoding and should not be created by coalesing. We need to fix our vsetvli handling to make sure this can never happen any other way, but removing the unneeded select is still a worthwhile optimization.	2021-09-01 12:37:52 -07:00
Philip Reames	29fa37ec9f	[SCEV] If max BTC is zero, then so is the exact BTC [2 of 2] This extends D108921 into a generic rule applied to constructing ExitLimits along all paths. The remaining paths (primarily howFarToZero) don't have the same reasoning about UB sensitivity as the howManyLessThan ones did. Instead, the remain cause for max counts being more precise than exact counts is that we apply context sensitive loop guards on the max path, and not on the exact path. That choice is mildly suspect, but out of scope of this patch. The MVETailPredication.cpp change deserves a bit of explanation. We were previously figuring out that two SCEVs happened to be equal because the happened to be identical. When we optimized one with context sensitive information, but not the other, we lost the ability to prove them equal. So, cover this case by subtracting and then applying loop guards again. Without this, we see changes in test/CodeGen/Thumb2/mve-blockplacement.ll Differential Revision: https://reviews.llvm.org/D109015	2021-09-01 11:51:48 -07:00
Arthur Eubanks	b9b419a13c	[NFC] Remove redundant code added in `04ce2de3`	2021-09-01 11:30:07 -07:00
Thomas Lively	fec4749200	[WebAssembly] Lower v2f32 to v2f64 extending loads with promote_low Previously extra wide v4f32 to v4f64 extending loads would be legalized to v2f32 to v2f64 extending loads, which would then be scalarized by legalization. (v2f32 to v2f64 extending loads not produced by legalization were already being emitted correctly.) Instead, mark v2f32 to v2f64 extending loads as legal and explicitly lower them using promote_low. This regresses the addressing modes supported for the extloads not produced by legalization, but that's a fine trade off for now. Differential Revision: https://reviews.llvm.org/D108496	2021-09-01 10:27:42 -07:00
Amara Emerson	a86bbe1e31	[AArch64][GlobalISel] Handle any-extending FPR loads in manual selection code. When we have an any-extending FPR bank load, none of the tablegen patterns match and we fall back to the C++ selector. Like with the truncating stores that were fixed recently, the C++ wasn't able to handle it and ended up generating invalid copies between different size regclasses. This change adds handling for this case, splitting the load into a regular load and a SUBREG_TO_REG to extend it into the original wide destination reg.	2021-09-01 10:19:22 -07:00
hsmahesha	97688bfd3d	Revert "Revert "Disable ReplaceLDS pass, patch up tests to match"" This reverts commit `5ae6804d17`.	2021-09-01 21:52:50 +05:30
hsmahesha	5ae6804d17	Revert "Disable ReplaceLDS pass, patch up tests to match" This reverts commit `50ad3478bd`. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D109062	2021-09-01 21:19:39 +05:30
Alexander Kornienko	893ac53afc	Fix -Wunused-variable	2021-09-01 11:29:30 +02:00
Christudasan Devadasan	4dab15288d	[AMDGPU] Introduce RC flags for vector register classes Configure and use the TSFlags in TargetRegisterClass to have unique flags for VGPR and AGPR register classes. The vector register class queries like `hasVGPRs` will now become more efficient with just a bitwise operation. Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D108815	2021-09-01 02:55:45 -04:00
Kai Luo	5eaebd5d64	[PowerPC] Implement quadword atomic load/store Add support to load/store i128 atomically. Reviewed By: jsji Differential Revision: https://reviews.llvm.org/D105612	2021-09-01 06:55:40 +00:00
Luke	a78dd726f4	[SLP][RISCV] Implement unsigned getMinVectorRegisterBitWidth() for RISCV Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D108973	2021-09-01 14:25:15 +08:00
hsmahesha	98f4713122	[AMDGPU] Split entry basic block after alloca instructions. While initializing the LDS pointers within entry basic block of kernel(s), make sure that the entry basic block is split after alloca instructions. Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D108971	2021-09-01 10:18:44 +05:30
Wang, Pengfei	74043caef2	[X86] Enable half type support in inline assembly constraints Reviewed By: LuoYuanke Differential Revision: https://reviews.llvm.org/D105799	2021-09-01 09:29:31 +08:00
Nick Desaulniers	e9b3f25730	[RISCVISelLowering] avoid emitting libcalls to __mulodi4() and __multi3() Similar to D108842, D108844, D108926, D108928, and D108936. __has_builtin(builtin_mul_overflow) returns true for 32b RISCV targets, but Clang is deferring to compiler RT when encountering long long types. If the semantics of __has_builtin mean "the compiler resolves these, always" then we shouldn't conditionally emit a libcall. Link: https://bugs.llvm.org/show_bug.cgi?id=28629 Reviewed By: craig.topper Differential Revision: https://reviews.llvm.org/D108939	2021-08-31 11:23:56 -07:00
Nick Desaulniers	d8b6ae072d	[PPCISelLowering] avoid emitting libcalls to __mulodi4() Similar to D108842, D108844, and D108926. __has_builtin(builtin_mul_overflow) returns true for 32b PPC targets, but Clang is deferring to compiler RT when encountering long long types. This breaks ppc44x_defconfig + CONFIG_BLK_DEV_NBD=y builds of the Linux kernel that are using builtin_mul_overflow with these types for these targets. If the semantics of __has_builtin mean "the compiler resolves these, always" then we shouldn't conditionally emit a libcall. This will still need to be worked around in the Linux kernel in order to continue to support these builds of the Linux kernel for this target with older releases of clang. Link: https://bugs.llvm.org/show_bug.cgi?id=28629 Link: https://github.com/ClangBuiltLinux/linux/issues/1438 Reviewed By: nemanjai Differential Revision: https://reviews.llvm.org/D108936	2021-08-31 11:09:58 -07:00
Joe Nash	c96839265a	[AMDGPU] Enable ds_min/ds_max on more subtargets Adds patterns for f64 ds_min/ds_max. Shrinks HasLDSFPAtomics scope to enable f32. Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D108994 Change-Id: Id890b677841ee588b20d42b1bb3f4cdbf6e9ba1a	2021-08-31 13:22:31 -04:00
David Green	22c384129e	[ARM] Add missing validForTailPredication for VMINNM/VMAXNM Apparently this was missing, preventing the generation of tail predication loops containing VMINNM, VMAXNM, VMINNMA and VMAXNMA.	2021-08-31 18:19:03 +01:00
Simon Pilgrim	9e2d14c285	[X86] Copy X86SchedSkylakeServer.td to X86SchedIceLake.td Icelake, Rocketlake and Tigerlake targets currently use the SkylakeServer scheduler model, despite being a later microarchitecture, leading to both reported bugs (PR48110) and discrepancies when comparing llvm-mca reports to other profiling tools (OSACA, uops, uica, etc.). And tbh I'm getting sick of llvm-mca getting blamed for what are backend scheduler model issues :-( This patch doesn't attempt to fix any of these discrepancies - there should be no changes in codegen - its a setup patch that copies the skx model, renames all the resources, adds the additional ports (but doesn't reference them yet) and updates the llvm-exegesis pfm counter mappings (based off https://sourceforge.net/p/perfmon2/libpfm4/ci/master/tree/lib/events/intel_icl_events.h). This should make it trivial for anyone with hardware access to use llvm-exegesis reports to iteratively improve the model (my attempts to get hold of a cheap tiger lake box haven't been fruitful yet....). I will copy the SkylakeServer llvm-mca resource tests as follow up commits - the diff should entirely be the resource renames. Differential Revision: https://reviews.llvm.org/D108914	2021-08-31 11:57:20 +01:00
Simon Wallis	f417b660ee	[Arm] Add assert in T2 Imm7s code emitter Add assert to provoke failure in object file output, not just in disassembly output. Reviewed By: yroux Differential Revision: https://reviews.llvm.org/D107259	2021-08-31 08:16:48 +01:00
Alexander Pivovarov	eb946cc5b6	Fix typo in comments Reviewed By: MaskRay, jsji Differential Revision: https://reviews.llvm.org/D108857	2021-08-31 11:55:40 +05:30
Heejin Ahn	3419e85b15	[WebAssembly] Free setjmpTable before exiting calls in EmSjLj This is an improvement over D107852. We don't need to enumerate specific function names; we can just check for `noreturn` attribute. This also requires us to make sure `__resumeExeption` and `emscripten_longjmp` have `noreturn` attribute too; one of them is a JS function and the other calls a JS function so Clang does not have a way to deduce they don't return. This is effectively NFC, because I'm not sure if there is an additional case this case covers; if we add a custom function call that has `noreturn` attribute, it will be processed within the SjLj handling and turned into `__invoke` call. So this really applies to some special functions like `emscripten_longjmp`. Reviewed By: dschuff Differential Revision: https://reviews.llvm.org/D108955	2021-08-30 21:46:25 -07:00
Heejin Ahn	b8fc71b7ae	[WebAssembly] Share rethrowing BBs in LowerEmscriptenEHSjLj There are three kinds of "rethrowing" BBs in this pass: 1. In Emscripten SjLj, after a possibly longjmping function call, we check if the thrown longjmp corresponds to one of setjmps within the current function. If not, we rethrow the longjmp by calling `emscripten_longjmp`. 2. In Emscripten EH, after a possibly throwing function call, we check if the thrown exception corresponds to the current `catch` clauses. If not, we rethrow the exception by calling `__resumeException`. 3. When both Emscripten EH and SjLj are used, when we check for an exception after a possibly throwing function call, it is possible that we get not an exception but a longjmp. In this case, we shouldn't swallow it; we should rethrow the longjmp by calling `emscripten_longjmp`. 4. When both Emscripten EH and SjLj are used, when we check for a longjmp after a possibly longjmping function call, it is possible that we get not a longjmp but an exception. In this case, we shouldn't swallot it; we should rethrow the exception by calling `__resumeException`. Case 1 is in Emscripten SjLj, 2 is in Emscripten EH, and 3 and 4 are relevant when both Emscripten EH and SjLj are used. 3 and 4 were first implemented in D106525. We create BBs for 1, 3, and 4 in this pass. We create those BBs for every throwing/longjmping function call, along with other BBs that contain condition checks. What this CL does is to create a single BB within a function for each of 1, 3, and 4 cases. These BBs are exiting BBs in the function and thus don't have successors, so easy to be shared between calls. The names of BBs created are: Case 1: `call.em.longjmp` Case 3: `rethrow.exn` Case 4: `rethrow.longjmp` For the case 2 we don't currently create BBs; we only replace the existing `resume` instruction with `call @__resumeException`. And Clang already creates only a single `resume` BB per function and reuses it, so we don't need to optimize this case. Not sure what are good benchmarks for EH/SjLj, but this decreases the size of the object file for `grfmt_jpeg.bc` (presumably from opencv) we got from one of our users by 8.9%. Even after running `wasm-opt -O4` on them, there is still 4.8% improvement. Reviewed By: dschuff Differential Revision: https://reviews.llvm.org/D108945	2021-08-30 21:44:34 -07:00
Nikita Popov	c1b7540645	[TTI] Sink IVDescriptors.h include (NFC) Forward declare RecurrenceDescriptor and include IVDescritor.h only in implementation code that actually needs it.	2021-08-30 22:41:58 +02:00
Owen Anderson	db9de22f2b	Teach the AArch64 backend patterns to generate the EOR3 instruction. Adds patterns to match the EOR3 instruction. Reviewed By: dmgreen Differential Revision: https://reviews.llvm.org/D108793	2021-08-30 20:01:08 +00:00
David Green	efa340fbd2	[ARM] Workaround tailpredication min/max costmodel The min/max intrinsics are not yet canonical, but when they are the tail predications analysis will change from treating them like icmp to treating them like intrinsics. Unfortunately, they can currently produce better code by not being tail predicated thanks to the vectorizer picking higher VF's and the backend folding to better instructions (especially for saturate patterns). In the long run we will need to improve the vectorizers cost modelling, recognizing the instruction directly, but in the meantime this treats min/max as before to prevent performance regressions.	2021-08-30 19:19:51 +01:00
Nikita Popov	0529e2e018	[InstrInfo] Use 64-bit immediates for analyzeCompare() (NFCI) The backend generally uses 64-bit immediates (e.g. what MachineOperand::getImm() returns), so use that for analyzeCompare() and optimizeCompareInst() as well. This avoids truncation for targets that support immediates larger 32-bit. In particular, we can avoid the bugprone value normalization hack in the AArch64 target. This is a followup to D108076. Differential Revision: https://reviews.llvm.org/D108875	2021-08-30 19:46:04 +02:00
Kazu Hirata	c50faffb4e	[llvm] Remove redundant calls to str() and c_str() (NFC) Identified with readability-redundant-string-cstr.	2021-08-30 09:05:05 -07:00
Craig Topper	0560a4adb3	[RISCV] Enable CONCAT_VECTORS for fixed FP vectors. Reviewed By: frasercrmck Differential Revision: https://reviews.llvm.org/D108487	2021-08-30 08:47:45 -07:00
Simon Pilgrim	af2920ec6f	[TTI][X86] getArithmeticInstrCost - move opcode canonicalization before all target-specific costs. NFCI. The GLM/SLM special cases still get tested first but after the the MUL/DIV/REM pattern detection - this will be necessary for when we make the SLM vXi32 MUL canonicalization generic to improve PMULLW/PMULHW/PMADDDW cost support etc.	2021-08-30 12:24:59 +01:00
Wang, Pengfei	ab40dbfe03	[X86] AVX512FP16 instructions enabling 6/6 Enable FP16 complex FMA instructions. Ref.: https://software.intel.com/content/www/us/en/develop/download/intel-avx512-fp16-architecture-specification.html Reviewed By: LuoYuanke Differential Revision: https://reviews.llvm.org/D105269	2021-08-30 13:08:45 +08:00
Qiu Chaofan	3bdd850d0c	[PowerPC] Set branch/call instructions as no hasSideEffects PowerPC can model these instructions, so we don't need this flag set. Reviewed By: shchenz, jsji Differential Revision: https://reviews.llvm.org/D71983	2021-08-30 12:23:35 +08:00
Vince Bridgers	55ba1de7c5	[X86] Remove X86LowerAMXType::getRowFromCol from X86LowerAMXType.cpp Remove method X86LowerAMXType::getRowFromCol since it's not used, and it's causing a warning. Reviewed By: LuoYuanke Differential Revision: https://reviews.llvm.org/D108862	2021-08-29 12:27:34 -05:00
Yonghong Song	4948927058	[BPF] support btf_tag attribute in .BTF section A new kind BTF_KIND_TAG is added to .BTF to encode btf_tag attributes. The format looks like CommonType.name : attribute string CommonType.type : attached to a struct/union/func/var. CommonType.info : encoding BTF_KIND_TAG kflag == 1 to indicate the attribute is for CommonType.type, or kflag == 0 for struct/union member or func argument. one uint32_t : to encode which member/argument starting from 0. If one particular type or member/argument has more than one attribute, multiple BTF_KIND_TAG will be generated. Differential Revision: https://reviews.llvm.org/D106622	2021-08-28 21:02:27 -07:00
Nikita Popov	16086d47c0	[WebAssembly] Fix FastISel of condition in different block (PR51651) If the icmp is in a different block, then the register for the icmp operand may not be initialized, as it nominally does not have cross-block uses. Add a check that the icmp is in the same block as the branch, which should be the common case. This matches what X86 FastISel does: `5b6b090cf2/llvm/lib/Target/X86/X86FastISel.cpp (L1648)` The "not" transform that could have a similar issue is dropped entirely, because it is currently dead: The incoming value is a branch or select condition of type i1, but this code requires an i32 to trigger. Fixes https://bugs.llvm.org/show_bug.cgi?id=51651. Differential Revision: https://reviews.llvm.org/D108840	2021-08-28 10:28:24 +02:00
Nick Desaulniers	c8c176d999	[MipsISelLowering] avoid emitting libcalls to __mulodi4() __has_builtin(__builtin_mul_overflow) returns true for 32b MIPS targets, but Clang is deferring to compiler RT when encountering `long long` types. This breaks sanitizer builds of the Linux kernel that are using __builtin_mul_overflow with these types for these targets. If the semantics of __has_builtin mean "the compiler resolves these, always" then we shouldn't conditionally emit a libcall. This will still need to be worked around in the Linux kernel in order to continue to support malta_defconfig builds of the Linux kernel for this target with older releases of clang. Link: https://bugs.llvm.org/show_bug.cgi?id=28629 Link: https://github.com/ClangBuiltLinux/linux/issues/1438 Reviewed By: rengolin Differential Revision: https://reviews.llvm.org/D108844	2021-08-27 15:15:36 -07:00
Nick Desaulniers	5c91b98c5d	[ARMISelLowering] avoid emitting libcalls to __mulodi4() __has_builtin(__builtin_mul_overflow) returns true for 32b ARM targets, but Clang is deferring to compiler RT when encountering `long long` types. This breaks sanitizer builds of the Linux kernel that are using __builtin_mul_overflow with these types for these targets. If the semantics of __has_builtin mean "the compiler resolves these, always" then we shouldn't conditionally emit a libcall. This will still need to be worked around in the Linux kernel in order to continue to support allmodconfig builds of the Linux kernel for this target with older releases of clang. Link: https://bugs.llvm.org/show_bug.cgi?id=28629 Link: https://github.com/ClangBuiltLinux/linux/issues/1438 Reviewed By: rengolin Differential Revision: https://reviews.llvm.org/D108842	2021-08-27 15:14:47 -07:00
Craig Topper	dbf0d8118c	[RISCV] Use ~0ULL instead of ~0U when checking for invalid ErrorInfo. ErrorInfo is a uint64_t and is initialized to all 1s. Not sure how to test this. Noticed while working on .insn support.	2021-08-27 12:30:33 -07:00
Roman Lebedev	6734018041	[Codegen][X86] EltsFromConsecutiveLoads(): if only have AVX1, ensure that the "load" is actually foldable (PR51615) This fixes another reproducer from https://bugs.llvm.org/show_bug.cgi?id=51615 And again, the fix lies not in the code added in D105390 In this case, we completely don't check that the "broadcast-from-mem" we create can actually fold the load. In this case, it's operand was not a load at all: ``` Combining: t16: v8i32 = vector_shuffle<0,u,u,u,0,u,u,u> t14, undef:v8i32 Creating new node: t29: i32 = undef RepeatLoad: t8: i32 = truncate t7 t7: i64 = extract_vector_elt t5, Constant:i64<0> t5: v2i64,ch = load<(load (s128) from %ir.arg)> t0, t2, undef:i64 t2: i64,ch = CopyFromReg t0, Register:i64 %0 t1: i64 = Register %0 t4: i64 = undef t3: i64 = Constant<0> Combining: t15: v8i32 = undef ``` Reviewed By: RKSimon Differential Revision: https://reviews.llvm.org/D108821	2021-08-27 20:26:53 +03:00
Philipp Krones	54e8cae565	[MC][RISCV] Add RISCV MCObjectFileInfo This makes sure, that the text section will have a 2-byte alignment, if the +c extension is enabled. Reviewed By: MaskRay, luismarques Differential Revision: https://reviews.llvm.org/D102052	2021-08-27 18:23:29 +01:00
Craig Topper	0eeab8b282	[RISCV] Add -riscv-v-fixed-length-vector-elen-max to limit the ELEN used for fixed length vectorization. This adds an ELEN limit for fixed length vectors. This will scalarize any elements larger than this. It will also disable some fractional LMULs. For example, if ELEN=32 then mf8 becomes illegal, i32/f32 vectors can't use any fractional LMULs, i16/f16 can only use mf2, and i8 can use mf2 and mf4. We may also need something for the scalable vectors, but that has interactions with the intrinsics and we can't scalarize a scalable vector. Longer term this should come from one of the Zve* features	2021-08-27 10:17:35 -07:00
Jun Ma	15b2a8e7fa	[AArch64][SVE] Optimize ptrue predicate pattern with known sve register width. For vectors that are exactly equal to getMaxSVEVectorSizeInBits, just use AArch64SVEPredPattern::all, which can enable the use of unpredicated ptrue when available. TestPlan: check-llvm Differential Revision: https://reviews.llvm.org/D108706	2021-08-27 20:03:48 +08:00
Jun Ma	8c47103491	[AArch64][SVE] Add API for conversion between SVE predicate pattern and element number. NFC This patch solely moves convert operation between SVE predicate pattern and element number into two small functions. It's pre-commit patch for optimize pture with known sve register width. Differential Revision: https://reviews.llvm.org/D108705	2021-08-27 20:03:48 +08:00
Jun Ma	3f919dfe0d	[AArch64][SVE] Use getPTrue uniformly.NFC.	2021-08-27 20:03:48 +08:00
Serge Pavlov	cdbe569fb6	[X86] Implement llvm.isnan(x86_fp80) as unordered comparison x86_fp80 format allows values that do not fit any of IEEE-754 category. Previously they were recognized by intrinsic __builtin_isnan as NaNs. Now this intrinsic is implemented using instruction FXAM, which distinguish between NaNs and unsupported values. It can make some programs behave differently. As a solution, this fix changes lowering of the intrinsic. If floating point exceptions are ignored, llvm.isnan is lowered into unordered comparison, as __buildtin_isnan was implemented earlier. In strictfp functions the intrinsic is lowered using FXAM, which does not raise exceptions even for signaling NaN, as required by IEEE-754 and C standards. Differential Revision: https://reviews.llvm.org/D108037	2021-08-27 18:06:07 +07:00
Nathan Sidwell	199ac3a839	[NFC][X86] Sret return register cleanup There are no paths into LowerFormalParms that have already specified the sret register. We always materialize a virtual and then assign it to the physical reg at the point of the return. Differential Revision: https://reviews.llvm.org/D108762	2021-08-27 04:03:49 -07:00
Ricky Taylor	8d3f112f0c	[M68k] Update pointer data layout Fixes PR51626. The M68k requires that all instruction, word and long word reads are aligned to word boundaries. From the 68020 onwards, there is a performance benefit from aligning long words to long word boundaries. The M68k uses the same data layout for pointers and integers. In line with this, this commit updates the pointer data layout to match the layout already set for 32-bit integers: 32:16:32. Differential Revision: https://reviews.llvm.org/D108792	2021-08-27 11:47:27 +01:00
Roman Lebedev	d4d459e747	[X86] AMD Zen 3: MULX w/ mem operand has the same throughput as with reg op Exegesis is faulty and sometimes when measuring throughput^-1 produces snippets that have loop-carried dependencies, which must be what caused me to incorrectly measure it originally. After looking much more carefully, the inverse throughput should match that of the MULX w/ reg op. As per llvm-exegesis measurements.	2021-08-27 13:27:05 +03:00
Roman Lebedev	0f04936a2d	[X86] AMD Zen 3: MULX produces low part of the result in 3cy, +1cy for high part As per llvm-exegesis measurements.	2021-08-27 13:27:05 +03:00
Matt Arsenault	04ce2de330	AMDGPU: Remove implicit argument attributes when introducing new calls In a future patch, a new set of amdgpu-no-* attributes will be introduced to indicate when a function does not need an implicitly passed input. This pass introduces new instances of these intrinsic calls, and should remove the attributes if they were present before.	2021-08-26 22:08:04 -04:00
Chen Zheng	324bd467a2	[PowerPC][ELF] make sure local variable space does not overlap with parameter save area Reviewed By: jsji Differential Revision: https://reviews.llvm.org/D105271	2021-08-27 01:58:41 +00:00
Matt Arsenault	088cc63640	AMDGPU: Invert AMDGPUAttributor Switch to using BitIntegerState for each of the inputs, and invert their meanings. This now diverges more from the old AMDGPUAnnotateKernelFeatures, but this isn't used yet anyway.	2021-08-26 21:32:13 -04:00
Matt Arsenault	46d82e7357	AMDGPU: Restrict attributor transforms We only really want this to add the custom attributes. Theoretically the regular transforms were already run at this point. Touching undefined behavior breaks a lot of tests when this is enabled by default, many of which are expecting to test handling of undef operations.	2021-08-26 21:08:51 -04:00
Matt Arsenault	cf32d61a05	AMDGPU: Remove hacky attribute deduction from AMDGPUAttributor amdgpu-calls and amdgpu-stack-objects don't really belong as attributes, and are currently a hacky way of passing an analysis into the DAG. These don't really belong in the IR, and don't really fit in with the other attributes. Remove these to facilitate inverting the pass. I don't exactly understand the indirect call test changes. These tests are using calls which are trivially replacable with a direct call, so I'm not sure what the point is.	2021-08-26 20:31:14 -04:00
Matt Arsenault	98d7aa435f	AMDGPU: Stop inferring use of llvm.amdgcn.kernarg.segment.ptr We no longer use this intrinsic outside of the backend and no longer support using it outside of kernels.	2021-08-26 20:30:03 -04:00
Heejin Ahn	f5cff292e2	[WebAssembly] Fix PHI when relaying longjmps When doing Emscritpen EH, if SjLj is also enabled and used and if the thrown exception has a possiblity being a longjmp instead of an exception, we shouldn't swallow it; we should rethrow, or relay it. It was done in D106525 and the code is here: `8441a8eea8/llvm/lib/Target/WebAssembly/WebAssemblyLowerEmscriptenEHSjLj.cpp (L858-L898)` Here is the pseudocode of that part: (copied from comments) ``` if (%__THREW__.val == 0 \|\| %__THREW__.val == 1) goto %tail else goto %longjmp.rethrow longjmp.rethrow: ;; This is longjmp. Rethrow it %__threwValue.val = __threwValue emscripten_longjmp(%__THREW__.val, %__threwValue.val); tail: ;; Nothing happened or an exception is thrown ... Continue exception handling ... ``` If the current BB (where the `invoke` is created) has successors that has the current BB as its PHI incoming node, now that has to change to `tail` in the pseudocode, because `tail` is the latest BB that is connected with the next BB, but this was missing. Reviewed By: tlively Differential Revision: https://reviews.llvm.org/D108785	2021-08-26 17:25:26 -07:00
Matt Arsenault	ce51c5d4a9	AMDGPU: Fix crashing on kernel declarations when lowering LDS This was trying to insert the used marker into a declaration.	2021-08-26 19:01:10 -04:00
Kirill Stoimenov	2e83a0efb9	[asan] Fixed a runtime crash. Looks like the NoRegister has some effect on the final code that is generated. My guess is that some optimization kicks in at the end? When I use -S to dump the assembly I get the correct version with 'shrq $3, %r8': movq %r9, %r8 shrq $3, %r8 movsbl 2147450880(%r8), %r8d But, when I disassemble the final binary I get RAX in stead of R8: mov %r9,%r8 shr $0x3,%rax movsbl 0x7fff8000(%r8),%r8d Reviewed By: vitalybuka Differential Revision: https://reviews.llvm.org/D108745	2021-08-26 20:30:25 +00:00
Jessica Paquette	2363a20001	[AArch64][GlobalISel] Optimize G_BUILD_VECTOR of undef + 1 elt -> SUBREG_TO_REG This pattern ``` %elt = ... something ... %undef = G_IMPLICIT_DEF %vec = G_BUILD_VECTOR %elt, %undef, %undef, ... %undef ``` Can be selected to a SUBREG_TO_REG, assuming `%elt` and `%vec` have the same register bank. We don't care about any of the bits in `%vec` aside from those in `%elt`, which just happens to be the 0th element. This is preferable to emitting `mov` instructions for every index. This gives minor code size improvements on the test suite at -Os. Differential Revision: https://reviews.llvm.org/D108773	2021-08-26 11:45:11 -07:00
Craig Topper	1b9417454e	[RISCV] Insert a sext_inreg when type legalizing i32 shl by constant on RV64. Similar to what we do for add/sub/mul. This can help remove some sext.w. There are some regressions on some bswap tests, but I have an idea how to fix that for a follow up. A new PACKW pattern is added to handle the new sext_inreg placement. Differential Revision: https://reviews.llvm.org/D108663	2021-08-26 10:20:19 -07:00
Stanislav Mekhanoshin	827dd17e26	[AMDGPU] Invert partial vgpr to agpr spill lane order On targets requiring VGPR alignment we may end up spilling an unaligned register if we were partially spilled odd number of leading lanes. The reminder will start with an odd register. This problem is solved by inverting the order of lanes to be spillied so that we start from the end. Differential Revision: https://reviews.llvm.org/D108732	2021-08-26 09:39:03 -07:00
Roman Lebedev	a8125bf4a8	[X86][Codegen] PR51615: don't replace wide volatile load with narrow broadcast-from-memory Even though https://bugs.llvm.org/show_bug.cgi?id=51615 appears to be introduced by D105390, the fix lies here. We can not replace a wide volatile load with a broadcast-from-memory, because that would narrow the load, which isn't legal for volatiles. Reviewed By: spatel Differential Revision: https://reviews.llvm.org/D108757	2021-08-26 18:46:49 +03:00
Jacob Bramley	05f3219b38	[AArch64] Lower fptoi.sat intrinsics for NEON. Following on from D102353, extend the fptoi.sat intrinsics to use NEON fcvt* instructions. Differential Revision: https://reviews.llvm.org/D108460	2021-08-26 15:37:00 +01:00
Simon Pilgrim	c17f5afa88	[X86] getShape - don't dereference dyn_cast<> dyn_cast can return nullptr, use cast<> to assert we have the correct type.	2021-08-26 15:08:13 +01:00
Simon Pilgrim	47f2affa08	Fix MSVC "result of 32-bit shift implicitly converted to 64 bits" warning. NFCI.	2021-08-26 15:08:12 +01:00
Jessica Clarke	8f89e2f6c9	[AMDGPU] Remove dead and broken ComplexPatterns SelectADDRParam was discovered as being dead 5 years ago and removed in `7b4ef068c6` but the unused ComplexPattern definition was left behind. SelectADDRDWord has never existed as far as I can tell, even back when AMDGPU was R600-only and called that. Reviewed By: foad Differential Revision: https://reviews.llvm.org/D108758	2021-08-26 12:48:32 +01:00
Andrea Di Biagio	4a5b191703	[X86][MCA] Address the latest issues with MULX reported in PR51495. It turns out that SchedWrite WriteIMulH was always assigned to the low half of the result of a MULX (rather than to the high half). To avoid confusion, this patch swaps the two MULX writes in the tablegen definition of MULX32/64. That way, write names better describe what they actually refer to; this also avoids further complications if in future we decide to reuse the same MulH writes to also model other scalar integer multiply instructions. I also had to swap the latency values for the two MULX writes to make sure that the change is effectively an NFC. In fact, none of the existing x86 tests were affected by this small refactoring. This patch also fixes a bug in MCA: a wrong latency value was propagated for instructions that perform multiple writes to a same register. This last issue was found by Roman while testing MULX on targets that define a different latency for the Low/High part of the result. Differential Revision: https://reviews.llvm.org/D108727	2021-08-26 12:08:20 +01:00
Matthew Devereau	9b830c798e	[AArch64][SVE] Teach cost model masked gathers/scatters are cheap Tell the cost model to use the scalable calculation for non-neon fixed vector. This results in a cheaper cost for fixed-length SVE masked gathers/scatters allowing the vectorizor to emit them more frequently.	2021-08-26 11:17:47 +01:00
David Green	6ffc6951a3	[AArch64] Remove unpredictable from narrowing instructions. Like other similar instructions the xtn2 family do not have side effects, and explicitly marking them as such can help improve scheduling freedom.	2021-08-26 09:43:44 +01:00
Heejin Ahn	e849d99df1	[WebAssembly] Use entry block only for initializations in EmSjLj Emscripten SjLj transformation is done in four steps. This will be mostly the same for the soon-to-be-added Wasm SjLj; the step 1, 3, and 4 will be shared and there will be separate way of doing step 2. 1. Initialize `setjmpTable` and `setjmpTableSize` in the entry BB 2. Handle `setjmp` callsites 3. Handle `longjmp` callsites 4. Cleanup and update SSA We initialize `setjmpTable` and `setjmpTableSize` in the entry BB. But if the entry BB contains a `setjmp` call, some `setjmp` handling transformation will also happen in the entry BB, such as calling `saveSetjmp`. This is fine for Emscripten SjLj but not for Wasm SjLj, because in Wasm SjLj we will add a dispatch BB that contains a `switch` right after the entry BB, from which we jump to one of post-`setjmp` BBs. And this dispatch BB should precede all `setjmp` calls. Emscripten SjLj (current): ``` entry: %setjmpTable = ... %setjmpTableSize = ... ... call @saveSetjmp(...) ``` Wasm SjLj (follow-up): ``` entry: %setjmpTable = ... %setjmpTableSize = ... setjmp.dispatch: ... ; Jump to the right post-setjmp BB, if we are returning from a ; longjmp. If this is the first setjmp call, go to %entry.split. switch i32 %no, label %entry.split [ i32 1, label %post.setjmp1 i32 2, label %post.setjmp2 ... i32 N, label %post.setjmpN ] entry.split: ... call @saveSetjmp(...) ``` So in Wasm SjLj we split the entry BB to make the entry block only for `setjmpTable` and `setjmpTableSize` initialization and insert a `setjmp.dispatch` BB. (This part is not in this CL. This will be a follow-up.) But note that Emscripten SjLj and Wasm SjLj share all steps except for the step 2. If we only split the entry BB only for Wasm SjLj, there will be one more `if`-`else` and the code will be more complicated. So this CL splits the entry BB in Emscripten SjLj and put only initialization stuff there as follows: Emscripten SjLj (this CL): ``` entry: %setjmpTable = ... %setjmpTableSize = ... br %entry.split entry.split: ... call @saveSetjmp(...) ``` This is just done to share code with Wasm SjLj. It adds an unnecessary branch but this will be removed in later optimization passes anyway. This is in effect NFC, meaning the program behavior will not change, but existing ll tests files have changed because the entry block was split. The reason I upload this in a separate CL is to make the Wasm SjLj diff tidier, because this changes many existing Emscripten SjLj tests, which can be confusing for the follow-up Wasm SjLj CL. Reviewed By: tlively Differential Revision: https://reviews.llvm.org/D108729	2021-08-25 15:46:57 -07:00
Heejin Ahn	2f88a30ca6	[WebAssembly] Extract longjmp handling in EmSjLj to a function (NFC) Emscripten SjLj and (soon-to-be-added) Wasm SjLj transformation share many steps: 1. Initialize `setjmpTable` and `setjmpTableSize` in the entry BB 2. Handle `setjmp` callsites 3. Handle `longjmp` callsites 4. Cleanup and update SSA 1, 3, and 4 are identical for Emscripten SjLj and Wasm SjLj. Only the step 2 is different. This CL extracts the current Emscripten SjLj's longjmp callsites handling into a function. The reason to make this a separate CL is, without this, the diff tool cannot compare things well in the presence of moved code and added code in the followup Wasm SjLj CL, and it ends up mixing them together, making the diff unreadable. Also fixes some typos and variable names. So far we've been calling the buffer argument to `setjmp` and `longjmp` `jmpbuf`, but the name used in the man page for those functions is `env`, so updated them to be consistent. Reviewed By: tlively Differential Revision: https://reviews.llvm.org/D108728	2021-08-25 15:45:38 -07:00
Ricky Taylor	f659b6b1fa	[M68k][NFC] Rename M68kOperand::Kind to KindTy Rename the M68kOperand::Type enumeration to KindTy to avoid ambiguity with the Kind field when referencing enumeration values e.g. `Kind::Value`. This works around a compilation error under GCC 5, where GCC won't lookup enum class values if you have a similarly named field (see https://gcc.gnu.org/bugzilla/show_bug.cgi?id=60994). The error in question is: `M68kAsmParser.cpp:857:8: error: 'Kind' is not a class, namespace, or enumeration` Differential Revision: https://reviews.llvm.org/D108723	2021-08-25 22:24:43 +01:00
Heejin Ahn	c2c9a3fd9c	[WebAssembly] Rename wasm.catch.exn intrinsic back to wasm.catch The plan was to use `wasm.catch.exn` intrinsic to catch exceptions and add `wasm.catch.longjmp` intrinsic, that returns two values (setjmp buffer and return value), later to catch longjmps. But because we decided not to use multivalue support at the moment, we are going to use one intrinsic that returns a single value for both exceptions and longjmps. And even if it's not for that, I now think the naming of `wasm.catch.exn` is a little weird, because the intrinsic can still take a tag immediate, which means it can be used for anything, not only exceptions, as long as that returns a single value. This partially reverts D107405. Reviewed By: tlively Differential Revision: https://reviews.llvm.org/D108683	2021-08-25 14:19:22 -07:00
Nathan Sidwell	ab55cc6cef	[X86] pr51000 in-register struct return tailcalling In-register structure returns are not special, and handled by lowering to multiple-value tuples. We can tail-call from non-sret fns to structure-returning functions, except on i686 where the sret pointer is callee-pop. Differential Revision: https://reviews.llvm.org/D105807	2021-08-25 10:15:50 -07:00
Stanislav Mekhanoshin	11b7ee974a	[AMDGPU] Avoid assert for saved FP With spilling into AGPRs enabled we cannot reliably predict if we need to save FP or not. We may finally spill everything into AGPRs and never touch stack. In this case we still may save FP. This is deficiency but not an error, so avoid the assert. Differential Revision: https://reviews.llvm.org/D107404	2021-08-25 09:50:59 -07:00
Kirill Stoimenov	832aae738b	[asan] Implemented intrinsic for the custom calling convention similar used by HWASan for X86. The implementation uses the int_asan_check_memaccess intrinsic to instrument the code. The intrinsic is replaced by a call to a function which performs the access check. The generated function names encode the input register name as a number using Reg - X86::NoRegister formula. Reviewed By: vitalybuka Differential Revision: https://reviews.llvm.org/D107850	2021-08-25 15:31:46 +00:00
alex-t	ed0f4415f0	[AMDGPU] Divergence-driven compare operations instruction selection Description: This change enables the compare operations to be selected to SALU/VALU form dependent of the SDNode divergence flag. Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D106079	2021-08-25 18:30:49 +03:00
Neumann Hon	6b94777be5	[SystemZ] [NFC] Replace SpecialRegisters field with a unique_ptr instead of a raw pointer. This patch replaces the SpecialRegisters field with a unique_ptr instead of a raw pointer. This is better practice, and allows us to remove the definition of the dtor for the SystemZSubtarget class. Reviewed By: uweigand, Kai Differential Revision: https://reviews.llvm.org/D108639	2021-08-25 11:28:18 -04:00
Andrea Di Biagio	5f848b311f	[X86][SchedModel] Fix latency the Hi register write of MULX (PR51495). Before this patch, WriteIMulH reported a latency value which is correct for the RR variant of MULX, but not for the RM variant. This patch fixes the issue by introducing a new WriteIMulHLd, which is meant to be used only by the RM variant of MULX. Differential Revision: https://reviews.llvm.org/D108701	2021-08-25 16:12:09 +01:00
Thomas Johnson	8c3886b0ec	[ARC] Add ADC (addition with carry) and SBC (subtraction with carry) instructions Differential Revision: https://reviews.llvm.org/D108672	2021-08-25 07:46:15 -07:00
Nicholas Guy	36fcf47fc8	[AArch64] Generate SMOV in place of sext(fmov(...)) A single smov instruction is capable of moving from a vector register while performing the sign-extend during said move, rather than each step being performed by separate instructions. Differential Revision: https://reviews.llvm.org/D108633	2021-08-25 15:23:22 +01:00
Joe Nash	e381833ba5	[AMDGPU] Support global_atomic_fmin/max on gfx10 Makes patterns added for gfx90a usable with the gfx10 versions of the insts. Reviewed By: rampitec Differential Revision: https://reviews.llvm.org/D108654 Change-Id: I86167bf6b4823f975f74ccb619bd6190331ba16b	2021-08-25 09:35:10 -04:00
Daniil Fukalov	48958d02d2	[NFC][AMDGPU] Reduce includes dependencies. 1. Splitted out some parts of R600 target to separate modules/headers. 2. Reduced some include lists in headers. 3. Found and fixed issue with override `GCNTargetMachine::getSubtargetImpl()` and `R600TargetMachine::getSubtargetImpl()` had different return value type than base class. 4. Minor forward declarations cleanup. Reviewed By: foad Differential Revision: https://reviews.llvm.org/D108596	2021-08-25 12:01:55 +03:00
Vang Thao	549f6a819a	[MachineCopyPropagation] Check CrossCopyRegClass for cross-class copys On some AMDGPU subtargets, copying to and from AGPR registers using another AGPR register is not possible. A intermediate VGPR register is needed for AGPR to AGPR copy. This is an issue when machine copy propagation forwards a COPY $agpr, replacing a COPY $vgpr which results in $agpr = COPY $agpr. It is removing a cross class copy that may have been optimized by previous passes and potentially creating an unoptimized cross class copy later on. To avoid this issue, check CrossCopyRegClass if a different register class will be needed for the copy. If so then avoid forwarding the copy when the destination does not match the desired register class and if the original copy already matches the desired register class. Issue seen while attempting to optimize another AGPR to AGPR issue: Live-ins: $agpr0 $vgpr0 = COPY $agpr0 $agpr1 = V_ACCVGPR_WRITE_B32 $vgpr0 $agpr2 = COPY $vgpr0 $agpr3 = COPY $vgpr0 $agpr4 = COPY $vgpr0 After machine-cp: $vgpr0 = COPY $agpr0 $agpr1 = V_ACCVGPR_WRITE_B32 $vgpr0 $agpr2 = COPY $agpr0 $agpr3 = COPY $agpr0 $agpr4 = COPY $agpr0 Machine-cp propagated COPY $agpr0 to replace $vgpr0 creating 3 AGPR to AGPR copys. Later this creates a cross-register copy from AGPR->VGPR->AGPR for each copy when the prior VGPR->AGPR copy was already optimal. Reviewed By: lkail, rampitec Differential Revision: https://reviews.llvm.org/D108011	2021-08-24 21:22:36 -07:00
Thomas Lively	977eeb0c38	[WebAssembly] Fix some UB from `ca541aa319`	2021-08-24 19:44:03 -07:00
Heejin Ahn	d5244fb160	[WebAssembly] Use SSAUpdaterBulk in LowerEmscriptenSjLj We update SSA in two steps in Emscripten SjLj: 1. Rewrite uses of `setjmpTable` and `setjmpTableSize` variables and place `phi`s where necessary, which are updated where we call `saveSetjmp`. 2. Do a whole function level SSA update for all variables, because we split BBs where `setjmp` is called and there are possibly variable uses that are not dominated by a def. (See `955b91c19c/llvm/lib/Target/WebAssembly/WebAssemblyLowerEmscriptenEHSjLj.cpp (L1314-L1324)`) We have been using `SSAUpdater` to do this, but `SSAUpdaterBulk` class was added after this pass was first created, and for the step 2 it looks like a better alternative with a possible performance benefit. Not sure the author is aware of it, but `SSAUpdaterBulk` seems to have a limitation: it cannot handle a use within the same BB as a def but before it. For example: ``` ... = %a + 1 %a = foo(); ``` or ``` %a = %a + 1 ``` The uses `%a` in RHS should be rewritten with another SSA variable of `%a`, most likely one generated from a `phi`. But `SSAUpdaterBulk` thinks all uses of `%a` are below the def of `%a` within the same BB. (`SSAUpdater` has two different functions of rewriting because of this: `RewriteUse` and `RewriteUseAfterInsertions`.) This doesn't affect our usage in the step 2 because that deals with possibly non-dominated uses by defs after block splitting. But it does in the step 1, which still uses `SSAUpdater`. But this CL also simplifies the step 1 by using `make_early_inc_range`, removing the need to advance the iterator before rewriting a use. This is NFC; the test changes are just the order of PHI nodes. Reviewed By: dschuff Differential Revision: https://reviews.llvm.org/D108583	2021-08-24 18:23:51 -07:00
Heejin Ahn	77b921b870	[WebAssembly] Tidy up EH/SjLj options This CL is small, but the description can be a little long because I'm trying to sum up the status quo for Emscripten/Wasm EH/SjLj options. First, this CL adds an option for Wasm SjLj (`-wasm-enable-sjlj`), which handles SjLj using Wasm EH. The implementation for this will be added as a followup CL, but this adds the option first to do error checking. This also adds an option for Wasm EH (`-wasm-enable-eh`), which has been already implemented. Before we used `-exception-model=wasm` as the same meaning as enabling Wasm EH, but after we add Wasm SjLj, it will be possible to use Wasm EH instructions for Wasm SjLj while not enabling EH, so going forward, to use Wasm EH, `opt` and `llc` will need this option. This only affects `opt` and `llc` command lines and does not affect Emscripten user interface. Now we have two modes of EH (Emscripten/Wasm) and also two modes of SjLj (also Emscripten/Wasm). The options corresponding to each of are: - Emscripten EH: `-enable-emscripten-cxx-exceptions` - Emscripten SjLj: `-enable-emscripten-sjlj` - Wasm EH: `-wasm-enable-eh -exception-model=wasm` `-mattr=+exception-handling` - Wasm SjLj: `-wasm-enable-sjlj -exception-model=wasm` `-mattr=+exception-handling` The reason Wasm EH/SjLj's options are a little complicated are `-exception-model` and `-mattr` are common LLVM options ane not under our control. (`-mattr` can be omitted if it is embedded within the bitcode file.) And we have the following rules of the option composition: - Emscripten EH and Wasm EH cannot be turned on at the same itme - Emscripten SjLj and Wasm SjLj cannot be turned on at the same time - Wasm SjLj should be used with Wasm EH Which means we now allow these combinations: - Emscripten EH + Emscripten SjLj: the current default in `emcc` - Wasm EH + Emscripten SjLj: This is allowed, but only as an interim step in which we are testing Wasm EH but not yet have a working implementation of Wasm SjLj. This will error out (D107687) in compile time if `setjmp` is called in a function in which Wasm exception is used. - Wasm EH + Wasm SjLj: This will be the default mode later when using Wasm EH. Currently Wasm SjLj implementation doesn't exist, so it doesn't work. - Emscripten EH + Wasm SjLj will not work. This CL moves these error checking routines to `WebAssemblyPassConfig::addIRPasses`. Not sure if this is an ideal place to do this, but I couldn't find elsewhere. Currently some checking is done within LowerEmscriptenEHSjLj, but these checks only run if LowerEmscriptenEHSjLj runs so it may not run when Wasm EH is used. This moves that to `addIRPasses` and adds some more checks. Currently LowerEmscriptenEHSjLj pass is responsible for Emscripten EH and Emscripten SjLj. Wasm EH transformations are done in multiple places, including WasmEHPrepare, LateEHPrepare, and CFGStackify. But in the followup CL, LowerEmscriptenEHSjLj pass will be also responsible for a part of Wasm SjLj transformation, because WasmSjLj will also be using several Emscripten library functions, and we will be sharing more than half of the transformation to do that between Emscripten SjLj and Wasm SjLj. Currently we have `-enable-emscripten-cxx-exceptions` and `-enable-emscripten-sjlj` but these only work for `llc`, because for `llc` we feed these options to the pass but when we run the pass using `opt` the pass will be created with no options and the default options will be used, which turns both Emscripten EH and Emscripten SjLj on. Now we have one more SjLj option to care for, LowerEmscriptenEHSjLj pass needs a finer way to control these options. This CL removes those default parameters and make LowerEmscriptenEHSjLj pass read directly from command line options specified. So if we only run `opt -wasm-lower-em-ehsjlj`, currently both Emscripten EH and Emscripten SjLj will run, but with this CL, none will run unless we additionally pass `-enable-emscripten-cxx-exceptions` or `-enable-emscripten-sjlj`, or both. This does not affect users; this only affects our `opt` tests because `emcc` will not call either `opt` or `llc`. As a result of this, our existing Emscripten EH/SjLj tests gained one or both of those options in their `RUN` lines. Reviewed By: dschuff Differential Revision: https://reviews.llvm.org/D107685	2021-08-24 17:54:39 -07:00
Thomas Lively	ca541aa319	[WebAssembly] Fix up out-of-range BUILD_VECTOR lane constants Fixes PR51605 in which a DAG combine and legalization sequence generated out-of-range constants in BUILD_VECTOR lanes. In the v16i8 case, the constants were 255, which would be in range if DAG ISel used unsigned constants, but it is out of range because DAG ISel uses signed constants. Differential Revision: https://reviews.llvm.org/D108669	2021-08-24 17:24:03 -07:00
Amara Emerson	2ed8053d46	Revert "[AArch64][GlobalISel] Don't contract cross-bank copies into truncating stores." This reverts commit `67bf3ac744`. The reason is that this change is now superseded by `04fb9b729a` which fixes the underlying problem in the selector. Now it's fine to generate truncating FP stores since the selector code will just generate subreg copies to handle them.	2021-08-24 16:26:56 -07:00
Amara Emerson	04fb9b729a	[AArch64][GlobalISel] Fix incorrect handling of fp truncating stores. When the tablegen patterns fail to select a truncating scalar FPR store, our manual selection code also failed to handle it silently, trying to generate an invalid copy. Fix this by adding support in the manual code to generate a proper subreg copy before selecting a non-truncating store.	2021-08-24 16:07:00 -07:00
Min-Yih Hsu	d2bb6d512c	[X86] Add explicit library dependency on LLVMInstrumentation Patch `9588b685c6` introduced dependency on ASAN. But it didn't explicitly put LLVMInstrumentation as one of the library dependencies such that the build will fail if we're building LLVM as shared libraries (i.e. -DBUILD_SHARED_LIBS=ON). This patch explicitly links X86CodeGen against the Instrumentation component. Differential Revision: https://reviews.llvm.org/D108662	2021-08-24 14:08:17 -07:00
Jessica Paquette	ef8707574b	[AArch64][GlobalISel] Legalize narrow scalar FP arithmetic Widen narrow fp arithmetic ops (e.g. G_FADD). When we don't have full FP16 support, widen to s32. Otherwise widen to s16. https://godbolt.org/z/TbT9Pqa7e Differential Revision: https://reviews.llvm.org/D108660	2021-08-24 13:54:28 -07:00
Patrick Holland	e4ebfb5786	[MCA] Adding an AMDGPUCustomBehaviour implementation. This implementation allows mca to model the desired behaviour of the s_waitcnt instruction. This patch also adds the RetireOOO flag to the AMDGPU instructions within the scheduling model. This flag is only used by mca and allows instructions to finish out-of-order which helps mca's simulations more closely model the actual device. Differential Revision: https://reviews.llvm.org/D104730	2021-08-24 13:33:58 -07:00
Kirill Stoimenov	b97ca3aca1	Revert "[asan] Implemented intrinsic for the custom calling convention similar used by HWASan for X86." This reverts commit `9588b685c6`. Breaks a bunch of builds. Reviewed By: GMNGeoffrey Differential Revision: https://reviews.llvm.org/D108658	2021-08-24 13:21:20 -07:00
Kirill Stoimenov	9588b685c6	[asan] Implemented intrinsic for the custom calling convention similar used by HWASan for X86. The implementation uses the int_asan_check_memaccess intrinsic to instrument the code. The intrinsic is replaced by a call to a function which performs the access check. The generated function names encode the input register name as a number using Reg - X86::NoRegister formula. Reviewed By: vitalybuka Differential Revision: https://reviews.llvm.org/D107850	2021-08-24 19:34:34 +00:00
Thomas Johnson	ce1dc9d647	[ARC] Add codegen for the readcyclecounter intrinsic along with disassembly for associated instructions Differential Revision: https://reviews.llvm.org/D108598	2021-08-24 11:53:20 -07:00
Jessica Paquette	db232de193	[AArch64][GlobalISel] Legalize + select v2p0 -> v264 G_PTRTOINT 1) Just mark this case as legal because it can just be a copy. 2) Ensure the copy in the existing code actually gets selected. Without doing this, we'll crash because the destination won't have a register class. This fell back 35 times in a build of clang with GISel for AArch64. Differential Revision: https://reviews.llvm.org/D108610	2021-08-24 11:02:01 -07:00
Jessica Paquette	67d4dd5c07	[AArch64][GlobalISel] Select @llvm.aarch64.neon.ld4.* Reuse the selection code from the ld2 case. This is similar to how SDAG handles things in AArch64ISelDAGToDAG. (See SelectLoad) This fell back ~100 times while building clang with GISel enabled for AArch64. Factoring out the gross subreg copy part ought to make selecting the rest of this family fairly easy. Differential Revision: https://reviews.llvm.org/D108600	2021-08-24 09:03:49 -07:00
Simon Pilgrim	307890f85b	[X86] Freeze vXi8 shl(x,1) -> add(x,x) vector fold (PR50468) We don't have any vXi8 shift instructions (other than on XOP which is handled separately), so replace the shl(x,1) -> add(x,x) fold with shl(x,1) -> add(freeze(x),freeze(x)) to avoid the undef issues identified in PR50468. Split off from D106675 as I'm still looking at whether we can fix the vXi16/i32/i64 issues with the D106679 alternative. Differential Revision: https://reviews.llvm.org/D108139	2021-08-24 16:08:24 +01:00
Ricky Taylor	47f52f989b	[M68k][AsmParser] Support parsing register masks & fix printing them Fixes PR51580. Register masks will now be printed as 'movem.l (%sp), %a0-%a5/%d5' for example and can now be parsed in the same format. Previously the printed syntax was 'movem.l (%sp), %a0-%a5,%d', which didn't match prior art and was too ambiguous to easily parse. Differential Revision: https://reviews.llvm.org/D108597	2021-08-24 10:40:02 +01:00
Jeremy Morse	992e21eeee	[DebugInfo][InstrRef] Fix over-droppage of locations in X86FloatingPoint Over in D105657, we started dropping instruction numbers (that become variable locations) from call instructions, as we can't correctly represent the x87 FP stack. Unfortunately, it turns out that the "special FP instructions" that this pass transforms includes "every call instruction" [0]. Thus, we've ended up dropping all return values from all calls. Ouch. This patch adds a filter: only drop instruction numbers from calls if they return something on the FP stack. Seeing how LLVM only allows a single return value, this should drop instruction numbers on anything that returns a float, and nothing else. Rather than writing a new test, I've modified the original one to have a positive and negative case: drop instruction number on a call with an FP-stack modification, keep it on a plain call. Differential Revision: https://reviews.llvm.org/D108580	2021-08-24 10:24:07 +01:00
Cullen Rhodes	e9c8973f1c	[AArch64][SME] Fix v8.6a bf16 NEON instruction predication In streaming mode on SME targets only the scalar BFCVT armv8.6-a instruction is legal, predicate the illegal instructions on NEON to disable them in streaming mode (see D107902). BFCVT is predicated on HasNEONorStreamingSVE. The reference can be found here: https://developer.arm.com/documentation/ddi0602/2021-06/SIMD-FP-Instructions Reviewed By: paulwalker-arm Differential Revision: https://reviews.llvm.org/D108279	2021-08-24 08:13:57 +00:00
Martin Storsjö	039b469b85	[ARM] Allow using ';' as asm statement separator in MSVC mode This does the same as D96259, but for ARM, just like AArch64, using the same comment char as for ELF and MinGW mode. As the assembly input/output of LLVM is GAS style, trying to match what MS armasm.exe does isn't needed (because the comment char used is the least concern when it comes to that; all directives differ too). If a separate armasm compatible mode is implemented, it can use its own comment style (just like llvm-ml implements MS ml.exe compatible assembly parsing). This fixes building compiler-rt assembly files for ARM in MSVC mode. The updated testcase literals-comments.s was only intended to make sure that '#' isn't interpreted as a comment char. Differential Revision: https://reviews.llvm.org/D107251	2021-08-24 11:01:49 +03:00
Liu, Chen3	b7795eb646	[X86] Building constant vector which element type is half will cause assertion fail. Fix assertion fail when building con constant vector which element type is half. Differential Revision: https://reviews.llvm.org/D108612	2021-08-24 14:34:30 +08:00
Wang, Pengfei	c728bd5bba	[X86] AVX512FP16 instructions enabling 5/6 Enable FP16 FMA instructions. Ref.: https://software.intel.com/content/www/us/en/develop/download/intel-avx512-fp16-architecture-specification.html Reviewed By: LuoYuanke Differential Revision: https://reviews.llvm.org/D105268	2021-08-24 09:07:19 +08:00
Jessica Paquette	2ec2b25fba	[AArch64][GlobalISel] Select @llvm.aarch64.neon.ld2.* This is pretty similar to the ST2 selection code in `AArch64InstructionSelector::selectIntrinsicWithSideEffects`. This is a GISel equivalent of the ld2 case in `AArch64DAGToDAGISel::Select`. There's some weirdness there that appears here too (e.g. using ld1 for scalar cases, which are 1-element vectors in SDAG.) It's a little gross that we have to create the copy and then select it right after, but I think we'd need to refactor the existing copy selection code quite a bit to do better. This was falling back while building llvm-project with GISel for AArch64. Differential Revision: https://reviews.llvm.org/D108590	2021-08-23 17:15:53 -07:00
Fangrui Song	ba6e15d8cc	[TargetMachine] Move COFF special case for ExternalSymbolSDNode from shouldAssumeDSOLocal to X86Subtarget Intended to be NFC. ARM/AArch64 don't appear to need adjustment. TargetMachine::shouldAssumeDSOLocal is expected to be very simple, ideally matching isDSOLocal(). The IR producers are expected to set dso_local correctly. (While some may think this function can make producers' work easier, the function is really not in a good position to set dso_local. See the various special cases we duplicate from clang CodeGenModule.cpp.) Reviewed By: mstorsjo Differential Revision: https://reviews.llvm.org/D108514	2021-08-23 13:54:40 -07:00
Artem Belevich	49d982d8cb	[CUDA] Add support for CUDA-11.4 Differential Revision: https://reviews.llvm.org/D108239	2021-08-23 13:24:46 -07:00
David Green	50f4ae58eb	[AArch64] Correct store ReadAdrBase operand It appears that the Read operand for stores was being placed on the first operand (the stored value) not the address base. This adds a ReadST for the stored value operand, allowing the ReadAdrBase to correctly act upon the address. Differential Revision: https://reviews.llvm.org/D108287	2021-08-23 21:07:55 +01:00
Zarko Todorovski	b575bbd0c7	[PowerPC][AIX] Set the HasAlloca flag in the AIX Traceback Table only if R31 is used as a frame pointer After `c063946476` usage of R31 doesn't necessarily mean that alloca is used. The `TracebackTable::IsAllocaUsedMask` flag should be set only when R31 is used as a frame pointer. On AIX the `function calls alloca' bit seems to be set whenever R31 is set up as a frame pointer, even when there is no alloca call. Reviewed By: lkail Differential Revision: https://reviews.llvm.org/D108141	2021-08-23 15:20:41 -04:00
Jessica Paquette	a2c8e17658	[AArch64][GlobalISel] Add regbankselect support for G_LLROUND Same as G_LROUND: destination should always be a GPR, source should always be a FPR. Differential Revision: https://reviews.llvm.org/D108566	2021-08-23 10:32:20 -07:00
Jessica Paquette	fe51f9098b	[AArch64][GlobalISel] Legalize G_LLROUND for s64 + s32 Same as G_LROUND. Also add a TODO for full fp16 legalization. Differential Revision: https://reviews.llvm.org/D108564	2021-08-23 09:45:23 -07:00
Florian Hahn	d024a01511	Recommit "[LoopVectorize][AArch64] Enable ordered reductions by default for AArch64" This reverts the revert `ab9296f13b`. The issue causing the revert should be fixed in `9baed023b4`.	2021-08-23 11:25:27 +01:00
Jay Foad	7a967d9011	[AMDGPU] Try to fix a GCC 11 warning Apparently GCC 11 was warning: AMDGPURegisterBankInfo.cpp:2543:33: warning: enumerated and non-enumerated type in conditional expression [-Wextra]	2021-08-23 10:51:37 +01:00
Cullen Rhodes	fb82b836b7	[AArch64][SME] Support NEON scalar FP instructions in streaming mode The following scalar FP instructions are legal in streaming mode: 0101 1110 xx1x xxxx 11x1 11xx xxxx xxxx # FMULX/FRECPS/FRSQRTS (scalar) 0101 1110 x10x xxxx 00x1 11xx xxxx xxxx # FMULX/FRECPS/FRSQRTS (scalar, FP16) 01x1 1110 1x10 0001 11x1 10xx xxxx xxxx # FRECPE/FRSQRTE/FRECPX (scalar) 01x1 1110 1111 1001 11x1 10xx xxxx xxxx # FRECPE/FRSQRTE/FRECPX (scalar, FP16) Predicate them on `HasNEONorStreamingSVE`. Full list of affected instructions: FMULX16, FMULX32, FMULX64, FRECPS16, FRECPS32, FRECPS64, FRSQRTS16, FRSQRTS32, FRSQRTS64, FRECPEv1f16, FRECPEv1i32, FRECPEv1i64, FRECPXv1f16, FRECPXv1i32, FRECPXv1i64, FRSQRTEv1f16, FRSQRTEv1i32, FRSQRTEv1i64 Depends on D107902. The reference can be found here: https://developer.arm.com/documentation/ddi0602/2021-06/SIMD-FP-Instructions Execution of NEON instructions that are illegal in streaming mode will cause a trap or exception. Using FMULX [1] as an example, this check is at the top of the pseudocode: if elements == 1 then CheckFPEnabled64(); else CheckFPAdvSIMDEnabled64(); For the legal scalar variants it calls `CheckFPEnabled64`, whereas for the illegal vector variants it calls `CheckFPAdvSIMDEnabled64` which traps. This is useful for observing which instructions are/aren't legal in streaming mode. [1] https://developer.arm.com/documentation/ddi0602/2021-06/SIMD-FP-Instructions/FMULX--Floating-point-Multiply-extended- Reviewed By: david-arm Differential Revision: https://reviews.llvm.org/D108039	2021-08-23 08:48:34 +00:00
Cullen Rhodes	cf3c6cca9f	[AArch64][SME] Add predicate for NEON support in streaming mode Split out from D107903 to remove dependency for D108039 and D108279. Reviewed By: paulwalker-arm Differential Revision: https://reviews.llvm.org/D108293	2021-08-23 08:48:33 +00:00
Kai Luo	7165e6713f	[PowerPC] Use int64_t to represent stack object offset and frame size This is the first step to enable PPC64 support huge frame size(>2G). Also fix an assertion error for frame size, i.e.,`int x; !isInt<32>(x);` should be always evaluated false, so the guard code for frame size is impossible to hit. Reviewed By: jsji Differential Revision: https://reviews.llvm.org/D107435	2021-08-23 02:13:21 +00:00
Simon Pilgrim	805fb1f6c1	[X86] combineMul - move MUL_IMM comment inside function. NFC. combineMul is now used for other things as well as the mul-with-constant expansion - move the comment to where its actually relevant.	2021-08-22 18:27:03 +01:00
Simon Pilgrim	352df10a23	[X86][AVX] matchShuffleAsBlend - use isElementEquivalent to help match broadcast/repeated elements Extend matchShuffleAsBlend to not only match against known in-place elements for BLEND shuffles, but use isElementEquivalent to determine if the shuffle mask's referenced element is the same as the in-place element. This allows us to replace a number of insertps instructions with more general blendps instructions (better opportunities for commutation, concatenation etc.).	2021-08-22 15:26:17 +01:00
Simon Pilgrim	96fb3eef66	Fix signed/unsigned comparison warning. NFCI.	2021-08-22 15:02:19 +01:00
Simon Pilgrim	a1c892b439	[X86][SSE] lowerVECTOR_SHUFFLE - canonicalize with horizontal ops. Before lowering shuffles, see if we can merge horizontal ops or canonicalize the shuffle mask to point to the same LHS/RHS of the HOps when an HOp's args are repeated.	2021-08-22 14:54:48 +01:00
Simon Pilgrim	8533e782ef	[X86] Try to sync HSW + BDW model class defs to simplify comparisons. NFC. Broadwell is mainly a die shrink of Haswell, but the model had many of the scheduling classes in different orders, making side-by-side comparisons very difficult. The InstRW overrides are still quite different, but at least that part of the side-by-side diff is now in the same position. This was noticed while I was trying to investigate diffs between llvm-mca and other perf analyzers in https://uica.uops.info/ - we used to be able to do diffs between most of the models very easily, but we seem to have lost that simplicity as classes have been altered, models have been refined and other models have rotted.	2021-08-22 13:02:51 +01:00
Ben Shi	f69fb7ac72	[DAGCombiner] Add target hook function to decide folding (mul (add x, c1), c2) Reviewed by: lebedev.ri, spatel, craig.topper, luismarques, jrtc27 Differential Revision: https://reviews.llvm.org/D107711	2021-08-22 16:53:32 +08:00
Wang, Pengfei	b088536ce9	[X86] AVX512FP16 instructions enabling 4/6 Enable FP16 unary operator instructions. Ref.: https://software.intel.com/content/www/us/en/develop/download/intel-avx512-fp16-architecture-specification.html Reviewed By: LuoYuanke Differential Revision: https://reviews.llvm.org/D105267	2021-08-22 08:59:35 +08:00
Fangrui Song	0473e9f41a	[AArch64] Replace unneeded CCAssignToRegWithShadow with CCAssignToReg CCState::AllocateReg handles aliased registers.	2021-08-21 16:33:29 -07:00
Fangrui Song	a83d99c55e	[TargetMachine] Drop special case for *-win32-macho clang CodeGenModule shouldAssumeDSOLocal has set dso_local.	2021-08-21 13:59:17 -07:00
Fangrui Song	c5ee312368	[TargetMachine] Simplify shouldAssumeDSOLocal. NFC	2021-08-21 12:37:29 -07:00
David Green	605489d593	[ARM] Fix VQDMULH fold for scalar smin Add a variant of mve-vqdmulh tests that uses min/max intrinsics directly, including a scalar test that shows it misbehaving for min intrinsics and a fix for the combine to prevent it from misbehaving.	2021-08-21 16:33:18 +01:00
Amara Emerson	3187a4f3f1	[AArch64][GlobalISel] Add legalizer support for the @llvm.get.dynamic.area.offset intrinsic. This is just 0 on AArch64.	2021-08-20 17:13:34 -07:00
Amara Emerson	67bf3ac744	[AArch64][GlobalISel] Don't contract cross-bank copies into truncating stores. Truncating stores with GPR bank sources shouldn't be mutated into using FPR bank sources, since those aren't supported. Ideally this should be a selection failure in the tablegen patterns, but for now avoid generating them.	2021-08-20 16:36:23 -07:00
Jessica Paquette	9e9d70591e	[AArch64][GlobalISel] Legalize non-register-sized scalar G_BITREVERSE Clamp types to [s32, s64] and make them a power of 2. This matches SDAG's behaviour. https://godbolt.org/z/vTeGqf4vT Differential Revision: https://reviews.llvm.org/D108344	2021-08-20 14:44:03 -07:00
Jessica Paquette	7e91c59844	[AArch64][GlobalISel] Legalize 32-bit + narrow G_SMULO + G_UMULO SDAG lowers 32-bit and 64-bit G_SMULO + G_UMULO. We were missing the 32-bit case. For other sizes, make the 0th type a power of 2 and clamp it to either 32 bits or 64 bits. Right now, this will allow us to handle narrow types (e.g. s4, s24, etc.). The LegalizerHelper doesn't support narrowing G_SMULO or G_UMULO right now. I think we want clamping behaviour either way, so we might as well include it now to be explicit. Differential Revision: https://reviews.llvm.org/D108240	2021-08-20 14:37:46 -07:00
Jessica Paquette	16caf6321c	[AArch64][GlobalISel] Clamp vectors of p0 when legalizing G_LOAD/G_STORE We had a rule for <n x s64> but not one for <n x p0>. As a result, we'd fall back on like <5 x p0> or whatever. Differential Revision: https://reviews.llvm.org/D108484	2021-08-20 14:34:49 -07:00
Jessica Paquette	470c74f181	[AArch64][GlobalISel] Add regbankselect support for G_LROUND Destination is always a GPR, since the result is always an integer. Source is always a FPR, since the source is always floating point. Differential Revision: https://reviews.llvm.org/D108419	2021-08-20 14:31:14 -07:00
Jessica Paquette	44bf0dc625	[AArch64][GlobalISel] Mark G_LROUND as legal for s64 dst + s32/s64 src. Matches SDAG's behaviour for these types. Differential Revision: https://reviews.llvm.org/D108420	2021-08-20 14:22:58 -07:00
Florian Hahn	ab9296f13b	Revert "[LoopVectorize][AArch64] Enable ordered reductions by default for AArch64" This reverts commit `f4122398e7` to investigate a crash exposed by it. The patch breaks building the code below with `clang -O2 --target=aarch64-linux` int a; double b, c; void d() { for (; a; a++) { b += c; c = a; } }	2021-08-20 21:24:28 +01:00
Andrea Di Biagio	35d4292a73	[X86][SchedModels] Fix missing ReadAdvance for MULX and ADCX/ADOX (PR51494) Before this patch, instructions MULX32rm and MULX64rm were missing a ReadAdvance for the implicit read of register EDX/RDX. This patch fixes the issue, and it also introduces a new SchedWrite for the two variants of MULX. The general idea behind this last change is to eventually decrease the number of InstRW in the scheduling models. This patch also adds a ReadAdvance for the implicit read of EFLAGS in ADCX/ADOX. Differential Revision: https://reviews.llvm.org/D108372	2021-08-20 17:39:51 +01:00
Thomas Lively	88962cea46	[WebAssembly] Restore builtins and intrinsics for pmin/pmax Partially reverts `85157c0079`, which had removed these builtins and intrinsics in favor of normal codegen patterns. It turns out that it is possible for the patterns to be split over multiple basic blocks, however, which means that DAG ISel is not able to select them to the pmin/pmax instructions. To make sure the SIMD intrinsics generate the correct instructions in these cases, reintroduce the clang builtins and corresponding LLVM intrinsics, but also keep the normal pattern matching as well. Differential Revision: https://reviews.llvm.org/D108387	2021-08-20 09:21:31 -07:00
Ben Shi	5b6c9a5ab0	[RISCV] Optimize add in the zba extension with SHADD Optimize (add x, c) to (SHADD (c>>b), x) if c is not simm12 while (c>>b) is simm12 and c has b trailing zeros. Reviewed By: luismarques Differential Revision: https://reviews.llvm.org/D108193	2021-08-20 22:41:49 +08:00
Simon Pilgrim	9efda541bf	[CostModel][X86] Add costs for f32/f64 scalar and vector types. The f16 half types are still pretty useless as we don't have it as a legal type (we treat them as i16 most of the time)	2021-08-20 14:31:12 +01:00
Tim Northover	3d41ef68e7	AArch64: don't form indexed paired ops if base reg overlaps operands. The registers involved might not be identical, but can still overlap (e.g. "str w0, [x0, #4]!").	2021-08-20 11:39:38 +01:00
Jingu Kang	94c4952951	[AArch64] Enable Upper bound unrolling universally Differential Revision: https://reviews.llvm.org/D105996	2021-08-20 11:25:38 +01:00
Fraser Cormack	5b06cbac11	[RISCV] Fix reporting of incorrect commutable operand indices This patch fixes an issue where RISCV's `findCommutedOpIndices` would incorrectly return the pseudo `CommuteAnyOperandIndex` as a commutable operand index, rather than fixing a specific index. Reviewed By: rogfer01 Differential Revision: https://reviews.llvm.org/D108206	2021-08-20 10:27:15 +01:00
Sebastian Neubauer	f3fe44fa05	[AMDGPU] Fix too many constants with flat scratch Prevent SIFoldOperands from creating SALU instructions with a constant and a frame index. Previously, only one operand was checked to be a frame index, leading to too many constants when flat scratch is enabled and stack offsets are large. Differential Revision: https://reviews.llvm.org/D108368	2021-08-20 08:21:36 +02:00
Anshil Gandhi	508b06699a	[Remarks] [AMDGPU] Emit optimization remarks for atomics generating hardware instructions Produce remarks when atomic instructions are expanded into hardware instructions in SIISelLowering.cpp. Currently, these remarks are only emitted for atomic fadd instructions. Differential Revision: https://reviews.llvm.org/D108150	2021-08-19 20:51:19 -06:00
Amara Emerson	95ac3d15e9	[AArch64][GlobalISel] Add G_VECREDUCE fewerElements support for full scalarization. For some reductions like G_VECREDUCE_OR on AArch64, we need to scalarize completely if the source is <= 64b. This change adds support for that in the legalizer. If the source has a pow-2 num elements, then we can do a tree reduction using the scalar operation in the individual elements. Otherwise, we just create a sequential chain of operations. For AArch64, we only need to scalarize if the input is <64b. If it's great than 64b then we can first do a fewElements step to 64b, taking advantage of vector instructions until we reach the point of scalarization. I also had to relax the verifier checks for reductions because the intrinsics support <1 x EltTy> types, which we lower to scalars for GlobalISel. Differential Revision: https://reviews.llvm.org/D108276	2021-08-19 16:38:52 -07:00
Amara Emerson	a0051f7149	[AArch64][GlobalISel] Fix miscompile of <16 x s8> G_EXTRACT_VECTOR_ELT. When support for copying vector s8 lanes was added recently, this also had the side effect of fixing a fallback for <16 x s8> extracts since both used the same helper. However, there was a bug in another helper to get the regclass for a specific FPR-native type, which was assigning FPR16 to s8 instead of FPR8.	2021-08-19 16:22:32 -07:00
Thomas Lively	be6c49e743	[WebAssembly] Add explicit casts to silence -Wc++11-narrowing	2021-08-19 16:00:07 -07:00
Thomas Lively	fd0557dbf1	[WebAssembly] More convert_low and promote_low codegen The convert_low and promote_low instructions can widen the lower two lanes of a four-lane vector, but we were previously scalarizing patterns that widened lanes besides the low two lanes. The commit adds a shuffle to move the widened lanes into the low lane positions so the convert_low and promote_low instructions can be used instead of scalarizing. Depends on D108266. Differential Revision: https://reviews.llvm.org/D108341	2021-08-19 15:37:12 -07:00
Thomas Lively	b311a040ef	[WebAssembly] Pattern match SIMD convert_low and promote_low during ISel Since the simplest DAG patterns for convert_low and promote_low instructions involved v2i32, v2f32, v4i64, and v4f64 types, which are not legal in the WebAssembly backend and would be eliminated by type legalization, we were previously matching those patterns in a DAG combine before the type legalization stage. However in cases where the vectors were wider than 128 bits, the patterns we matched were not created until the type legalization stage when the wide vectors were split up. Type legalization would continue to eliminate the illegal types we were matching as well, so the code ended up scalarized. To make the ISel for these instructions more robust, match the scalarized patterns rather than the patterns containing illegal types. Add tests with double-wide vectors to show that this works as intended. Fixes PR51098. Depends on D107502. Differential Revision: https://reviews.llvm.org/D108266	2021-08-19 15:24:28 -07:00
Arthur Eubanks	44a3241f10	[NFC] Replace some attribute methods that use confusing indexes	2021-08-19 14:10:26 -07:00
Thomas Lively	b69374ca58	[WebAssembly] Legalize vector types by widening The default legalization of unsupported vector types is to promote the integers in each lane, which leads to extra sign or zero extending and masking when moving data into and out of vectors. Switch our preferred type legalization from the default to vector widening, which keeps the data in the low lanes of the vector rather than in the low bits of each lane. The unused high lanes can be ignored. Half-wide vectors are now loaded from memory into the low 64 bits of the v128 rather than spread out among the lanes. As a result, v128.load64_splat is a much more common operation, so add new patterns to support it. Differential Revision: https://reviews.llvm.org/D107502	2021-08-19 12:07:33 -07:00
Stanislav Mekhanoshin	8d7d89b081	[AMDGPU] Add alias.scope metadata to lowered LDS struct Alias analysis is unable to disambiguate accesses to the structure fields without it unlike distinct variables. As a result we cannot combine ds_read and ds_write operations in a case of any store in between which always considered clobbering. Differential Revision: https://reviews.llvm.org/D108315	2021-08-19 11:40:30 -07:00
Tim Northover	edab411ee6	AArch64: copy all parts of the mem operand across when combining a store In particular we were dropping volatility, which can lead to unwanted transformations.	2021-08-19 18:26:39 +01:00
Owen Anderson	06a4c85890	Use v16i8 rather than v2i64 as the VT for memset expansion on AArch64. This allows the instruction selector to realize that it can directly broadcast the low byte of the memset value, rather than replicating it to a 64-bit GPR before broadcasting. This fixes PR50985. Differential Revision: https://reviews.llvm.org/D108354	2021-08-19 16:54:07 +00:00
David Green	d10f23a25d	[ISel] Expand saddsat and ssubsat via asr and xor This changes the lowering of saddsat and ssubsat so that instead of using: r,o = saddo x, y c = setcc r < 0 s = c ? INTMAX : INTMIN ret o ? s : r into using asr and xor to materialize the INTMAX/INTMIN constants: r,o = saddo x, y s = ashr r, BW-1 x = xor s, INTMIN ret o ? x : r https://alive2.llvm.org/ce/z/TYufgD This seems to reduce the instruction count in most testcases across most architectures. X86 has some custom lowering added to compensate for cases where it can increase instruction count. Differential Revision: https://reviews.llvm.org/D105853	2021-08-19 16:08:07 +01:00
Simon Pilgrim	9419729b6a	[CostModel][X86] Add VPOPCNTDQ/BITALG ctpop costs VPOPCNTDQ + BITALG add ctpop instructions for vXi64/vXi32 + vXi16/vXi8 vector types respectively	2021-08-19 15:40:09 +01:00
Craig Topper	36d8316cc8	[RISCV] Reduce duplicate code for calling SimplifyDemandedBits. This encapsulates the APInt creation and worklist management into a helper function. To keep one common interface I've use Log2_32 in places that previously created a mask by subtracting 1 from a power of 2. Differential Revision: https://reviews.llvm.org/D108324	2021-08-19 07:09:38 -07:00
Matthew Devereau	734708e04f	[AArch64][SVE] Teach cost model that masked loads/stores are cheap Reduce the cost of VLS masked loads/stores to make the vectorizor emit them more frequently.	2021-08-19 13:01:33 +01:00
David Sherwood	f4122398e7	[LoopVectorize][AArch64] Enable ordered reductions by default for AArch64 I have added a new TTI interface called enableOrderedReductions() that controls whether or not ordered reductions should be enabled for a given target. By default this returns false, whereas for AArch64 it returns true and we rely upon the cost model to make sensible vectorisation choices. It is still possible to override the new TTI interface by setting the command line flag: -force-ordered-reductions=true\|false I have added a new RUN line to show that we use ordered reductions by default for SVE and Neon: Transforms/LoopVectorize/AArch64/strict-fadd.ll Transforms/LoopVectorize/AArch64/scalable-strict-fadd.ll Differential Revision: https://reviews.llvm.org/D106653	2021-08-19 09:29:40 +01:00
Jessica Paquette	c22b64ef66	[AArch64][GlobalISel] Don't allow s128 for G_ISNAN getAPFloatFromSize doesn't support s128, so we can't lower this without asserting right now. To fix the buildbots, don't allow any scalars other than s16, s32, and s64.	2021-08-18 13:59:00 -07:00
Jessica Paquette	3d91d5b757	[AArch64][GlobalISel] Mark G_FMINNUM/G_FMAXNUM as floating point opcodes We need to ensure that these end up on FPR to allow imported patterns to select them. This will also ensure that we get good regbank selection when dealing with instructions like G_PHI/G_LOAD/G_STORE which deduce their banks from their uses/users. Differential Revision: https://reviews.llvm.org/D108260	2021-08-18 13:32:19 -07:00
Jessica Paquette	45e1a6bd25	[AArch64][GlobalISel] Legalize scalar G_FMINNUM + G_FMAXNUM For subtargets with full FP16, this is legal for s16, s32, and s64. Without full FP16, it's legal for s32 and s64. For s128, this is a libcall. We also support some vector types, but for now, let's just support scalars. Differential Revision: https://reviews.llvm.org/D108259	2021-08-18 13:30:03 -07:00
Joe Nash	9dbc968ed9	[AMDGPU] Fix atomic float max/min intrinsics Hooked up raw.buffer.atomic.fmin/max.f64 This instruction should be available on GFX6, GFX7, and GFX10. It was implemented for GFX90a with a different name. Added intrinsic def for image_atomic_fmin/fmax; the instruction defs were already there. Reviewed By: arsenm Differential Revision: https://reviews.llvm.org/D108208 Change-Id: I473f98d28b2afbeeb2c27822d9686b5e86634e2f	2021-08-18 14:12:42 -04:00
Craig Topper	3f9b37ccb1	[RISCV] Remove sext_inreg+add/sub/mul/shl isel patterns. Let the sext_inreg be selected to sext.w. Remove unneeded sext.w during PostProcessISelDAG. This gives opportunities for some other isel patterns to match like the ADDIPair or matching mul with immediate to shXadd. This becomes possible after D107658 started selecting W instructions based on users. The sext.w will be considered a W user so isel will often select a W instruction for the sext.w input and we can just remove the sext.w. Otherwise we can combine the sext.w with a ADD/SUB/MUL/SLLI to create a new W instruction in parallel to the the original instruction. Reviewed By: luismarques Differential Revision: https://reviews.llvm.org/D107708	2021-08-18 11:07:11 -07:00
Jessica Paquette	791006fb8c	[GlobalISel] Implement lowering for G_ISNAN + use it in AArch64 GlobalISel equivalent to `TargetLowering::expandISNAN`. Use it in AArch64 and add a testcase. Differential Revision: https://reviews.llvm.org/D108227	2021-08-18 10:54:25 -07:00
Craig Topper	6d7ea597ef	[RISCV] Insert sext_inreg when type legalizing add/sub/mul with constant LHS. We already do this for non-constants RHS. This just removes the special case. I believe the special case may have been needed because the ANY_EXTEND of a constant used to create zero extended constants, but we recently changed that to produce sign extended constants. D107658 is needed to prevent some regressions. Reviewed By: luismarques Differential Revision: https://reviews.llvm.org/D107697	2021-08-18 10:44:25 -07:00
Craig Topper	20e6265873	[RISCV] Improve constant materialization for stores of i16 or i32 negative constants. DAGCombiner::visitStore can clear the upper bits of constants used by stores. This leads prevents them from being recognized as sign extended negative values making them more expensive to materialize. This patch uses the hasAllNBitUsers method from D107658 to make a negative constant if none of the users care about the upper bits. Reviewed By: luismarques Differential Revision: https://reviews.llvm.org/D108052	2021-08-18 10:25:12 -07:00
Craig Topper	d9ba1a9c5c	[RISCV] Teach isel to select ADDW/SUBW/MULW/SLLIW when only the lower 32-bits are used. We normally select these when the root node is a sext_inreg, but SimplifyDemandedBits can sometimes bypass the sext_inreg for some users. This can create situation where sext_inreg+add/sub/mul/shl is selected to a W instruction, and then the add/sub/mul/shl is separately selected to a non-W instruction with the same inputs. This patch tries to detect when it would still be ok to use a W instruction without the sext_inreg by checking the direct users. This can allow the W instruction to CSE with one created for a sext_inreg+add/sub/mul/shl. To minimize complexity and cost of checking, we make no attempt to determine if the CSE will happen and just always use a W instruction when we can. Differential Revision: https://reviews.llvm.org/D107658	2021-08-18 10:22:00 -07:00
Craig Topper	f70238914a	[RISCV] Add zext.h/zext.w to RISCVTTIImpl::getIntImmCostInst. If we have these instructions, we don't need to hoist the immediate for an AND that would match them. Reviewed By: luismarques Differential Revision: https://reviews.llvm.org/D107783	2021-08-18 09:40:40 -07:00
David Sherwood	219d4518fc	[Analysis][AArch64] Make fixed-width ordered reductions slightly more expensive For tight loops like this: float r = 0; for (int i = 0; i < n; i++) { r += a[i]; } it's better not to vectorise at -O3 using fixed-width ordered reductions on AArch64 targets. Although the resulting number of instructions in the generated code ends up being comparable to not vectorising at all, there may be additional costs on some CPUs, for example perhaps the scheduling is worse. It makes sense to deter vectorisation in tight loops. Differential Revision: https://reviews.llvm.org/D108292	2021-08-18 17:01:56 +01:00
Bing1 Yu	ffe58de393	[X86] [AMX] Fix the test case failure caused by D107544. The issue can be duplicated when EXPENSIVE_CHECKS is specified for llvm build. Thank Simon report this issue at https://bugs.llvm.org/show_bug.cgi?id=51513. We need return correct value for the changed IR. Reviewed By: RKSimon, LuoYuanke Differential Revision: https://reviews.llvm.org/D108269	2021-08-18 22:27:22 +08:00
Tim Northover	8eb054a87d	AArch64: compare correct type for multi-valued SDNode. If Orig produces more than one value (rare) with different types (rarer) then we need to make sure we check against the one that Orig actually represents, not just the first type. Unfortunately because of the combination of things that need to happen I wasn't able to produce a test.	2021-08-18 09:35:31 +01:00
Amara Emerson	284006079e	[AArch64][GlobalISel] Add support for selection of s8:fpr = G_UNMERGE <8 x s8>	2021-08-18 00:34:06 -07:00

... 3 4 5 6 7 ...

64273 Commits