llvm-project

Commit Graph

Author	SHA1	Message	Date
Eric Christopher	c5a85af3b2	Cache the Function dependent subtarget on the MachineFunction. As preparation for removing the getSubtargetImpl() call from TargetMachine go ahead and flip the switch on caching the function dependent subtarget and remove the bare getSubtargetImpl call from the X86 port. As part of this add a few tests that show we can generate code and assemble on X86 based on features/cpu on the Function. llvm-svn: 232879	2015-03-21 03:13:10 +00:00
Eric Christopher	cba722f8c1	Grab the cached subtarget off of the MachineFunction. llvm-svn: 232878	2015-03-21 03:13:07 +00:00
Eric Christopher	948bdf996b	Grab a subtarget off of a MipsTargetMachine rather than a bare target machine in preparation for the TargetMachine bare getSubtarget/getSubtargetImpl calls going away. llvm-svn: 232877	2015-03-21 03:13:05 +00:00
Eric Christopher	5c3dffc459	Simplify the query for a subtarget in the NVPTX pass manager. llvm-svn: 232876	2015-03-21 03:13:03 +00:00
Eric Christopher	cd53d6eda7	Change getISAEncoding to use the target triple to determine thumb-ness similar to the rest of the Module level asm printing infrastructure as debug info finalization happens after the function may be missing. llvm-svn: 232875	2015-03-21 03:13:01 +00:00
Eric Christopher	23a7d1e6f4	Make the Hexagon ISelDAGToDAG pass set the subtarget dynamically on each runOnMachineFunction invocation. llvm-svn: 232874	2015-03-21 03:12:59 +00:00
Ahmed Bougacha	e6bb09ac3f	[AArch64] Prefer UZP for concat_vector of illegal truncs. Follow-up to r232459: prefer a UZP shuffle to the intermediate truncs. llvm-svn: 232871	2015-03-21 01:08:39 +00:00
Sanjay Patel	c88f724fed	[X86] Prefer blendps over insertps codegen for one special case With this patch, for this one exact case, we'll generate: blendps %xmm0, %xmm1, $1 instead of: insertps %xmm0, %xmm1, $0 If there's a memory operand available for load folding and we're optimizing for size, we'll still generate the insertps. The detailed performance data motivation for this may be found in D7866; in summary, blendps has 2-3x throughput vs. insertps on widely used chips. Differential Revision: http://reviews.llvm.org/D8332 llvm-svn: 232850	2015-03-20 21:19:52 +00:00
Benjamin Kramer	063667cea2	X86: Make helper functions static. NFC. llvm-svn: 232848	2015-03-20 21:07:30 +00:00
Rafael Espindola	36a15cb975	Don't declare all text sections at the start of the .s The code this patch removes was there to make sure the text sections went before the dwarf sections. That is necessary because MachO uses offsets relative to the start of the file, so adding a section can change relaxations. The dwarf sections were being printed at the start just to produce symbols pointing at the start of those sections. The underlying issue was fixed in r231898. The dwarf sections are now printed when they are about to be used, which is after we printed the text sections. To make sure we don't regress, the patch makes the MachO streamer assert if CodeGen puts anything unexpected after the DWARF sections. llvm-svn: 232842	2015-03-20 20:00:01 +00:00
Rafael Espindola	bdfbde56e0	Reorganize the x86 ELF relocation selection logic. The main differences are: * Split in 32 and 64 bit functions. * First switch on the Modifier so that we have only one non fully covered switch. * Map the fixup kind first to a x86_64 (or i386) specific enum, to make it easy to handle cases like X86::reloc_riprel_4byte_movq_load. * Switch on IsPCRel last, which reduces code duplication. Fixes pr22308. llvm-svn: 232837	2015-03-20 19:48:54 +00:00
John Brawn	1f26a47630	[ARM] Fix handling of thumb1 out-of-range frame offsets LocalStackSlotPass assumes that isFrameOffsetLegal doesn't change its answer when the base register changes. Unfortunately this isn't true in thumb1, where SP-based loads allow a larger offset than non-SP-based loads, and this causes the base register reuse code to generate instructions that are unencodable, causing an assertion failure. Solve this by adding a BaseReg parameter to isFrameOffsetLegal, which ARMBaseRegisterInfo can then make use of to give the correct answer. Differential Revision: http://reviews.llvm.org/D8419 llvm-svn: 232825	2015-03-20 17:20:07 +00:00
Simon Pilgrim	180cad2e57	Stripped trailing whitespace. NFC. llvm-svn: 232822	2015-03-20 16:08:17 +00:00
Tom Stellard	3b0dab9f3f	R600/SI: Refactor VOP2 instruction defs llvm-svn: 232817	2015-03-20 15:14:23 +00:00
Tom Stellard	23c2c3d0f4	R600/SI: Refactor VOP1 instruction defs llvm-svn: 232816	2015-03-20 15:14:21 +00:00
Rafael Espindola	8c8d15879f	Reduce indentation after return. NFC. llvm-svn: 232814	2015-03-20 14:33:25 +00:00
Rafael Espindola	2d74274017	Use early returns. NFC. llvm-svn: 232813	2015-03-20 14:23:46 +00:00
Rafael Espindola	9e77cba164	Fold a llvm_unreachable into an assert. NFC. llvm-svn: 232811	2015-03-20 13:50:15 +00:00
Rafael Espindola	5f5e24bb92	clang-format a function. NFC. llvm-svn: 232810	2015-03-20 13:47:40 +00:00
Craig Topper	3a8eb896c9	[Tablegen] Attempt to add support for patterns containing nodes with multiple results. This is needed for AVX512 masked scatter/gather support. The R600 change is necessary to remove a hack that was working around the lack of multiple results. llvm-svn: 232798	2015-03-20 05:09:06 +00:00
Alexei Starovoitov	f049a68d78	[bpf] fix build fix BPF backend build broken by r232699 llvm-svn: 232795	2015-03-20 02:35:29 +00:00
Sanjay Patel	803fb7c85c	move insert, extract, concat helper functions closer to related helper functions; NFCI llvm-svn: 232781	2015-03-19 23:04:25 +00:00
Eric Christopher	12cf76fe26	Add an MCSubtargetInfo variable to the TargetMachine. This enables us to remove calls to the subtarget from the TargetMachine and with a small hack for backends that require global subtarget information for module level code generation, e.g. mips abi flags, as mentioned in a fixme in the code. llvm-svn: 232776	2015-03-19 22:36:37 +00:00
Eric Christopher	72e23a219c	Add a TargetMachine local MCRegisterInfo and MCInstrInfo so that they can be used without a subtarget in constructing subtarget independent passes. llvm-svn: 232775	2015-03-19 22:36:32 +00:00
Sanjay Patel	d5c2d287f9	[X86, AVX] use blends instead of insert128 with index 0 Another case of x86-specific shuffle strength reduction: avoid generating insert*128 instructions with index 0 because they are slower than their non-lane-changing blend equivalents. Shuffle lowering already catches most of these cases, but the zero vector case and some other paths such as in the modified test in vector-shuffle-256-v32.ll were getting through. Differential Revision: http://reviews.llvm.org/D8366 llvm-svn: 232773	2015-03-19 22:29:40 +00:00
Artem Belevich	9e8a039318	Add support for __nvvm_reflect changes in libdevice in CUDA-7.0 Summary: CUDA 7.0's libdevice uses slightly different IR to call __nvvm_reflect and that triggers an assertion in nvvm_reflect optimization pass. This change allows nvvm_reflect pass to deal with both old and new ways to pass an argument to __nvvm_reflect. Test Plan: ninja check-all Reviewers: eliben, echristo Subscribers: jholewinski, llvm-commits Differential Revision: http://reviews.llvm.org/D8399 llvm-svn: 232732	2015-03-19 17:05:35 +00:00
Krzysztof Parzyszek	421133470f	[Hexagon] Add support for vector instructions llvm-svn: 232728	2015-03-19 16:33:08 +00:00
Krzysztof Parzyszek	c6f19333cf	[Hexagon] ENDLOOP is a non-reversible conditional branch llvm-svn: 232725	2015-03-19 15:18:57 +00:00
Daniel Sanders	b1fbacab5f	[sparc] Small fix to r232719 to make 2007-12-17-InvokeAsm.ll pass on the buildbot. llvm-svn: 232720	2015-03-19 11:27:23 +00:00
Daniel Sanders	f5d1110075	[sparc] Only support the 'm' inline assembly memory constraint. NFC. Summary: SPARC doesn't seem to support any additional constraints. Therefore remove the target hook. No functional change intended. Reviewers: venkatra Subscribers: llvm-commits Differential Revision: http://reviews.llvm.org/D8214 llvm-svn: 232719	2015-03-19 11:26:05 +00:00
Rafael Espindola	cd584a809d	Split the object streamer callback in one per file format. There are two main advantages to doing this * Targets that only need to handle one of the formats specially don't have to worry about the others. For example, x86 now only registers a constructor for the COFF streamer. * Changes to the arguments passed to one format constructor will not impact the other formats. llvm-svn: 232699	2015-03-19 01:50:16 +00:00
Rafael Espindola	69244c3e78	two or more, use a for. llvm-svn: 232688	2015-03-18 23:15:49 +00:00
Simon Pilgrim	5ec5c9cafe	[X86][SSE] Avoid scalarization of v2i64 vector shifts (REAPPLIED) Fixed broken tests. Differential Revision: http://reviews.llvm.org/D8416 llvm-svn: 232682	2015-03-18 22:18:51 +00:00
Bill Schmidt	1723525e01	[PowerPC] Correct typo in PPCInstrAltivec.td llvm-svn: 232681	2015-03-18 22:13:03 +00:00
Eric Christopher	050f590a0c	Revert "[X86][SSE] Avoid scalarization of v2i64 vector shifts" as it appears to have broken tests/bots. This reverts commit r232660. llvm-svn: 232670	2015-03-18 21:01:00 +00:00
Eric Christopher	5ac4e120be	Revert "Add a TargetMachine local MCRegisterInfo and MCInstrInfo so that" Committed too early. This reverts commit r232666. llvm-svn: 232667	2015-03-18 20:41:44 +00:00
Eric Christopher	4e80e18e78	Add a TargetMachine local MCRegisterInfo and MCInstrInfo so that they can be used without a subtarget in constructing subtarget independent passes. llvm-svn: 232666	2015-03-18 20:37:36 +00:00
Eric Christopher	a0de253d27	Revert "Migrate the AArch64 TargetRegisterInfo to its TargetMachine" as we don't necessarily need to do this yet - though we could move the base class to the TargetMachine as it isn't subtarget dependent. This reverts commit r232103. llvm-svn: 232665	2015-03-18 20:37:30 +00:00
Simon Pilgrim	5c837edc2a	[X86][SSE] Avoid scalarization of v2i64 vector shifts Currently v2i64 vectors shifts (non-equal shift amounts) are scalarized, costing 4 x extract, 2 x x86-shifts and 2 x insert instructions - and it gets even more awkward on 32-bit targets. This patch separately shifts the vector by both shift amounts and then shuffles the partial results back together, costing 2 x shuffles and 2 x sse-shifts instructions (+ 2 movs on pre-AVX hardware). Note - this patch only improves the SHL / LSHR logical shifts as only these are supported in SSE hardware. Differential Revision: http://reviews.llvm.org/D8416 llvm-svn: 232660	2015-03-18 19:35:31 +00:00
Rafael Espindola	105270f68c	Add a default implementation of createObjectStreamer. This removes duplicated code from backends that don't need to do anything fancy. llvm-svn: 232658	2015-03-18 19:08:20 +00:00
Krzysztof Parzyszek	36ccfa5779	[Hexagon] Use pseudo-instructions for true/false predicate values llvm-svn: 232657	2015-03-18 19:07:53 +00:00
Krzysztof Parzyszek	7a9cd80f54	Revert "[Hexagon] Use pseudo-instructions for true/false predicate values" This reverts r232650. Missed a piece of code in the previous commit. llvm-svn: 232656	2015-03-18 18:50:06 +00:00
Rafael Espindola	38438bae21	Handle X86::reloc_riprel_4byte in 32 bits mode. We can get there with .code64. Fixes pr22349. llvm-svn: 232651	2015-03-18 17:33:40 +00:00
Krzysztof Parzyszek	5d7e8fcd52	[Hexagon] Use pseudo-instructions for true/false predicate values llvm-svn: 232650	2015-03-18 17:20:51 +00:00
Krzysztof Parzyszek	47ab1f2007	[Hexagon] Intrinsics for circular and bit-reversed loads and stores llvm-svn: 232645	2015-03-18 16:23:44 +00:00
Krzysztof Parzyszek	78cc36fed7	[Hexagon] Handle ENDLOOP0 in InsertBranch and RemoveBranch llvm-svn: 232643	2015-03-18 15:56:43 +00:00
John Brawn	0dbcd65442	[ARM] Align stack objects passed to memory intrinsics Memcpy, and other memory intrinsics, typically tries to use LDM/STM if the source and target addresses are 4-byte aligned. In CodeGenPrepare look for calls to memory intrinsics and, if the object is on the stack, 4-byte align it if it's large enough that we expect that memcpy would want to use LDM/STM to copy it. Differential Revision: http://reviews.llvm.org/D7908 llvm-svn: 232627	2015-03-18 12:01:59 +00:00
Kai Nacke	c44b162c4b	[mips] Add itineraries for ext and ins instructions. Currently, there are no itineraries defined for ext and ins instructions. This patch adds these itineraries and uses them in the instruction definitions. Reviewed By: dsanders Differential Revision: http://reviews.llvm.org/D7209 llvm-svn: 232613	2015-03-18 06:28:38 +00:00
Alexei Starovoitov	6d93de6439	[bpf] fix build fix BPF backend build broken by r232429 Patch by Brenden Blanco llvm-svn: 232581	2015-03-18 01:39:40 +00:00
Krzysztof Parzyszek	8c1cab9a27	Generate bit manipulation instructions on Hexagon llvm-svn: 232577	2015-03-18 00:43:46 +00:00

1 2 3 4 5 ...

32448 Commits