llvm-project

Commit Graph

Author	SHA1	Message	Date
Justin Lebar	2e4ecfdebe	[CUDA] Implement __ldg using intrinsics. Summary: Previously it was implemented as inline asm in the CUDA headers. This change allows us to use the [addr+imm] addressing mode when executing ld.global.nc instructions. This translates into a 1.3x speedup on some benchmarks that call this instruction from within an unrolled loop. Reviewers: tra, rsmith Subscribers: jhen, cfe-commits, jholewinski Differential Revision: http://reviews.llvm.org/D19990 llvm-svn: 270150	2016-05-19 22:49:13 +00:00
Benjamin Kramer	504c01cc67	Don't rely on value numbers in test, those are fragile and change in Release (no asserts) builds. llvm-svn: 270085	2016-05-19 17:57:35 +00:00
Artem Belevich	ffa5fc51b8	[CUDA] Allow sm_50,52,53 GPUs LLVM accepts them since r233575. Differential Revision: http://reviews.llvm.org/D20405 llvm-svn: 270084	2016-05-19 17:47:47 +00:00
Simon Pilgrim	9b3729b043	[X86][SSE] Sync with llvm/test/CodeGen/X86/sse-intrinsics-fast-isel.ll sse-builtins.c now just covers SSE1 intrinsics llvm-svn: 270083	2016-05-19 17:11:31 +00:00
Simon Pilgrim	bcf8846be5	[X86][SSE2] Fixed shuffle of results in _mm_cmpnge_sd/_mm_cmpngt_sd tests llvm-svn: 270079	2016-05-19 16:48:59 +00:00
Ranjeet Singh	b631aafee3	[ARM] Fix cdp intrinsic - Fixed cdp intrinsic to only accept compile time constant values previously you could pass in a variable to the builtin which would result in illegal llvm assembly output Differential Revision: http://reviews.llvm.org/D20394 llvm-svn: 270058	2016-05-19 13:04:34 +00:00
Michael Zuckerman	178113e8cc	[Clang][AVX512][intrinsics] continue completing missing set intrinsics Differential Revision: http://reviews.llvm.org/D20160 llvm-svn: 270047	2016-05-19 12:07:49 +00:00
Simon Pilgrim	97728dfb39	[X86][SSE2] Added _mm_move_* tests llvm-svn: 270043	2016-05-19 11:18:49 +00:00
Simon Pilgrim	cddcd2bd45	[X86][SSE2] Added _mm_cast* and _mm_set* tests llvm-svn: 270042	2016-05-19 11:03:48 +00:00
Simon Pilgrim	3f64bb9618	[X86][SSE2] Sync with llvm/test/CodeGen/X86/sse2-intrinsics-fast-isel.ll llvm-svn: 270034	2016-05-19 09:52:59 +00:00
Simon Pilgrim	063c57c1f9	Revert r269967 (SSE2 builtin checks) due to failed buildbots llvm-svn: 269970	2016-05-18 18:22:20 +00:00
Simon Pilgrim	8beed747ce	[X86][SSE2] Sync with llvm/test/CodeGen/X86/sse2-intrinsics-fast-isel.ll llvm-svn: 269967	2016-05-18 18:12:34 +00:00
Michael Zuckerman	2cacc35343	[Clang][AVX512] completing missing intrinsics [pandnd]. Differential Revision: http://reviews.llvm.org/D20101 llvm-svn: 269939	2016-05-18 15:25:53 +00:00
Krzysztof Parzyszek	e0026e4e21	[Hexagon] Recognize "q" and "v" in inline-asm as register constraints Clang follow-up to r269933. llvm-svn: 269934	2016-05-18 14:56:14 +00:00
Simon Pilgrim	a090864762	Removed duplicate SSE42 builtin tests from avx-builtins.c llvm-svn: 269932	2016-05-18 14:32:16 +00:00
Simon Pilgrim	519c78f3ae	[X86][SSE42] Sync with llvm/test/CodeGen/X86/sse42-intrinsics-fast-isel.ll llvm-svn: 269931	2016-05-18 14:29:55 +00:00
Simon Pilgrim	7a4d7d47c9	[X86][SSE41] Sync with llvm/test/CodeGen/X86/sse41-intrinsics-fast-isel.ll llvm-svn: 269926	2016-05-18 13:47:16 +00:00
Simon Pilgrim	7e148a94a4	[X86][SSE3] Sync with llvm/test/CodeGen/X86/sse3-intrinsics-fast-isel.ll llvm-svn: 269921	2016-05-18 13:17:39 +00:00
Ashutosh Nema	51c9dd0081	Add new intrinsic support for MONITORX and MWAITX instructions Summary: MONITORX/MWAITX instructions provide similar capability to the MONITOR/MWAIT pair while adding a timer function, such that another termination of the MWAITX instruction occurs when the timer expires. The presence of the MONITORX and MWAITX instructions is indicated by CPUID 8000_0001, ECX, bit 29. The MONITORX and MWAITX instructions are intercepted by the same bits that intercept MONITOR and MWAIT. MONITORX instruction establishes a range to be monitored. MWAITX instruction causes the processor to stop instruction execution and enter an implementation-dependent optimized state until occurrence of a class of events. Opcode of MONITORX instruction is "0F 01 FA". Opcode of MWAITX instruction is "0F 01 FB". These opcode information is used in adding tests for the disassembler. These instructions are enabled for AMD's bdver4 architecture. Patch by Ganesh Gopalasubramanian! Reviewers: echristo, craig.topper Subscribers: RKSimon, joker.eph, llvm-commits, cfe-commits Differential Revision: http://reviews.llvm.org/D19796 llvm-svn: 269907	2016-05-18 11:56:23 +00:00
Craig Topper	39c871038a	[X86] Add immediate range checks for many of the builtins. This time allow -128 to 255 for builtins that use a char type immediate." llvm-svn: 269878	2016-05-18 03:18:12 +00:00
Simon Pilgrim	2d1decf7cb	[X86][SSE] Tidied up MMX/SSE/SSE2 builtin tests to the correct test file llvm-svn: 269852	2016-05-17 22:03:31 +00:00
Filipe Cabecinhas	09fbfcafc3	Revert "[X86] Add immediate range checks for many of the builtins." This reverts commit r269619. llvm-svn: 269765	2016-05-17 14:07:43 +00:00
Craig Topper	dbbe4a5542	[AVX512] Fix return types in several test cases to match the intrinsic they're testing. llvm-svn: 269738	2016-05-17 04:41:32 +00:00
Craig Topper	8ca5373c72	[X86] Fix a few intrinsic tests to use the return type that matches the intrinsic they're testing. llvm-svn: 269735	2016-05-17 03:42:37 +00:00
Michael Zuckerman	bf05a4589e	[Clang][AVX512] completing missing intrinsics for [vpabs] instruction set Differential Revision: http://reviews.llvm.org/D20069 llvm-svn: 269680	2016-05-16 18:57:24 +00:00
Nico Weber	379a1952b3	[ms] Reintroduce feature guards in intrinsic headers in Microsoft mode Visual Studio's C++ standard library headers include intrin.h, so the intrinsic headers get included a lot more often in Microsoft mode than elsewhere. The AVX512 intrinsics are a lot of code (0.7 MB, causing 30% compile time overhead for small programs including e.g. <string> and 6% compile time overhead for larger projects like e.g. v8). Since multiversioning can't be relied on in Microsoft mode (cl.exe doesn't support it), having faster compiles seems like the much better tradeoff until we have a better intrinsic story going forward (which we'll need for e.g. PR19898). Actually using intrinsics on Windows already requires the right /arch: settings, so this patch should have no big behavior change. See also thread "The intrinsics headers (especially avx512) are too big. What to do about it?" on cfe-dev. http://reviews.llvm.org/D20291 llvm-svn: 269675	2016-05-16 18:14:07 +00:00
Michael Zuckerman	cb85677471	[Clang][AVX512] completing missing intrinsics [vsqrt\|vrsqrt\|vrcp14 ]. Differential Revision: http://reviews.llvm.org/D20068 llvm-svn: 269649	2016-05-16 11:42:01 +00:00
Craig Topper	9c6c85f1ad	[AVX512] Add typecasts to some intrinsics to avoid doing operations on the __m512/__m512i/__m512d types. llvm-svn: 269631	2016-05-16 06:38:36 +00:00
Craig Topper	e5cc18054a	[AVX512] Use correct types in test case. llvm-svn: 269622	2016-05-16 01:09:19 +00:00
Craig Topper	0f7ea93541	[X86] Add immediate range checks for many of the builtins. llvm-svn: 269619	2016-05-15 22:18:00 +00:00
Craig Topper	dca1f230ae	[AVX512] Add intrinsics for 512-bit insertf32x8/insertf32x4/inserti32x4. llvm-svn: 269617	2016-05-15 21:26:20 +00:00
Oleg Ranevskyy	7232f66051	[CodeGen] Clang does not choose aapcs-vfp calling convention for ARM bare metal target with hard float (EABIHF) Summary: Clang does not detect `aapcs-vfp` for the EABIHF environment. The reason is that only GNUEABIHF is considered while choosing calling convention, EABIHF is ignored. This causes clang to use `aapcs` for EABIHF and add the `arm_aapcscc` specifier to functions in generated IR. The modified `arm-cc.c` test checks that no calling convention specifier is added to functions for EABIHF, which means the default one is used (`CallingConv::ARM_AAPCS_VFP`). Reviewers: rengolin, compnerd, t.p.northover Subscribers: aemerson, rengolin, asl, cfe-commits Differential Revision: http://reviews.llvm.org/D20219 llvm-svn: 269419	2016-05-13 14:45:57 +00:00
Filipe Cabecinhas	ab731f7e86	[ubsan] Add -fsanitize-undefined-strip-path-components=N Summary: This option allows the user to control how much of the file name is emitted by UBSan. Tuning this option allows one to save space in the resulting binary, which is helpful for restricted execution environments. With a positive N, UBSan skips the first N path components. With a negative N, UBSan only keeps the last N path components. Reviewers: rsmith Subscribers: cfe-commits Differential Revision: http://reviews.llvm.org/D19666 llvm-svn: 269309	2016-05-12 16:51:36 +00:00
Michael Zuckerman	13d3c002df	[clang][AVX512] completing missing set intrinsics Differential Revision: http://reviews.llvm.org/D20099 llvm-svn: 269172	2016-05-11 11:41:29 +00:00
Michael Zuckerman	5e2c6b6200	[clang][AVX512] completing missing intrinsics for [vpermt2d\|vptestm] instruction set. Differential Revision: http://reviews.llvm.org/D20096 llvm-svn: 269170	2016-05-11 11:21:18 +00:00
NAKAMURA Takumi	d4fbaef2b0	clang/test/CodeGen/avx512f-builtins.c: Fix for -Asserts. llvm-svn: 269079	2016-05-10 17:16:12 +00:00
Michael Zuckerman	e9e8e573e3	[Clang][AVX512] completing missing intrinsics [load/store] Differential Revision: http://reviews.llvm.org/D20063 llvm-svn: 269056	2016-05-10 13:13:54 +00:00
Michael Zuckerman	de860e5585	[Clang][AVX512] completing missing intrinsics [vmin/vmax]{sd\|sq\|uq\|ud}. Differential Revision: http://reviews.llvm.org/D20064 llvm-svn: 269042	2016-05-10 11:34:19 +00:00
Michael Zuckerman	2564d2f5fe	[Clang][AVX512] completing missing intrinsics [vextractf]. Differential Revision: http://reviews.llvm.org/D20061 llvm-svn: 269037	2016-05-10 10:14:50 +00:00
Michael Zuckerman	7360d8a9cc	[Clang][AVX512] completing missing intrinsics [roundscale, ceil, floor] Differential Revision: http://reviews.llvm.org/D20070 llvm-svn: 269022	2016-05-10 07:30:58 +00:00
George Burgess IV	3dc1669133	[Sema] Fix an overload resolution bug with enable_if. Currently, if clang::isBetterOverloadCandidate encounters an enable_if attribute on either candidate that it's inspecting, it will ignore all lower priority attributes (e.g. pass_object_size). This is problematic in cases like: ``` void foo(char c) __attribute__((enable_if(1, ""))); void foo(char c __attribute__((pass_object_size(0)))) __attribute__((enable_if(1, ""))); ``` ...Because we would ignore the pass_object_size attribute in the second `foo`, and consider any call to `foo` to be ambiguous. This patch makes overload resolution consult further tiebreakers (e.g. pass_object_size) if two candidates have equally good enable_if attributes. llvm-svn: 269005	2016-05-10 01:59:34 +00:00
Michael Zuckerman	f9be3bb1d5	[clang][AVX512] completing missing intrinsics [vmin/vmax]. Differential Revision: http://reviews.llvm.org/D20062 llvm-svn: 268910	2016-05-09 12:38:49 +00:00
Michael Zuckerman	f15447537f	[Clang][AVX512] completing missing intrinsics [CVT] Differential Revision: http://reviews.llvm.org/D20056 llvm-svn: 268903	2016-05-09 10:32:51 +00:00
Krzysztof Parzyszek	09ba254f10	[Hexagon] Add a testcase for __builtin_HEXAGON_A2_tfrpi llvm-svn: 268637	2016-05-05 15:55:54 +00:00
Marcin Koscielnicki	b31ee6db11	[SystemZ] Add -mbackchain option. This option, like the corresponding gcc option, is SystemZ-specific and enables storing frame backchain links, as specified in the ABI. Differential Revision: http://reviews.llvm.org/D19891 llvm-svn: 268575	2016-05-04 23:37:40 +00:00
Michael Zuckerman	e6f7389b5a	[Clang][Builtin][AVX512] Adding intrinsics fot cvt{u}si2s{d\|s} cvt{sd\|ss}2{ss\|sd} instruction set Differential Revision: http://reviews.llvm.org/D19765 llvm-svn: 268481	2016-05-04 08:55:11 +00:00
Reid Kleckner	8195f696e4	[X86] Add -malign-double support The -malign-double flag causes i64 and f64 types to have alignment 8 instead of 4. On x86-64, the behavior of -malign-double is enabled by default. Rebases and cleans phosek's work here: http://reviews.llvm.org/D12860 Patch by Sean Klein Reviewers: rnk Subscribers: rnk, jfb, dschuff, phosek Differential Revision: http://reviews.llvm.org/D19734 llvm-svn: 268473	2016-05-04 02:58:24 +00:00
Pete Cooper	71dfcb42eb	Change test to use regex instead of explicit value numbers. NFC. We were seeing an internal failure when running this test. I can't see a good reason for the difference, but the simple fix is to use %{{.*}} instead of %1. llvm-svn: 268416	2016-05-03 18:32:01 +00:00
Michael Zuckerman	c66770313a	[clang][AVX512][BuiltIn] Adding intrinsics for cast{pd\|ps\|si}128_{pd\|ps\|si}512 and castsi256_si512 instruction set Differential Revision: http://reviews.llvm.org/D19858 llvm-svn: 268387	2016-05-03 14:26:52 +00:00
Michael Zuckerman	e871785eb6	[Clang][avx512][Builtin] Adding intrinsics for cvtw2mask{128\|256\|512} instruction set Differential Revision: http://reviews.llvm.org/D19766 llvm-svn: 268385	2016-05-03 14:12:23 +00:00

1 2 3 4 5 ...

3630 Commits