llvm-project

Commit Graph

Author	SHA1	Message	Date
Tim Northover	0f6c9d0a9b	ARM NEON: add vcvtX (with rounding mode) intrinsics to v8 ARM. These instructions (well, the f32 ones) are supported on 32-bit ARMv8, not just AArch64. Now that the arm_neon.td refactoring is complete, adding them is surprisingly simple. rdar://problem/16035743 llvm-svn: 201661	2014-02-19 10:37:13 +00:00
Tim Northover	1994fa7d3d	ARM & AArch64 NEON: share the vabs implementation. This changes ARM to use @llvm.fabs for floating-point vabs. Patterns already existed in the backend, and it might help mid-end phases since it's more likely to be understood than @llvm.arm.neon.vabs. llvm-svn: 201313	2014-02-13 10:44:17 +00:00
Tim Northover	02b438754c	AArch64: share slgihtly more NEON implementation with ARM. The s64/u64 vcvt conversion operations are actually pretty much identical to the s32/u32 ones in implementation, and can be shared with just one extra variable. llvm-svn: 201145	2014-02-11 11:27:44 +00:00
Tim Northover	d23fc6cceb	ARM: move vshll NEON implementation to common code Now that both ARM backends use the same implementation for vshll operations, the code can be shared. This is also a necessary LLVM/Clang interface update. llvm-svn: 201094	2014-02-10 16:20:36 +00:00
Tim Northover	a2e0a27d26	ARM: implement vshrn NEON intrinsic in terms of shr/trunc Now the backend supports the natural LLVM IR, we can shamelessly steal the AArch64 front-end code to implement the vshrn intrinsic on 32-bit ARM. llvm-svn: 201086	2014-02-10 14:04:12 +00:00
Tim Northover	7ffb2c5523	ARM & AArch64: combine implementation of vcaXYZ intrinsics Now that the back-end intrinsics are more regular, there's no need for the special handling these got in the front-end, so they can be moved to EmitCommonNeonBuiltinExpr. llvm-svn: 200769	2014-02-04 14:55:52 +00:00
Tim Northover	02e38609e7	ARM: implement support for crypto intrinsics in arm_neon.h llvm-svn: 200708	2014-02-03 17:28:04 +00:00
Tim Northover	51ab388266	AArch64: use new non-polymorphic crypto intrinsics The LLVM backend now has invariant types on the various crypto-intrinsics, because in all cases there's only really one interpretation. llvm-svn: 200707	2014-02-03 17:28:00 +00:00
Tim Northover	5309111c22	ARM & AArch64: unify the rest of the completely shared NEON implementations This should be the last routine patch: AArch64 does still delegate to EmitARMBuiltinExpr, but the remaining instances have complications of one sort or another so some more cunning thought will be needed. llvm-svn: 200528	2014-01-31 10:46:52 +00:00
Tim Northover	ba1e344d90	ARM & AArch64: another block of miscellaneous NEON sharing. llvm-svn: 200527	2014-01-31 10:46:49 +00:00
Tim Northover	027b4ee607	ARM & AArch64: move shared vld/vst intrinsics to common implementation. llvm-svn: 200526	2014-01-31 10:46:45 +00:00
Tim Northover	9d3ab5fe9f	ARM & AArch64: more instructions into common block llvm-svn: 200525	2014-01-31 10:46:41 +00:00
Tim Northover	61fc835d6e	ARM & AArch64: merge another NEON block completely. llvm-svn: 200524	2014-01-31 10:46:36 +00:00
Tim Northover	58c4474dea	ARM & AArch64: extend shared NEON implementation to first block. This extends the refactoring to the whole of the first block of trivial correspondences (as a fairly arbitrary boundary). llvm-svn: 200472	2014-01-30 14:48:01 +00:00
Tim Northover	ac85c341ae	ARM & AArch64: fully share NEON implementation of permutation intrinsics As a starting point, this moves the CodeGen for NEON permutation instructions (vtrn, vzip, vuzp) into a new shared function. llvm-svn: 200471	2014-01-30 14:47:57 +00:00
Tim Northover	c322f838bc	ARM & AArch64: share the BI__builtin_neon enum defs. llvm-svn: 200470	2014-01-30 14:47:51 +00:00
Kevin Qin	ce1f0e85ba	[AArch64 NEON] Fix a bug about vcles_f32 and vcled_f64. As vcles_f32() and vcled_f64 are implemented by FCMGE, operands should make a swap. llvm-svn: 199866	2014-01-23 03:42:06 +00:00
Hao Liu	f96fd37888	[AArch64]The compare to zero intrinsics should be implemented by 'icmp/fcmp' and 'sext' not 'zext'. Modify the implementation by replacing zext with sext. llvm-svn: 197898	2013-12-23 02:44:00 +00:00
Chad Rosier	6030c84a2f	[AArch64] Refactor NEON floating-point Max/Min/Maxnm/Minnm across vector AArch64 intrinsics to use f32 types, rather than their vector equivalents. llvm-svn: 197091	2013-12-11 23:21:39 +00:00
Chad Rosier	c520fce72d	[AArch64] Add NEON scalar floating-point compare LLVM AArch64 intrinsics that use f32/f64 types, rather than their vector equivalents. llvm-svn: 197071	2013-12-11 21:03:56 +00:00
Chad Rosier	edd4403510	[AArch64] Refactor the NEON scalar floating-point reciprocal step and floating-point reciprocal square root step LLVM AArch64 intrinsics to use f32/f64 types, rather than their vector equivalents. llvm-svn: 197070	2013-12-11 21:03:54 +00:00
Chad Rosier	6ce4387c5c	[AArch64] Refactor the NEON scalar floating-point reciprocal estimate, floating- point reciprocal exponent, and floating-point reciprocal square root estimate LLVM AArch64 intrinsics to use f32/f64 types, rather than their vector equivalents. llvm-svn: 197069	2013-12-11 21:03:52 +00:00
Chad Rosier	17c248a7a2	[AArch64] Refactor the NEON floating-point absolute difference LLVM AArch64 intrinsic to use f32/f64 types, rather than their vector equivalents. llvm-svn: 196969	2013-12-10 21:34:23 +00:00
Chad Rosier	37051a80e9	[AArch64] Refactor the NEON signed/unsigned floating-point convert to fixed-point LLVM AArch64 intrinsics to use f32/f64, rather than their vector equivalents. llvm-svn: 196968	2013-12-10 21:34:21 +00:00
Chad Rosier	8f6f3d124c	[AArch64] Overload NEON signed/unsigned floating-point convert to fixed-point and fixed-point convert to floating-point LLVM AArch64 intrinsics. llvm-svn: 196967	2013-12-10 21:34:20 +00:00
Chad Rosier	11a78c86e1	[AArch64] Overload NEON signed/unsigned integer convert to floating-point LLVM AArch64 intrinsics. llvm-svn: 196966	2013-12-10 21:34:17 +00:00
Chad Rosier	8d96c803df	[AArch64] Refactor the redundant code in the EmitAArch64ScalarBuiltinExpr() function. No functional change intended. llvm-svn: 196936	2013-12-10 17:44:36 +00:00
Chad Rosier	58f6a1fee7	[AArch64] Refactor the Neon vector/scalar floating-point convert intrinsics so that they use float/double rather than the vector equivalents when appropriate. llvm-svn: 196931	2013-12-10 16:11:55 +00:00
Chad Rosier	ff3b79aead	[AArch64] Refactor the Neon vector/scalar floating-point convert implementation. Specifically, reuse the ARM intrinsics when possible. llvm-svn: 196927	2013-12-10 15:35:40 +00:00
Kevin Qin	fb79d7f843	[AArch64 NEON] Support poly128_t and implement relevant intrinsic. llvm-svn: 196888	2013-12-10 06:49:01 +00:00
Chad Rosier	ce511f2fcb	[AArch64] Refactor the NEON scalar reduce pairwise intrinsics so that they use float/double rather than the vector equivalents when appropriate. llvm-svn: 196836	2013-12-09 22:47:59 +00:00
Chad Rosier	01703584eb	[AArch64] Refactor the NEON scalar reduce pairwise front-end codegen to remove unnecessary patterns in tablegen. llvm-svn: 196835	2013-12-09 22:47:57 +00:00
Chad Rosier	ad3683c3cb	[AArch64] Remove q and non-q intrinsic definitions from the NEON scalar reduce pairwise implementation, using an overloaded definition instead. llvm-svn: 196834	2013-12-09 22:47:55 +00:00
Hao Liu	844a7da243	[AArch64]Add missing pair intrinsics such as: int32_t vminv_s32(int32x2_t a) which should be compiled into SMINP Vd.2S,Vn.2S,Vm.2S llvm-svn: 196750	2013-12-09 03:52:22 +00:00
Kevin Qin	ad53b87c70	[AArch64 NEON] Add ACLE intrinsic vceqz_f64. llvm-svn: 196361	2013-12-04 08:02:11 +00:00
Kevin Qin	8903f8df4b	[AArch64 NEON] Add missing compare intrinsics. llvm-svn: 196359	2013-12-04 07:53:09 +00:00
Hao Liu	a5246fde90	[AArch64]Add missing floating point convert, round and misc intrinsics. E.g. int64x1_t vcvt_s64_f64(float64x1_t a) -> FCVTZS Dd, Dn llvm-svn: 196211	2013-12-03 06:07:13 +00:00
Hao Liu	4b850c5e0d	revert r196152. This is a duplicate implementation. E.g. this patch defines: float64_t vabd_f64(float64_t a, float64_t b) But there is already a similar intrinsic "vabdd_f64" with the same types. Also, this intrinsic will be conflicted to the vector type intrinsic as following(Which is implemented by me and will be committed to trunk): float64x1_t vabd_f64(float64x1_t a, float64x1_t b). Two functions shouldn't have a same name in arm_neon.h. According to ARM ACLE document, such vabd_f64 with float64_t is not existing. So I revert this commit. llvm-svn: 196205	2013-12-03 05:35:17 +00:00
Hao Liu	ce258820ca	AArch64: Add missing scalar pair intrinsics. E.g. "float32_t vaddv_f32(float32x2_t a)" to be matched into "faddp s0, v1.2s". llvm-svn: 196199	2013-12-03 03:40:08 +00:00
Chad Rosier	b0574f3bf7	[AArch64] Add missing NEON scalar floating-point to integer convert ACLEs. llvm-svn: 196152	2013-12-02 21:07:24 +00:00
Hao Liu	8a0099e02c	Fix the problem that the range check for scalar narrow shift is too wide. E.g. the immediate value of vshrns_n_s16 is [1,16], which should be [1,8]. llvm-svn: 195942	2013-11-29 02:13:17 +00:00
Chad Rosier	9e59285cc8	[AArch64] Add support for NEON scalar floating-point absolute difference. llvm-svn: 195804	2013-11-27 01:46:19 +00:00
Chad Rosier	52e31b20cb	[AArch64] Add support for NEON scalar floating-point to integer convert instructions. llvm-svn: 195789	2013-11-26 22:17:51 +00:00
Ana Pazos	dbd1a22496	Implemented Neon scalar vdup_lane intrinsics. Fixed scalar dup alias and added test case. llvm-svn: 195329	2013-11-21 08:15:01 +00:00
Ana Pazos	2b02688fd9	Implemented Neon scalar by element intrinsics. Intrinsics implemented: vqdmull_lane, vqdmulh_lane, vqrdmulh_lane, vqdmlal_lane, vqdmlsl_lane scalar Neon intrinsics. llvm-svn: 195326	2013-11-21 07:36:33 +00:00
Hao Liu	171cedf61e	Implement AArch64 neon instructions class SIMD lsone and SIMD lone-post. llvm-svn: 195079	2013-11-19 02:17:31 +00:00
Hao Liu	5e4ce1ae9d	Implement the newly added AArch64 ACLE functions for ld1/st1 with 2/3/4 vectors. The functions are like: vst1_s8_x2 ... llvm-svn: 194991	2013-11-18 06:33:43 +00:00
Benjamin Kramer	847c1d90e1	Remove unused but set variable. llvm-svn: 194920	2013-11-16 11:47:52 +00:00
Ana Pazos	6f2a47a9e5	Implemented aarch64 Neon scalar vmulx_lane intrinsics Implemented aarch64 Neon scalar vfma_lane intrinsics Implemented aarch64 Neon scalar vfms_lane intrinsics Implemented legacy vmul_n_f64, vmul_lane_f64, vmul_laneq_f64 intrinsics (v1f64 parameter type) using Neon scalar instructions. Implemented legacy vfma_lane_f64, vfms_lane_f64, vfma_laneq_f64, vfms_laneq_f64 intrinsics (v1f64 parameter type) using Neon scalar instructions. llvm-svn: 194889	2013-11-15 23:33:31 +00:00
Chad Rosier	7aaee48bf0	[AArch64] Add support for legacy AArch32 NEON scalar shift right by immediate and accumulate instructions. llvm-svn: 194732	2013-11-14 22:02:24 +00:00
Kevin Qin	caac85e612	[AArch64 neon] support poly64 and relevant intrinsic functions. llvm-svn: 194660	2013-11-14 03:29:16 +00:00
Kevin Qin	1718af6f0a	Implement aarch64 neon instruction class misc. llvm-svn: 194657	2013-11-14 02:45:18 +00:00
Jiangning Liu	18b707cb3f	Implement AArch64 NEON instruction set AdvSIMD (table). llvm-svn: 194649	2013-11-14 01:57:55 +00:00
Reid Kleckner	59e4a6f5e2	-fms-extensions: Recognize _alloca as an alias for the alloca builtin Differential Revision: http://llvm-reviews.chandlerc.com/D1989 llvm-svn: 194617	2013-11-13 22:58:53 +00:00
Chad Rosier	e714a962b5	[AArch64] Tests for legacy AArch32 NEON scalar shift by immediate instructions. A number of non-overloaded intrinsics have been replaced by thier overloaded counterparts. llvm-svn: 194599	2013-11-13 20:05:44 +00:00
Chad Rosier	249c714bb4	[AArch64] Add support for NEON scalar floating-point convert to fixed-point instructions. llvm-svn: 194395	2013-11-11 18:04:22 +00:00
Jiangning Liu	c628af66c7	Implement AArch64 Neon instruction set Perm. llvm-svn: 194124	2013-11-06 03:35:53 +00:00
Jiangning Liu	37f5bb1b28	Implement AArch64 Neon instruction set Bitwise Extract. llvm-svn: 194119	2013-11-06 02:26:12 +00:00
Jiangning Liu	34a7109b47	Implement AArch64 Neon Crypto instruction classes AES, SHA, and 3 SHA. llvm-svn: 194086	2013-11-05 17:42:24 +00:00
Kevin Qin	9eece7b5e0	Implemented aarch64 neon intrinsic vcopy_lane with float type. llvm-svn: 194042	2013-11-05 02:05:44 +00:00
Chad Rosier	74329d6cff	[AArch64] Add support for NEON scalar fixed-point convert to floating-point instructions. llvm-svn: 193817	2013-10-31 22:37:08 +00:00
Chad Rosier	bdca387884	[AArch64] Add support for NEON scalar shift immediate instructions. llvm-svn: 193791	2013-10-31 19:29:05 +00:00
Mark Lacey	a8e7df3602	Add CodeGenABITypes.h for use in LLDB. CodeGenABITypes is a wrapper built on top of CodeGenModule that exposes some of the functionality of CodeGenTypes (held by CodeGenModule), specifically methods that determine the LLVM types appropriate for function argument and return values. I addition to CodeGenABITypes.h, CGFunctionInfo.h is introduced, and the definitions of ABIArgInfo, RequiredArgs, and CGFunctionInfo are moved into this new header from the private headers ABIInfo.h and CGCall.h. Exposing this functionality is one part of making it possible for LLDB to determine the actual ABI locations of function arguments and return values, making it possible for it to determine this for any supported target without hard-coding ABI knowledge in the LLDB code. llvm-svn: 193717	2013-10-30 21:53:58 +00:00
Chad Rosier	4d55e6e0a4	[AArch64] Add support for NEON scalar floating-point compare instructions. llvm-svn: 193692	2013-10-30 15:20:07 +00:00
Peter Collingbourne	b453cd64a7	Implement function type checker for the undefined behavior sanitizer. This uses function prefix data to store function type information at the function pointer. Differential Revision: http://llvm-reviews.chandlerc.com/D1338 llvm-svn: 193058	2013-10-20 21:29:19 +00:00
Chad Rosier	3c03dee1d1	[AArch64] Add support for NEON scalar extract narrow instructions. llvm-svn: 192971	2013-10-18 14:03:36 +00:00
Chad Rosier	e7465644c6	[AArch64] Add support for NEON scalar three register different instruction class. The instruction class includes the signed saturating doubling multiply-add long, signed saturating doubling multiply-subtract long, and the signed saturating doubling multiply long instructions. llvm-svn: 192909	2013-10-17 18:12:50 +00:00
Chad Rosier	00eef17dbe	[AArch64] Add support for NEON scalar negate instruction. llvm-svn: 192845	2013-10-16 21:04:53 +00:00
Chad Rosier	e904137c01	[AArch64] Add support for NEON scalar absolute value instruction. llvm-svn: 192844	2013-10-16 21:04:49 +00:00
Chad Rosier	2681b3fb61	Update comment. llvm-svn: 192807	2013-10-16 16:30:39 +00:00
Chad Rosier	069b90463d	[AArch64] Add support for NEON scalar signed saturating accumulated of unsigned value and unsigned saturating accumulate of signed value instructions. llvm-svn: 192801	2013-10-16 16:09:16 +00:00
Chad Rosier	a70fb7b716	[AArch64] Add support for NEON scalar signed saturating absolute value and scalar signed saturating negate instructions. llvm-svn: 192734	2013-10-15 21:19:02 +00:00
Chad Rosier	193573ec89	[AArch64] Add support for NEON scalar integer compare instructions. llvm-svn: 192597	2013-10-14 14:37:40 +00:00
Kevin Qin	f22bf50443	Implemented aarch64 SIMD copy related ACLE intrinsic : vget_lane, vset_lane, vcopy_lane, vcreate, vdup_n, vdup_lane, vmov_n. llvm-svn: 192411	2013-10-11 02:34:30 +00:00
Hao Liu	1eade6d927	Implement AArch64 vector load/store multiple N-element structure class SIMD(lselem). Including following 14 instructions: 4 ld1 insts: load multiple 1-element structure to sequential 1/2/3/4 registers. ld2/ld3/ld4: load multiple N-element structure to sequential N registers (N=2,3,4). 4 st1 insts: store multiple 1-element structure from sequential 1/2/3/4 registers. st2/st3/st4: store multiple N-element structure from sequential N registers (N = 2,3,4). llvm-svn: 192362	2013-10-10 17:01:49 +00:00
Tim Northover	72ace5cf12	Revert "Implement AArch64 vector load/store multiple N-element structure class SIMD(lselem). " This reverts commit r192351. The LLVM side broke the build and the Clang tests will inevitably fail without it. llvm-svn: 192356	2013-10-10 16:00:08 +00:00
Hao Liu	c319193636	Implement AArch64 vector load/store multiple N-element structure class SIMD(lselem). Including following 14 instructions: 4 ld1 insts: load multiple 1-element structure to sequential 1/2/3/4 registers. ld2/ld3/ld4: load multiple N-element structure to sequential N registers (N=2,3,4). 4 st1 insts: store multiple 1-element structure from sequential 1/2/3/4 registers. st2/st3/st4: store multiple N-element structure from sequential N registers (N = 2,3,4). E.g. ld1(3 registers version) will load 32-bit elements {A, B, C, D, E, F} sequentially into the three 64-bit vectors list {BA, DC, FE}. E.g. ld3 will load 32-bit elements {A, B, C, D, E, F} into the three 64-bit vectors list {DA, EB, FC}. llvm-svn: 192351	2013-10-10 14:59:36 +00:00
Chad Rosier	0a903478c6	[AArch64] Add support for NEON scalar floating-point reciprocal estimate, reciprocal exponent, and reciprocal square root estimate instructions. llvm-svn: 192243	2013-10-08 22:09:29 +00:00
Chad Rosier	0babda4b9c	[AArch64] Add support for NEON scalar signed/unsigned integer to floating-point convert instructions. llvm-svn: 192232	2013-10-08 20:43:46 +00:00
Matt Arsenault	2f15263807	Fix objectsize tests after r192117 llvm-svn: 192120	2013-10-07 19:00:18 +00:00
Chad Rosier	027dfade54	[AArch64] Add support for NEON scalar arithmetic instructions: SQDMULH, SQRDMULH, FMULX, FRECPS, and FRSQRTS. llvm-svn: 192112	2013-10-07 17:07:17 +00:00
Jiangning Liu	b96ebac02b	Implement aarch64 neon instruction set AdvSIMD (Across). llvm-svn: 192029	2013-10-05 08:22:55 +00:00
Amaury de la Vieuville	21bf6ed730	Do not emit undefined lsrh/ashr for NEON shifts These IR instructions are undefined when the amount is equal to operand size, but NEON right shifts support such shifts. Work around that by emitting a different IR in these cases. llvm-svn: 191953	2013-10-04 13:13:15 +00:00
Jiangning Liu	4617e9dc85	Implement aarch64 neon instruction set AdvSIMD (3V elem). llvm-svn: 191945	2013-10-04 09:21:17 +00:00
Joey Gouly	75987a65f3	[ARM] Add a builtin to allow you to use the 'sevl' instruction. llvm-svn: 191816	2013-10-02 10:00:18 +00:00
Benjamin Kramer	9b1dfe8b56	Mark an impossible path as unreachable to pacify GCC. llvm-svn: 191436	2013-09-26 16:36:08 +00:00
Benjamin Kramer	39c4924db9	Remove tabs. llvm-svn: 191427	2013-09-26 12:16:47 +00:00
NAKAMURA Takumi	788af10a8a	CGBuiltin.cpp: Prune a stray default: label. [-Wcovered-switch-default] llvm-svn: 191277	2013-09-24 04:37:50 +00:00
Jiangning Liu	036f16dc8c	Initial support for Neon scalar instructions. Patch by Ana Pazos. 1.Added support for v1ix and v1fx types. 2.Added Scalar Pairwise Reduce instructions. 3.Added initial implementation of Scalar Arithmetic instructions. llvm-svn: 191264	2013-09-24 02:48:06 +00:00
Eli Friedman	f9d8c6cebb	Add _mm_stream_si64 intrinsic. While I'm here, also fix the alignment computation for the whole family of intrinsics. PR17298. llvm-svn: 191243	2013-09-23 23:38:39 +00:00
Joey Gouly	1e8637b259	[ARMv8] Add builtins for CRC instructions. Patch by Bradley Smith! llvm-svn: 190931	2013-09-18 10:07:09 +00:00
Hal Finkel	28b2ae3692	Restore the sqrt -> llvm.sqrt mapping in fast-math mode This restores the sqrt -> llvm.sqrt mapping, but only in fast-math mode (specifically, when the UnsafeFPMath or NoNaNsFPMath CodeGen options are enabled). The @llvm.sqrt* intrinsics have slightly different semantics from the libm call, specifically, they are undefined when given a non-zero negative number (the libm calls will always return NaN for any negative number). This mapping was removed in r100613, and replaced with a TODO, but at that time the fast-math flags were not yet implemented. Now that we have these, restoring this mapping is important because it will enable autovectorization of sqrt calls in loops (at least in fast-math mode). llvm-svn: 190646	2013-09-12 23:57:55 +00:00
Jiangning Liu	1bda93a252	Implement aarch64 neon instruction set AdvSIMD (3V Diff), covering the following 26 instructions, SADDL, UADDL, SADDW, UADDW, SSUBL, USUBL, SSUBW, USUBW, ADDHN, RADDHN, SABAL, UABAL, SUBHN, RSUBHN, SABDL, UABDL, SMLAL, UMLAL, SMLSL, UMLSL, SQDMLAL, SQDMLSL, SMULL, UMULL, SQDMULL, PMULL llvm-svn: 190289	2013-09-09 02:21:08 +00:00
Hao Liu	b1852eed38	Inplement aarch64 neon instructions in AdvSIMD(shift). About 24 shift instructions: sshr,ushr,ssra,usra,srshr,urshr,srsra,ursra,sri,shl,sli,sqshlu,sqshl,uqshl,shrn,sqrshr$ and 4 convert instructions: scvtf,ucvtf,fcvtzs,fcvtzu llvm-svn: 189926	2013-09-04 09:29:13 +00:00
Tim Northover	550ce58312	ARM: comment on why vmull intrinsic has to exist for now. llvm-svn: 189464	2013-08-28 09:46:40 +00:00
Tim Northover	4ae9812283	ARM: Emit normal IR for vaddhn/vsubhn NEON intrinsics These operations "vector add high-half narrow" actually correspond to the sequence: %sum = add <4 x i32> %lhs, %rhs %high = lshr <4 x i32> %sum, <i32 16, i32 16, i32 16, i32 16> %res = trunc <4 x i32> %high to <4 x i16> Now that LLVM can spot this, Clang should emit the corresponding LLVM IR. llvm-svn: 189463	2013-08-28 09:46:37 +00:00
Tim Northover	4e423f724a	ARM: use vqdmull and vqadds/vqsubs to implement vqdmlal/vqdmlsl The NEON intrinsics vqdmlal and vqdmlsl are really just combinations of a saturating-doubling-multiply (vqdmull) and a saturating add/sub, so now that LLVM can spot those patterns Clang should emit them instead of specialised intrinsics. Feature already tested by existing ARM NEON intrinsics tests. llvm-svn: 189462	2013-08-28 09:46:34 +00:00
Juergen Ributzka	53e2f275d2	Fix last commit. llvm-svn: 188724	2013-08-19 23:08:53 +00:00
Juergen Ributzka	c6ab1f8bfd	Simplify code by using CreateMemTemp. No functional change intended. Reviewer: Eli llvm-svn: 188722	2013-08-19 22:20:37 +00:00
Juergen Ributzka	2c2dbf4542	Fix the name and the type of the argument for intrinisc _mm256_broadcastsi128_si256 to align with the Intel documentation. This fixes bug PR 16581 and rdar:14747994. llvm-svn: 188609	2013-08-17 16:40:09 +00:00

1 2 3 4 5 ...

516 Commits