This is an archive of the discontinued LLVM Phabricator instance.

[X86] Prefer reduced width multiplication over pmulld on Silvermont
ClosedPublic

Authored by zvi on Nov 29 2016, 4:58 AM.

Download Raw Diff

Details

Reviewers

delena
wmi
mkuper

Commits

rG8bc7e4da516c: [X86] Prefer reduced width multiplication over pmulld on Silvermont
rL288844: [X86] Prefer reduced width multiplication over pmulld on Silvermont

Summary

Prefer expansions such as: pmullw,pmulhw,unpacklwd,unpackhwd over pmulld.
On Silvermont [source: Optimization Reference Manual]:
PMULLD has a throughput of 1/11 [instruction/cycles].
PMULHUW/PMULHW/PMULLW have a throughput of 1/2 [instruction/cycles].

Fixes pr31202.

Analysis of this issue was done by Fahana Aleen.

Diff Detail

Repository: rL LLVM

Event Timeline

zvi updated this revision to Diff 79519.Nov 29 2016, 4:58 AM

zvi retitled this revision from to [X86] Prefer reduced width multiplication over pmulld on Silvermont.

zvi updated this object.

zvi added reviewers: mkuper, delena, wmi.

zvi set the repository for this revision to rL LLVM.

zvi added a subscriber: llvm-commits.

RKSimon added a subscriber: RKSimon.Nov 30 2016, 2:34 AM

mkuper added inline comments.Nov 30 2016, 10:24 AM

lib/Target/X86/X86Subtarget.cpp
232 ↗	(On Diff #79519)	(!isPMULLDSlow() \|\| hasSSE41()) is good enough, logically speaking.
test/CodeGen/X86/slow-pmulld.ll
1 ↗	(On Diff #79519)	Please add a check that we do generate a pmulld for non-slow targets. (Or, if we already have such a test, merge this one into that).

zvi added inline comments.Dec 4 2016, 12:42 AM

lib/Target/X86/X86Subtarget.cpp
232 ↗	(On Diff #79519)	Good catch
test/CodeGen/X86/slow-pmulld.ll
1 ↗	(On Diff #79519)	Ok, I will add tests for SSE4.1 targets w/o the slowpmulld feature

Fixes for Michael's comments.

delena accepted this revision.Dec 6 2016, 1:08 AM

delena edited edge metadata.

This revision is now accepted and ready to land.Dec 6 2016, 1:08 AM

@mkuper Anything to add?

Sorry, missed the update notification. LGTM.

Closed by commit rL288844: [X86] Prefer reduced width multiplication over pmulld on Silvermont (authored by zvi). · Explain WhyDec 6 2016, 11:45 AM

This revision was automatically updated to reflect the committed changes.

Revision Contents

Path

Size

llvm/

trunk/

lib/

Target/

X86/

3 lines

13 lines

5 lines

4 lines

test/

CodeGen/

X86/

slow-pmulld.ll

71 lines

Diff 80458

llvm/trunk/lib/Target/X86/X86.td

Show First 20 Lines • Show All 93 Lines • ▼ Show 20 Lines	def Feature64Bit : SubtargetFeature<"64bit", "HasX86_64", "true",
[FeatureCMOV]>;		[FeatureCMOV]>;
def FeatureCMPXCHG16B : SubtargetFeature<"cx16", "HasCmpxchg16b", "true",		def FeatureCMPXCHG16B : SubtargetFeature<"cx16", "HasCmpxchg16b", "true",
"64-bit with cmpxchg16b",		"64-bit with cmpxchg16b",
[Feature64Bit]>;		[Feature64Bit]>;
def FeatureSlowBTMem : SubtargetFeature<"slow-bt-mem", "IsBTMemSlow", "true",		def FeatureSlowBTMem : SubtargetFeature<"slow-bt-mem", "IsBTMemSlow", "true",
"Bit testing of memory is slow">;		"Bit testing of memory is slow">;
def FeatureSlowSHLD : SubtargetFeature<"slow-shld", "IsSHLDSlow", "true",		def FeatureSlowSHLD : SubtargetFeature<"slow-shld", "IsSHLDSlow", "true",
"SHLD instruction is slow">;		"SHLD instruction is slow">;
		def FeatureSlowPMULLD : SubtargetFeature<"slow-pmulld", "IsPMULLDSlow", "true",
		"PMULLD instruction is slow">;
// FIXME: This should not apply to CPUs that do not have SSE.		// FIXME: This should not apply to CPUs that do not have SSE.
def FeatureSlowUAMem16 : SubtargetFeature<"slow-unaligned-mem-16",		def FeatureSlowUAMem16 : SubtargetFeature<"slow-unaligned-mem-16",
"IsUAMem16Slow", "true",		"IsUAMem16Slow", "true",
"Slow unaligned 16-byte memory access">;		"Slow unaligned 16-byte memory access">;
def FeatureSlowUAMem32 : SubtargetFeature<"slow-unaligned-mem-32",		def FeatureSlowUAMem32 : SubtargetFeature<"slow-unaligned-mem-32",
"IsUAMem32Slow", "true",		"IsUAMem32Slow", "true",
"Slow unaligned 32-byte memory access">;		"Slow unaligned 32-byte memory access">;
def FeatureSSE4A : SubtargetFeature<"sse4a", "HasSSE4A", "true",		def FeatureSSE4A : SubtargetFeature<"sse4a", "HasSSE4A", "true",
▲ Show 20 Lines • Show All 288 Lines • ▼ Show 20 Lines	class SilvermontProc<string Name> : ProcessorModel<Name, SLMModel, [
FeaturePCLMUL,		FeaturePCLMUL,
FeatureAES,		FeatureAES,
FeatureSlowDivide64,		FeatureSlowDivide64,
FeatureCallRegIndirect,		FeatureCallRegIndirect,
FeaturePRFCHW,		FeaturePRFCHW,
FeatureSlowLEA,		FeatureSlowLEA,
FeatureSlowIncDec,		FeatureSlowIncDec,
FeatureSlowBTMem,		FeatureSlowBTMem,
		FeatureSlowPMULLD,
FeatureLAHFSAHF		FeatureLAHFSAHF
]>;		]>;
def : SilvermontProc<"silvermont">;		def : SilvermontProc<"silvermont">;
def : SilvermontProc<"slm">; // Legacy alias.		def : SilvermontProc<"slm">; // Legacy alias.

// "Arrandale" along with corei3 and corei5		// "Arrandale" along with corei3 and corei5
class NehalemProc<string Name> : ProcessorModel<Name, SandyBridgeModel, [		class NehalemProc<string Name> : ProcessorModel<Name, SandyBridgeModel, [
FeatureX87,		FeatureX87,
▲ Show 20 Lines • Show All 440 Lines • Show Last 20 Lines

llvm/trunk/lib/Target/X86/X86ISelLowering.cpp

This file is larger than 256 KB, so syntax highlighting is disabled by default.

	Show First 20 Lines • Show All 29,296 Lines • ▼ Show 20 Lines
	/// If %2 == sext32(trunc16(%2)), i.e., the scalar value range of %2 is			/// If %2 == sext32(trunc16(%2)), i.e., the scalar value range of %2 is
	/// -32768 to 32767, and the scalar value range of %4 is also -32768 to 32767,			/// -32768 to 32767, and the scalar value range of %4 is also -32768 to 32767,
	/// generate pmullw+pmulhw for it (MULS16 mode).			/// generate pmullw+pmulhw for it (MULS16 mode).
	/// If %2 == zext32(trunc16(%2)), i.e., the scalar value range of %2 is			/// If %2 == zext32(trunc16(%2)), i.e., the scalar value range of %2 is
	/// 0 to 65535, and the scalar value range of %4 is also 0 to 65535,			/// 0 to 65535, and the scalar value range of %4 is also 0 to 65535,
	/// generate pmullw+pmulhuw for it (MULU16 mode).			/// generate pmullw+pmulhuw for it (MULU16 mode).
	static SDValue reduceVMULWidth(SDNode *N, SelectionDAG &DAG,			static SDValue reduceVMULWidth(SDNode *N, SelectionDAG &DAG,
	const X86Subtarget &Subtarget) {			const X86Subtarget &Subtarget) {
	// pmulld is supported since SSE41. It is better to use pmulld			// Check for legality
	// instead of pmullw+pmulhw.
	// pmullw/pmulhw are not supported by SSE.			// pmullw/pmulhw are not supported by SSE.
	if (Subtarget.hasSSE41() \|\| !Subtarget.hasSSE2())			if (!Subtarget.hasSSE2())
				return SDValue();

				// Check for profitability
				// pmulld is supported since SSE41. It is better to use pmulld
				// instead of pmullw+pmulhw, except for subtargets where pmulld is slower than
				// the expansion.
				bool OptForMinSize = DAG.getMachineFunction().getFunction()->optForMinSize();
				if (Subtarget.hasSSE41() && (OptForMinSize \|\| !Subtarget.isPMULLDSlow()))
	return SDValue();			return SDValue();

	ShrinkMode Mode;			ShrinkMode Mode;
	if (!canReduceVMulWidth(N, DAG, Mode))			if (!canReduceVMulWidth(N, DAG, Mode))
	return SDValue();			return SDValue();

	SDLoc DL(N);			SDLoc DL(N);
	SDValue N0 = N->getOperand(0);			SDValue N0 = N->getOperand(0);
	▲ Show 20 Lines • Show All 4,777 Lines • Show Last 20 Lines

llvm/trunk/lib/Target/X86/X86Subtarget.h

Show First 20 Lines • Show All 172 Lines • ▼ Show 20 Lines	protected:
bool HasPFPREFETCHWT1;		bool HasPFPREFETCHWT1;

/// True if BT (bit test) of memory instructions are slow.		/// True if BT (bit test) of memory instructions are slow.
bool IsBTMemSlow;		bool IsBTMemSlow;

/// True if SHLD instructions are slow.		/// True if SHLD instructions are slow.
bool IsSHLDSlow;		bool IsSHLDSlow;

		/// True if the PMULLD instruction is slow compared to PMULLW/PMULHW and
		// PMULUDQ.
		bool IsPMULLDSlow;

/// True if unaligned memory accesses of 16-bytes are slow.		/// True if unaligned memory accesses of 16-bytes are slow.
bool IsUAMem16Slow;		bool IsUAMem16Slow;

/// True if unaligned memory accesses of 32-bytes are slow.		/// True if unaligned memory accesses of 32-bytes are slow.
bool IsUAMem32Slow;		bool IsUAMem32Slow;

/// True if SSE operations can have unaligned memory operands.		/// True if SSE operations can have unaligned memory operands.
/// This may require setting a configuration bit in the processor.		/// This may require setting a configuration bit in the processor.
▲ Show 20 Lines • Show All 258 Lines • ▼ Show 20 Lines	public:
bool hasADX() const { return HasADX; }		bool hasADX() const { return HasADX; }
bool hasSHA() const { return HasSHA; }		bool hasSHA() const { return HasSHA; }
bool hasPRFCHW() const { return HasPRFCHW; }		bool hasPRFCHW() const { return HasPRFCHW; }
bool hasRDSEED() const { return HasRDSEED; }		bool hasRDSEED() const { return HasRDSEED; }
bool hasLAHFSAHF() const { return HasLAHFSAHF; }		bool hasLAHFSAHF() const { return HasLAHFSAHF; }
bool hasMWAITX() const { return HasMWAITX; }		bool hasMWAITX() const { return HasMWAITX; }
bool isBTMemSlow() const { return IsBTMemSlow; }		bool isBTMemSlow() const { return IsBTMemSlow; }
bool isSHLDSlow() const { return IsSHLDSlow; }		bool isSHLDSlow() const { return IsSHLDSlow; }
		bool isPMULLDSlow() const { return IsPMULLDSlow; }
bool isUnalignedMem16Slow() const { return IsUAMem16Slow; }		bool isUnalignedMem16Slow() const { return IsUAMem16Slow; }
bool isUnalignedMem32Slow() const { return IsUAMem32Slow; }		bool isUnalignedMem32Slow() const { return IsUAMem32Slow; }
bool hasSSEUnalignedMem() const { return HasSSEUnalignedMem; }		bool hasSSEUnalignedMem() const { return HasSSEUnalignedMem; }
bool hasCmpxchg16b() const { return HasCmpxchg16b; }		bool hasCmpxchg16b() const { return HasCmpxchg16b; }
bool useLeaForSP() const { return UseLeaForSP; }		bool useLeaForSP() const { return UseLeaForSP; }
bool hasFastPartialYMMWrite() const { return HasFastPartialYMMWrite; }		bool hasFastPartialYMMWrite() const { return HasFastPartialYMMWrite; }
bool hasFastScalarFSQRT() const { return HasFastScalarFSQRT; }		bool hasFastScalarFSQRT() const { return HasFastScalarFSQRT; }
bool hasFastVectorFSQRT() const { return HasFastVectorFSQRT; }		bool hasFastVectorFSQRT() const { return HasFastVectorFSQRT; }
▲ Show 20 Lines • Show All 166 Lines • Show Last 20 Lines

llvm/trunk/lib/Target/X86/X86Subtarget.cpp

Show First 20 Lines • Show All 222 Lines • ▼ Show 20 Lines	void X86Subtarget::initSubtargetFeatures(StringRef CPU, StringRef FS) {

// Stack alignment is 16 bytes on Darwin, Linux, kFreeBSD and Solaris (both		// Stack alignment is 16 bytes on Darwin, Linux, kFreeBSD and Solaris (both
// 32 and 64 bit) and for all 64-bit targets.		// 32 and 64 bit) and for all 64-bit targets.
if (StackAlignOverride)		if (StackAlignOverride)
stackAlignment = StackAlignOverride;		stackAlignment = StackAlignOverride;
else if (isTargetDarwin() \|\| isTargetLinux() \|\| isTargetSolaris() \|\|		else if (isTargetDarwin() \|\| isTargetLinux() \|\| isTargetSolaris() \|\|
isTargetKFreeBSD() \|\| In64BitMode)		isTargetKFreeBSD() \|\| In64BitMode)
stackAlignment = 16;		stackAlignment = 16;

		assert((!isPMULLDSlow() \|\| hasSSE41()) &&
		"Feature Slow PMULLD can only be set on a subtarget with SSE4.1");
}		}

void X86Subtarget::initializeEnvironment() {		void X86Subtarget::initializeEnvironment() {
X86SSELevel = NoSSE;		X86SSELevel = NoSSE;
X863DNowLevel = NoThreeDNow;		X863DNowLevel = NoThreeDNow;
HasX87 = false;		HasX87 = false;
HasCMov = false;		HasCMov = false;
HasX86_64 = false;		HasX86_64 = false;
Show All 31 Lines	void X86Subtarget::initializeEnvironment() {
HasPKU = false;		HasPKU = false;
HasSHA = false;		HasSHA = false;
HasPRFCHW = false;		HasPRFCHW = false;
HasRDSEED = false;		HasRDSEED = false;
HasLAHFSAHF = false;		HasLAHFSAHF = false;
HasMWAITX = false;		HasMWAITX = false;
HasMPX = false;		HasMPX = false;
IsBTMemSlow = false;		IsBTMemSlow = false;
		IsPMULLDSlow = false;
IsSHLDSlow = false;		IsSHLDSlow = false;
IsUAMem16Slow = false;		IsUAMem16Slow = false;
IsUAMem32Slow = false;		IsUAMem32Slow = false;
HasSSEUnalignedMem = false;		HasSSEUnalignedMem = false;
HasCmpxchg16b = false;		HasCmpxchg16b = false;
UseLeaForSP = false;		UseLeaForSP = false;
HasFastPartialYMMWrite = false;		HasFastPartialYMMWrite = false;
HasFastScalarFSQRT = false;		HasFastScalarFSQRT = false;
▲ Show 20 Lines • Show All 72 Lines • Show Last 20 Lines

llvm/trunk/test/CodeGen/X86/slow-pmulld.ll

Property	Old Value	New Value
svn:eol-style	null	native

				; NOTE: Assertions have been autogenerated by utils/update_llc_test_checks.py
				; RUN: llc < %s -mtriple=i386-unknown-unknown -mcpu=silvermont \| FileCheck %s --check-prefix=CHECK32
				; RUN: llc < %s -mtriple=x86_64-unknown-unknown -mcpu=silvermont \| FileCheck %s --check-prefix=CHECK64
				; RUN: llc < %s -mtriple=i386-unknown-unknown -mattr=+sse4.1 \| FileCheck %s --check-prefix=SSE4-32
				; RUN: llc < %s -mtriple=x86_64-unknown-unknown -mattr=+sse4.1 \| FileCheck %s --check-prefix=SSE4-64

				define <4 x i32> @foo(<4 x i8> %A) {
				; CHECK32-LABEL: foo:
				; CHECK32: # BB#0:
				; CHECK32-NEXT: pshufb {{.*#+}} xmm0 = xmm0[0],zero,xmm0[4],zero,xmm0[8],zero,xmm0[12],zero,xmm0[u,u,u,u,u,u,u,u]
				; CHECK32-NEXT: movdqa {{.*#+}} xmm1 = <18778,18778,18778,18778,u,u,u,u>
				; CHECK32-NEXT: movdqa %xmm0, %xmm2
				; CHECK32-NEXT: pmullw %xmm1, %xmm0
				; CHECK32-NEXT: pmulhw %xmm1, %xmm2
				; CHECK32-NEXT: punpcklwd {{.*#+}} xmm0 = xmm0[0],xmm2[0],xmm0[1],xmm2[1],xmm0[2],xmm2[2],xmm0[3],xmm2[3]
				; CHECK32-NEXT: retl
				;
				; CHECK64-LABEL: foo:
				; CHECK64: # BB#0:
				; CHECK64-NEXT: pshufb {{.*#+}} xmm0 = xmm0[0],zero,xmm0[4],zero,xmm0[8],zero,xmm0[12],zero,xmm0[u,u,u,u,u,u,u,u]
				; CHECK64-NEXT: movdqa {{.*#+}} xmm1 = <18778,18778,18778,18778,u,u,u,u>
				; CHECK64-NEXT: movdqa %xmm0, %xmm2
				; CHECK64-NEXT: pmullw %xmm1, %xmm0
				; CHECK64-NEXT: pmulhw %xmm1, %xmm2
				; CHECK64-NEXT: punpcklwd {{.*#+}} xmm0 = xmm0[0],xmm2[0],xmm0[1],xmm2[1],xmm0[2],xmm2[2],xmm0[3],xmm2[3]
				; CHECK64-NEXT: retq
				;
				; SSE4-32-LABEL: foo:
				; SSE4-32: # BB#0:
				; SSE4-32-NEXT: pand {{\.LCPI.*}}, %xmm0
				; SSE4-32-NEXT: pmulld {{\.LCPI.*}}, %xmm0
				; SSE4-32-NEXT: retl
				;
				; SSE4-64-LABEL: foo:
				; SSE4-64: # BB#0:
				; SSE4-64-NEXT: pand {{.*}}(%rip), %xmm0
				; SSE4-64-NEXT: pmulld {{.*}}(%rip), %xmm0
				; SSE4-64-NEXT: retq
				%z = zext <4 x i8> %A to <4 x i32>
				%m = mul nuw nsw <4 x i32> %z, <i32 18778, i32 18778, i32 18778, i32 18778>
				ret <4 x i32> %m
				}

				define <4 x i32> @foo_os(<4 x i8> %A) minsize {
				; CHECK32-LABEL: foo_os:
				; CHECK32: # BB#0:
				; CHECK32-NEXT: pand {{\.LCPI.*}}, %xmm0
				; CHECK32-NEXT: pmulld {{\.LCPI.*}}, %xmm0
				; CHECK32-NEXT: retl
				;
				; CHECK64-LABEL: foo_os:
				; CHECK64: # BB#0:
				; CHECK64-NEXT: pand {{.*}}(%rip), %xmm0
				; CHECK64-NEXT: pmulld {{.*}}(%rip), %xmm0
				; CHECK64-NEXT: retq
				;
				; SSE4-32-LABEL: foo_os:
				; SSE4-32: # BB#0:
				; SSE4-32-NEXT: pand {{\.LCPI.*}}, %xmm0
				; SSE4-32-NEXT: pmulld {{\.LCPI.*}}, %xmm0
				; SSE4-32-NEXT: retl
				;
				; SSE4-64-LABEL: foo_os:
				; SSE4-64: # BB#0:
				; SSE4-64-NEXT: pand {{.*}}(%rip), %xmm0
				; SSE4-64-NEXT: pmulld {{.*}}(%rip), %xmm0
				; SSE4-64-NEXT: retq
				%z = zext <4 x i8> %A to <4 x i32>
				%m = mul nuw nsw <4 x i32> %z, <i32 18778, i32 18778, i32 18778, i32 18778>
				ret <4 x i32> %m
				}