Page MenuHomePhabricator

[X86] Custom lower ISD::FROUND with SSE4.1 to avoid a libcall.
ClosedPublic

Authored by craig.topper on Jan 29 2020, 12:01 AM.

Details

Summary

ISD::FROUND is defined to round to nearest with ties rounding
away from 0. This mode isn't supported in hardware on X86.

But as long as we aren't compiling with trapping math, we can
emulate this with floor(X + copysign(nextafter(0.5, 0.0), X)).

We have to use nextafter to avoid some corner cases that adding
0.5 would have. For example, if X is nextafter(0.5, 0.0) it should
round to 0.0, but adding 0.5 would need one extra bit of mantissa
than can be stored so it rounds to 1.0. Adding nextafter(0.5, 0.0)
instead will just increase the exponent by 1 and leave the mantissa
as all 1s. This would be nextafter(1.0, 0.0) which will floor to 0.0.

Techically this requires -fno-trapping-math which isn't our default.
But if we care about exceptions we should be using constrained
intrinsics. Constrained intrinsics would use STRICT_FROUND which
won't go through this code.

Fixes PR42195.

Diff Detail

Event Timeline

craig.topper created this revision.Jan 29 2020, 12:01 AM
Herald added a project: Restricted Project. · View Herald TranscriptJan 29 2020, 12:01 AM
Herald added a subscriber: hiraditya. · View Herald Transcript
spatel accepted this revision.Jan 29 2020, 6:53 AM

LGTM

llvm/lib/Target/X86/X86ISelLowering.cpp
20460

Include at least part of the text from this patch description as a block comment for this function, so:

/// ISD::FROUND is defined to round to nearest with ties rounding
/// away from 0. This mode isn't supported in hardware on X86.
/// But as long as we aren't compiling with trapping math, we can
/// emulate this with floor(X + copysign(nextafter(0.5, 0.0), X)).
/// ...
llvm/test/CodeGen/X86/vec_round.ll
0

This test/file was added with:
rL188048
...to show we wouldn't crash?
But it doesn't add value now that we have more thorough tests. I'd delete it.

This revision is now accepted and ready to land.Jan 29 2020, 6:53 AM
This revision was automatically updated to reflect the committed changes.