Download Raw Diff

Details

Reviewers

jholewinski
eliben
echristo

Commits

rGc29db8441959: [CUDA] Added a wrapper header for inclusion of stock CUDA headers.
rC253388: [CUDA] Added a wrapper header for inclusion of stock CUDA headers.
rL253388: [CUDA] Added a wrapper header for inclusion of stock CUDA headers.

Summary

Header files that come with CUDA are assuming split host/device compilation and are not usable by clang out of the box.
With a bit of preprocessor magic it's possible to twist them into a form usable by clang after D13170 and D13144 land.

Diff Detail

Repository: rL LLVM

Event Timeline

tra updated this revision to Diff 35739.Sep 25 2015, 10:58 AM

tra retitled this revision from to [CUDA] Added a wrapper header for inclusion of stock CUDA headers..

tra updated this object.

tra added reviewers: echristo, eliben, jholewinski.

tra added a subscriber: cfe-commits.

Fixed typos and whitespace nits.
use #pragma push_macro for CUDACC_RTC, too.

eliben added inline comments.Sep 25 2015, 12:14 PM

lib/Headers/clang_cuda_support.h
53 ↗	(On Diff #35747)	If this includes CUDA headers, maybe you can #error out if the CUDA version isn't good?

Include cuda.h and #error if we see wrong CUDA_VERSION.

lgtm

This revision is now accepted and ready to land.Sep 25 2015, 1:23 PM

Bikeshed: it's part of the clang headers, do we really need "clang" in the header name?

-eric

The (vague) idea was to make clear that the header is *not* part of cuda
distribution.

That said, the file could use a better name.

Do any of these sound better?

fix_cuda_headers.h
adapt_cuda_headers.h
cuda_shim.h

--Artem

Why not just call it cuda.h and use #include next for it and then fix it up?

-eric

cuda_runtime.h may be a better choice. nvcc -includes it during both host
and device compilation and this wrapper file is intended to serve a similar
purpose and will probably end up being -included by cc1 in the end.

--Artem

Renamed wrapper to cuda_runtime.h
Similarly to nvcc, automatically add "-include cuda_runtime.h" to CC1 invocations unless -nocudainc is specified.

Changed header wrapping strategy. Previous version was attempting to
make CUDA headers work for host/device compilations separately. In the
end host and device compilations ended up with different view of
CUDA-provided functions. While it mostly worked, that is not what we
really want. What we want is to have identical view of device-specific
functions in both cases and let function overloading handle name clashes
between host and device functions.

This wrapper now always includes CUDA headers exactly the same way during
host and device compilation passes and produces identical preprocessed
content during host and device side compilation for sm_35 GPUs. Device
compilation passes for older GPUs will see a smaller subset of device
functions supported by particular GPU.

As a bonus this wrapper works with CUDA 7.5 now.

Herald added a subscriber: klimek. · View Herald TranscriptOct 20 2015, 12:44 PM

I'm ignoring the content of the header, but this seems to be a not terrible way to do things. I gather that cuda_runtime.h is something that's typically included by the driver by nvidia and not the client?

Also, tests?

-eric

In D13171#272397, @echristo wrote:

I'm ignoring the content of the header, but this seems to be a not terrible way to do things. I gather that cuda_runtime.h is something that's typically included by the driver by nvidia and not the client?

Correct. cuda_runtime.h (and all it pulls in) is -include'd under the hood by nvcc.

Also, tests?

I'll add a test to verify that "-include cuda_runtime.h" shows up on cc1 command line where/when it's expected.

What would be a good way to test the wrapper itself within clang tree without real CUDA headers?

I've done fair amount of manual testing outside of clang source tree.

manual comparison of preprocessed output from cuda_runtime.h between host and device passes.
compiled 39 out of 46 thrust examples and verified that they produce output identical to nvcc-compiled binaries.

In D13171#272441, @tra wrote:

In D13171#272397, @echristo wrote:

I'm ignoring the content of the header, but this seems to be a not terrible way to do things. I gather that cuda_runtime.h is something that's typically included by the driver by nvidia and not the client?

Correct. cuda_runtime.h (and all it pulls in) is -include'd under the hood by nvcc.

Also, tests?

I'll add a test to verify that "-include cuda_runtime.h" shows up on cc1 command line where/when it's expected.

Ick.

What would be a good way to test the wrapper itself within clang tree without real CUDA headers?

Hrm. Maybe a set of inputs that stub out things? Hard really.

I've done fair amount of manual testing outside of clang source tree.

manual comparison of preprocessed output from cuda_runtime.h between host and device passes.

compiled 39 out of 46 thrust examples and verified that they produce output identical to nvcc-compiled binaries.

Cool.

LGTM with those changes and give a thought at how to test this in tree better.

Thanks!

-eric

Added test cases for force-including of cuda_runtime.h
Tweaked inclusion of one header due to use of default arguments.

tra mentioned this in D14370: [doc] Compile CUDA with LLVM.Nov 5 2015, 11:39 AM

jingyue added a subscriber: jingyue.Nov 6 2015, 8:58 AM

Closed by commit rL253388: [CUDA] Added a wrapper header for inclusion of stock CUDA headers. (authored by tra). · Explain WhyNov 17 2015, 2:31 PM

This revision was automatically updated to reflect the committed changes.

Diff 40435

cfe/trunk/lib/Headers/CMakeLists.txt

Show All 11 Lines	set(files
avx512vlintrin.h		avx512vlintrin.h
avx512dqintrin.h		avx512dqintrin.h
avx512vldqintrin.h		avx512vldqintrin.h
avxintrin.h		avxintrin.h
bmi2intrin.h		bmi2intrin.h
bmiintrin.h		bmiintrin.h
cpuid.h		cpuid.h
cuda_builtin_vars.h		cuda_builtin_vars.h
		cuda_runtime.h
emmintrin.h		emmintrin.h
f16cintrin.h		f16cintrin.h
float.h		float.h
fma4intrin.h		fma4intrin.h
fmaintrin.h		fmaintrin.h
fxsrintrin.h		fxsrintrin.h
htmintrin.h		htmintrin.h
htmxlintrin.h		htmxlintrin.h
▲ Show 20 Lines • Show All 78 Lines • Show Last 20 Lines

cfe/trunk/lib/Headers/cuda_runtime.h

				/*===---- cuda_runtime.h - CUDA runtime support ----------------------------===
				*
				* Permission is hereby granted, free of charge, to any person obtaining a copy
				* of this software and associated documentation files (the "Software"), to deal
				* in the Software without restriction, including without limitation the rights
				* to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
				* copies of the Software, and to permit persons to whom the Software is
				* furnished to do so, subject to the following conditions:
				*
				* The above copyright notice and this permission notice shall be included in
				* all copies or substantial portions of the Software.
				*
				* THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
				* IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
				* FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
				* AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
				* LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
				* OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
				* THE SOFTWARE.
				*
				*===-----------------------------------------------------------------------===
				*/

				#ifndef __CLANG_CUDA_RUNTIME_H__
				#define __CLANG_CUDA_RUNTIME_H__

				#if defined(__CUDA__) && defined(__clang__)

				// Include some standard headers to avoid CUDA headers including them
				// while some required macros (like __THROW) are in a weird state.
				#include <stdlib.h>

				// Preserve common macros that will be changed below by us or by CUDA
				// headers.
				#pragma push_macro("__THROW")
				#pragma push_macro("__CUDA_ARCH__")

				// WARNING: Preprocessor hacks below are based on specific of
				// implementation of CUDA-7.x headers and are expected to break with
				// any other version of CUDA headers.
				#include "cuda.h"
				#if !defined(CUDA_VERSION)
				#error "cuda.h did not define CUDA_VERSION"
				#elif CUDA_VERSION < 7000 \|\| CUDA_VERSION > 7050
				#error "Unsupported CUDA version!"
				#endif

				// Make largest subset of device functions available during host
				// compilation -- SM_35 for the time being.
				#ifndef __CUDA_ARCH__
				#define __CUDA_ARCH__ 350
				#endif

				#include "cuda_builtin_vars.h"

				// No need for device_launch_parameters.h as cuda_builtin_vars.h above
				// has taken care of builtin variables declared in the file.
				#define __DEVICE_LAUNCH_PARAMETERS_H__

				// {math,device}_functions.h only have declarations of the
				// functions. We don't need them as we're going to pull in their
				// definitions from .hpp files.
				#define __DEVICE_FUNCTIONS_H__
				#define __MATH_FUNCTIONS_H__

				#undef __CUDACC__
				#define __CUDABE__
				// Disables definitions of device-side runtime support stubs in
				// cuda_device_runtime_api.h
				#define __CUDADEVRT_INTERNAL__
				#include "host_config.h"
				#include "host_defines.h"
				#include "driver_types.h"
				#include "common_functions.h"
				#undef __CUDADEVRT_INTERNAL__

				#undef __CUDABE__
				#define __CUDACC__
				#include_next "cuda_runtime.h"

				#undef __CUDACC__
				#define __CUDABE__

				// CUDA headers use __nvvm_memcpy and __nvvm_memset which clang does
				// not have at the moment. Emulate them with a builtin memcpy/memset.
				#define __nvvm_memcpy(s,d,n,a) __builtin_memcpy(s,d,n)
				#define __nvvm_memset(d,c,n,a) __builtin_memset(d,c,n)

				#include "crt/host_runtime.h"
				#include "crt/device_runtime.h"
				// device_runtime.h defines __cxa_* macros that will conflict with
				// cxxabi.h.
				// FIXME: redefine these as __device__ functions.
				#undef __cxa_vec_ctor
				#undef __cxa_vec_cctor
				#undef __cxa_vec_dtor
				#undef __cxa_vec_new2
				#undef __cxa_vec_new3
				#undef __cxa_vec_delete2
				#undef __cxa_vec_delete
				#undef __cxa_vec_delete3
				#undef __cxa_pure_virtual

				// We need decls for functions in CUDA's libdevice woth __device__
				// attribute only. Alas they come either as __host__ __device__ or
				// with no attributes at all. To work around that, define __CUDA_RTC__
				// which produces HD variant and undef __host__ which gives us desided
				// decls with __device__ attribute.
				#pragma push_macro("__host__")
				#define __host__
				#define __CUDACC_RTC__
				#include "device_functions_decls.h"
				#undef __CUDACC_RTC__

				// Temporarily poison __host__ macro to ensure it's not used by any of
				// the headers we're about to include.
				#define __host__ UNEXPECTED_HOST_ATTRIBUTE

				// device_functions.hpp and math_functions*.hpp use 'static
				// __forceinline__' (with no __device__) for definitions of device
				// functions. Temporarily redefine __forceinline__ to include
				// __device__.
				#pragma push_macro("__forceinline__")
				#define __forceinline__ __device__ __inline__ __attribute__((always_inline))
				#include "device_functions.hpp"
				#include "math_functions.hpp"
				#include "math_functions_dbl_ptx3.hpp"
				#pragma pop_macro("__forceinline__")

				// For some reason single-argument variant is not always declared by
				// CUDA headers. Alas, device_functions.hpp included below needs it.
				static inline __device__ void __brkpt(int c) { __brkpt(); }

				// Now include *.hpp with definitions of various GPU functions. Alas,
				// a lot of thins get declared/defined with __host__ attribute which
				// we don't want and we have to define it out. We also have to include
				// {device,math}_functions.hpp again in order to extract the other
				// branch of #if/else inside.

				#define __host__
				#undef __CUDABE__
				#define __CUDACC__
				#undef __DEVICE_FUNCTIONS_HPP__
				#include "device_functions.hpp"
				#include "device_atomic_functions.hpp"
				#include "sm_20_atomic_functions.hpp"
				#include "sm_32_atomic_functions.hpp"
				#include "sm_20_intrinsics.hpp"
				// sm_30_intrinsics.h has declarations that use default argument, so
				// we have to include it and it will in turn include .hpp
				#include "sm_30_intrinsics.h"
				#include "sm_32_intrinsics.hpp"
				#undef __MATH_FUNCTIONS_HPP__
				#include "math_functions.hpp"
				#pragma pop_macro("__host__")

				#include "texture_indirect_functions.h"

				// Restore state of __CUDA_ARCH__ and __THROW we had on entry.
				#pragma pop_macro("__CUDA_ARCH__")
				#pragma pop_macro("__THROW")

				// Set up compiler macros expected to be seen during compilation.
				#undef __CUDABE__
				#define __CUDACC__
				#define __NVCC__

				#if defined(__CUDA_ARCH__)
				// We need to emit IR declaration for non-existing __nvvm_reflect to
				// let backend know that it should be treated as const nothrow
				// function which is implicitly assumed by NVVMReflect pass.
				extern "C" __device__ __attribute__((const)) int __nvvm_reflect(const void *);
				static __device__ __attribute__((used)) int __nvvm_reflect_anchor() {
				return __nvvm_reflect("NONE");
				}
				#endif

				#endif // __CUDA__
				#endif // __CLANG_CUDA_RUNTIME_H__

This is an archive of the discontinued LLVM Phabricator instance.

[CUDA] Added a wrapper header for inclusion of stock CUDA headers.
ClosedPublic

Details

Diff Detail

Event Timeline

Revision Contents

Diff 40435

cfe/trunk/lib/Headers/CMakeLists.txt

cfe/trunk/lib/Headers/cuda_runtime.h

This is an archive of the discontinued LLVM Phabricator instance.

[CUDA] Added a wrapper header for inclusion of stock CUDA headers.ClosedPublic

Details

Diff Detail

Event Timeline

Revision Contents

Diff 40435

cfe/trunk/lib/Headers/CMakeLists.txt

cfe/trunk/lib/Headers/cuda_runtime.h

[CUDA] Added a wrapper header for inclusion of stock CUDA headers.
ClosedPublic