[AArch64][SVE] Add intrinsics for non-temporal scatters/gathers
This patch adds the following intrinsics for non-temporal gather loads
and scatter stores:
- aarch64_sve_ldnt1_gather_index
- aarch64_sve_stnt1_scatter_index
These intrinsics implement the "scalar + vector of indices" addressing
mode.
As opposed to regular and first-faulting gathers/scatters, there's no
instruction that would take indices and then scale them. Instead, the
indices for non-temporal gathers/scatters are scaled before the
intrinsics are lowered to ldnt1 instructions.
The new ISD nodes, GLDNT1_INDEX and SSTNT1_INDEX, are only used as
placeholders so that we can easily identify the cases implemented in
this patch in performGatherLoadCombine and performScatterStoreCombined.
Once encountered, they are replaced with:
- GLDNT1_INDEX -> SPLAT_VECTOR + SHL + GLDNT1
- SSTNT1_INDEX -> SPLAT_VECTOR + SHL + SSTNT1
The patterns for lowering ISD::SHL for scalable vectors (required by
this patch) were missing, so these are added too.
This change is needed to accommodate for scalable types.