[CIR] Implement __builtin_cpu_supports, _init, and _is (#212900)
These are pretty trivial checks to a builtin variable, so this
implements it for x86, as this shows up in self-build. The tests are
pulled from classic-codegen and shows that we do the reasonable thing
for each of them.
[AMDGPU] Exclude CDNA parts from POPS exiting wave id pattern (#210892)
POPS hardware is graphics-pipe-only and absent on compute-only CDNA
targets (gfx908, gfx90a, gfx940, gfx942, gfx950), which incorrectly
matched the isGFX9GFX10 predicate and selected a nonexistent register
[AMDGPU] Fix S_ADD_I32 frame index folding emitting COPY with immediate (#212440)
When eliminating a frame index in `S_ADD_I32 %fi, imm`,
`SIRegisterInfo::eliminateFrameIndex` can simplify `0 + offset` into a
COPY or S_MOV_B32. The pass used a `MachineOperand` reference captured
before `removeOperand()`, which becomes stale after operands are removed
and shifted. That caused the wrong opcode to be selected (`COPY` instead
of `S_MOV_B32`), producing invalid machine IR such as `copy s4, 4`.
The invalid COPY is later hit by Machine Copy Propagation, which asserts
when calling `getReg()` on the immediate source operand.
Co-authored-by: Matt Arsenault <Matthew.Arsenault at amd.com>
[AArch64][SLP][NFC] Precommit scalar fmul extract cost test (#212837)
Precommit a baseline SLP test for the AArch64 scalar fmul extract cost.
The current cost model overestimates the cost of the lane-1 use, so only
one of the two fadd pairs is vectorized at the selected SLP threshold.
A follow-up patch will correct this #212739 should correct this.
[Flang][OpenMP] Remove present modifier application on descriptor (#211856)
This was a minor change upstreamed in the original PR:
https://github.com/llvm/llvm-project/pull/208133
However, it is a modification that needs a little more thought from a
specification perspective before it is rolled out, there's a number of
code bases that depend on the presence modifier being applied only to
the underlying data. However, this leads to inconsistencies when a user
makes use of any reference semantic modifiers when mapping as they
SHOULD be allowed to specify present applying to the descriptor. So, we
need to work out what the correct defualt behaviour is, and regardless
of the default support a user intentionally specifying presence
application on a descriptor via reference semantics.
For now, we will revert to previous state.
[SLP][modularisation][NFC] Extract InstructionsState (2/3) (#211461)
Move the InstructionsState class out of SLPVectorizer.cpp into the
existing private module SLPVectorizer/SLPCompatibilityAnalysis.{h,cpp}.
The class declaration (with trivial accessors) lives in the header; the
non-trivial method bodies are defined out-of-line in the .cpp:
isSameOperation
getMatchingMainOpOrAltOp
isMulDivLikeOp
isAddSubLikeOp
isCopyableElement
isExpandedBinOp
isExpandedOperand
isNonSchedulable
Part of the effort to modularize SLPVectorizer.cpp. See the RFC:
https://discourse.llvm.org/t/modularizing-slpvectorizer-cpp/90922
[libc++] Granularize `<optional>` (#206644)
Certain headers require `optional<T&>`, so it may be beneficial to split
out `optional<T>` and `optional<T&>`. This can allow consumers to only
bring in the `optional` flavour it needs..
- Parcel out the respective pieces into their own header.
- Certain sources rely on transitive includes brought in by
`<optional>`, so they're kept there for now, and only tests have been
fixed.
---------
Co-authored-by: Louis Dionne <ldionne.2 at gmail.com>
[X86] Apply the data32 mode switch in the Intel matcher (#212417)
In .code16, `data32 push 8` in Intel syntax assembled as `pushw $8` with
the 66 prefix dropped, and `data32 push 0x1234` truncated the immediate
to 16 bits. AT&T syntax gets both right.
`ForcedDataPrefix` is set while parsing either syntax, but only
`matchAndEmitATTInstruction` switched mode on it, so the Intel path took
the operand size from the mode and never saw the prefix.
Do the same switch in `matchAndEmitIntelInstruction`. The mode has to go
back to 16-bit before the instruction is emitted, otherwise the 32-bit
form is emitted without its 66 prefix. That function has several error
returns partway through matching, so a scope guard covers those.
Encodings after the change match both AT&T syntax and GNU as:
```
data32 push 8 [0x6a,0x08] -> [0x66,0x6a,0x08]
[3 lines not shown]
[SandboxIR] Fix notifyEraseInstr to skip scheduled neighbors
Guard both loops with !PredN->scheduled() / !SuccN->scheduled() so
scheduled neighbors are left untouched, and add a unit test that erases
a node with one scheduled and one unscheduled predecessor to cover the
fix.
[Clang][AIX] Restrict -mloadtime-comment-vars to file/namespace scope
Support only file- and namespace-scope variables. Name-matched static
data members, variable template specializations (explicit ones
included), and function-local statics are now diagnosed with
-Wloadtime-comment-var instead of being silently ignored. Implicit
instantiations are diagnosed via the template-instantiation path, once
per instantiating TU. Automatic locals have no symbol to match and
remain out of scope.
[NFC][HLSL] Fix msan errors in tests (#212901)
Some of the tests were not initializing all the fields prior to
generating metadata, this caused a read of uninitialized memory and
caused the sanitizer to fail.
Assisted by: Claude Opus 5
Caught here: https://lab.llvm.org/buildbot/#/builders/94/builds/19767
[libc] Modernize and extend dirent.h header. (#212902)
Extend the `<dirent.h>` header with macro and types specified in recent
POSIX.1-2024:
* Add `posix_dent` structure, which has more fields than `dirent`, that
are actually used in practice. This struct would be identical to
`dirent` that we have on Linux
* Add `reclen_t` type for `d_reclen` field.
* Add macro `DT_BLK` and friends
Also, extend the tests to verify the values of `d_type` field, now that
we have the proper macro defined.
Assisted by: Gemini, human-verified
[llvm-ml] make TEXTEQU directive not to eagerly expand macros in arguments (#209526)
We observed a crash in `TEXTEQU` that pastes two macros into one.
```masm
.data
part1 TEXTEQU <1>
part2 TEXTEQU <0>
joined TEXTEQU part1, part2 ; crash
```
`part1` is immediately rewritten into `1` as integer, which is rejected
by `TEXTEQU` parser. We need to keep `part1` as identifier for `TEXTEQU`
to pick up later.
[X86] Make WinEH crash test reliable under ASan (#212820)
The WinEH unwind `v2 error` test uses `not --crash` for malformed MIR
inputs.
In AddressSanitizer builds, the default `abort_on_error=0` can prevent
the
expected fatal error from being reported as a crash, causing FileCheck
to
receive no diagnostic output.
Set `ASAN_OPTIONS=abort_on_error=1` for the expected-crash invocations
in
`win64-eh-unwindv2-errors.mir`.
## Testing
- ASan build focused test passes.
- Non-ASan build focused test passes.
- Five additional serial ASan reruns pass.
[4 lines not shown]
[CIR][CUDA] Add support for NVVM xchg builtins (#211815)
Adds codegen support for the scoped and unscoped NVVM atomic exchange
builtins:
`atom_xchg,` `atom_cta_xchg,` and `atom_sys_xchg.`
These are lowered to the corresponding CIR `cir.atomic.xchg` operations
and subsequently lowered to LLVM `atomicrmw xchg` instructions.
[CIR] Implement complex rvalues NYI (#211645)
emitReturnOfRValue had an NYI for _Complex types, so returning one as an
Rvalue (see example of a lambda invoker) would NYI. Since the logic to
the store is already handled in the lower-to-LLVM, this ended up being a
pretty trivial patch.
Note; There are some differences in how this lowers, because our calling
convention for the ret is different here, and we maintain the 'complex'
type differently even through LLVM-IR. However, the IR looks to be
equivilent.
[MacroFusion] Add SDep param to predicates(NFC) (#212255)
This patch aims to extend the API for macro fusion predicates with an
additional SDep param which allows each predicate to individually decide
wether a pair should be macro fused based on the kind of dependency
between the 2 instructions.
A followup patch https://github.com/llvm/llvm-project/pull/212603
introduces a real user in AArch64.
[HLSL] Add sema for use of samplers and gathers on textures of doubles and ints (#212613)
Fixes https://github.com/llvm/llvm-project/issues/198882 and
https://github.com/llvm/llvm-project/issues/198883
This PR:
- Implements sema checks to reject the use of samplers and gathers on
textures of doubles.
- Implements sema checks to reject use of samplers on textures of
integers before shader model 6.7
Assisted by: Claude Opus 5