[X86] Fix ADOX miscompile by restricting COND_O optimization when EFLAGS are used (#220117)
This patch fixes a miscompile where the X86 DAGCombiner aggressively
folds an `ADD` node into an `ADOX` instruction even when the Zero Flag
(ZF) produced by the `ADD` is used by a subsequent branch (e.g., `je`).
[VE] Create TS1AM with getMemIntrinsicNode (#220151)
isa<AtomicSDNode> returns false on the built node: the lookup key and
SDNode::Profile disagree and the node never CSEs (see #219911).
CodeGen/VE/Scalar/atomic_swap.ll will break when we add an assert to
FoldingSet::insert that the `nodeProfile()` output matches `Token`.
Fix with getMemIntrinsicNode (`def ts1am` carries SDNPMemOperand).
[Clang][CodeGen][X86] Fix crash on __int128 bit-field access units (#216777)
Fixes #202205
The x86-64 SysV classifier skipped every unnamed bit-field as padding,
so an eightbyte holding nothing but a non-zero-width unnamed bit-field
stayed `NO_CLASS`. GCC treats that storage as INTEGER. The crash falls
out of this: a run of `__int128` bit-fields is lowered to a single
`i128` access unit spanning both eightbytes, but with only one of them
classified INTEGER the `i128` gets queried at offset 8 and hits
`assert(IROffset == 0)` in `GetINTEGERTypeAtOffset` — or `assert(Hi ==
Integer)` in the callers, depending on which eightbyte holds the named
field. It also silently diverges from GCC on ordinary shapes like
`struct { long : 64; long a; }`, which clang passed in one register
where GCC uses two.
The fix skips only zero-length bit-fields and classifies the rest like
named ones, matching GCC. Both eightbytes of an `__int128` bit-field run
then come out `INTEGER`, so the existing asserts hold unchanged. The
[9 lines not shown]
[flang][OpenMP] Fix use_device_addr handling for COMMON blocks in target data (#217105)
Related to #217112
Flang already handles `use_device_addr` for regular variables on `target
data`,
but named COMMON blocks had two gaps: combining a COMMON block in `map`
and
`use_device_addr` was incorrectly rejected, and whole-COMMON
`use_device_addr`
could leave references in the region bound to the host COMMON instead of
the
returned device address.
This fixes the COMMON-block semantic handling and lowering so the valid
`map`/`use_device_addr` combination is accepted and COMMON members use
the
correct target-data bindings.
[3 lines not shown]
[NFC][DirectX] Fix memory leaks exposed by ASAN (#220152)
This PR should fix ASAN errors from
https://lab.llvm.org/buildbot/#/builders/52/builds/19811
There are three sources of leaks:
1. TargetPassConfig not being `PM.add()`ed
```
Direct leak of 136 byte(s) in 1 object(s) allocated from:
#1 ...createPassConfig(...) DirectXTargetMachine.cpp:216:10 <- return new DirectXPassConfig(*this, PM);
#2 ...addPassesToEmitFile(...) DirectXTargetMachine.cpp:177:34 <- TargetPassConfig *PassConfig = createPassConfig(PM);
Indirect leak of 144 byte(s) in 1 object(s) allocated from:
#1 llvm::TargetPassConfig::TargetPassConfig(...) TargetPassConfig.cpp:604:10 <- Impl = new PassConfigImpl();
#2 DirectXPassConfig DirectXTargetMachine.cpp:112:9
```
2. MachineModuleInfoWrapperPass not being `PM.add()`ed
```
[18 lines not shown]
[OpenMP][libomp] Fix dist barrier arrival synchronization (#213845)
Make distributedBarrier::stillNeed atomic and use release/acquire
ordering for distributed barrier gather arrival flags.
The old volatile stillNeed flag did not synchronize an arriving thread's
pre-barrier writes with the thread that observed its arrival. On weakly
ordered architectures, a group leader could observe stillNeed == 0
before the arriving thread's pre-barrier writes were visible. This
allowed another thread to pass the barrier and read stale data written
before the barrier.
Use release stores when publishing stillNeed == 0. Keep the spin loops
on relaxed loads, then perform one acquire fence after all expected zero
values have been observed. This connects the arriving threads'
pre-barrier writes to the observer through the standard release/acquire
happens-before chain, without using acquire loads on every poll.
The same pattern is used when a group leader publishes its own stillNeed
[106 lines not shown]
Fix layering violation from #220073 (#220150)
That PR introduced circular dependencies LLVMTransformUtils <->
LLVMPasses. Move Trigger*CrashPasses into LLVMPasses.
Move TriggerCrashFunctionLegacyPass alongside the one usage
`-codegen-pipeline-trigger-crash` just to prevent a tiny .cpp file.
[SPIR-V] Always reject OpTypeVector with a non-standard width (#212685)
SPV_EXT_long_vector does not widen OpTypeVector past 4 components. Those
widths need OpTypeVectorIdEXT, so requesting the extension here was
wrong
[ADT][TableGen] Add UniquingSet, a FoldingSet with typed keys (#219630)
FoldingSet serializes a key into a FoldingSetNodeID to look a node up
and rebuilds the stored node's profile to compare against it. Where a
key can be read out of a node, neither is necessary.
UniquingSet reuses FoldingSetBase's storage, growth, removal and insert
token and replaces only the key: the node's `getKey()` supplies it, the
key type's `operator==` compares it, and DenseMapInfo hashes it inline.
An Info parameter overrides the key type or its hash. The hash cached on
each node keeps growth and erasure from calling `getKey()`, which a
DenseSet cannot avoid.
Prefer `UniquingSet` where a key can be read out of a node in O(1) and
the lookup key is built beside `getKey()`; keep FoldingSet for keys that
are wide, polymorphic or assembled at many call sites, where one Profile
helper keeps both sides consistent. insert asserts that a node hashes as
its lookup did.
[18 lines not shown]
[docs] [C++20] [Modules] Mentioning tricks to use std module without touching the code (#220147)
This commit introduces two tricks to use std module without changing
user's code.
Revert "[CSSPGO] Don't let pseudo probes block early-exit vectorization" (#220146)
Reverts llvm/llvm-project#219872
This was accidentally merged without approval
[clang][docs] Suppress Sphinx highlighting failure warnings in conf.py (#220118)
Following the upgrade to Sphinx 8.2 (#219299), Pygments syntax
highlighting fallbacks (e.g. on custom C++ attribute syntaxes like
`[[clang::...]]`) emit `[misc.highlighting_failure]` warnings when
retrying in relaxed mode. Because sphinx-build runs with `-W`, these
warnings abort the documentation build.
Mirror the configuration in `llvm/docs/conf.py` by setting
`suppress_warnings = ["misc.highlighting_failure"]`.
AI tool usage: An AI assistant was used to help research and draft the
documentation updates.
[mlir][arith] Expand ops for F8E4M3FN and F8E5M2 type. (#216653)
Patch to add support for arith op (`arith.truncf`) to truncate `f32/f16`
type to `f8E4M3FN/f8E5M2`.
[lld][MachO] Compare referent offsets in ICF (#219960)
ICF::equalsConstant() compared only the addends of relocations that
reference symbols in ConcatInputSections, not the symbols' offsets
within their sections. The ICF hash merely sums the referent symbols'
offsets, so two code sections whose relocations reference the same input
section at swapped offsets have equal hashes, pass both the constant and
the variable comparison, and are folded together even though they
reference different data.
Compare the symbol value plus the addend instead, as the ELF backend
does.
This issue was found by LLM while investigating why identical LSDAs were
not being folded. It has not been observed in a real-world link, and the
regression test is synthetic.
[flang][cuda] Propagate CUDA attrs from parent variable to component deallocs (#220059)
This is a follow up to
[#206614](https://github.com/llvm/llvm-project/pull/206614), which made
`allocate(foo(i)%arr(...))` inherit `foo`'s CUDA memory attribute, but
did not add the equivalent inheritance for `deallocate(foo(i)%arr)`. The
deallocation still lowered to an inlined `fir.freemem`, so memory
obtained from a CUDA allocator was released with libc `free()`.
The allocate side already walked the `DataRef` chain for a
CUDA-attributed parent, but the helpers were private to
`AllocateStmtHelper` and unreachable from the deallocate path. This
patch hoists `findCUDAAttrInDataRef` to file scope, adds
`getCUDAAttrParentSymbol(AllocateObject)` beside it, and reduces the
existing member to a thin wrapper. The allocate behavior is unchanged.
`genDeallocate` gains an optional `cudaSymbol` used only for the CUDA
decisions (`isCudaSymbol` and the `genCudaDeallocate` call).
`genDeallocateStmt` supplies the parent symbol, which is non-null only
[9 lines not shown]
[WebAssembly] Mark SIMD min and max as commutable (#219812)
Vector add/mul and scalar floating-point min/max are already marked as
commutable. This extends the same property to floating-point vector
min/max, allowing better WebAssembly register stackification.
Should avoid any locals as per what's happening now
```
.local v128
call red
local.set 0
call green
local.get 0
f32x4.min
```
[SelectionDAG] Fix CSE keys that disagree with SDNode::Profile (#219911)
A getNode helper builds its lookup ID by hand; matching it later
rebuilds one from the node with AddNodeIDCustom. Where the two disagree
the compare always fails and the node never CSEs. Fix whichever side is
wrong: the labels, DEACTIVATION_SYMBOL, GET/SET_FPENV_MEM and
EXPERIMENTAL_VECTOR_HISTOGRAM have no case; getLifetimeNode keys on a
frame index operand 1 already carries, getStridedLoadVP on the result
type instead of the memory type, and getPseudoProbeNode drops the
attributes its case profiles.
Ask AtomicSDNode instead of an opcode list stale since ATOMIC_LOAD_FADD,
and add the two opcodes its own classof was missing.
AddNodeIDCustom now takes the opcode to profile under, so MorphNodeTo's
pre-morph lookup keys on what the morph produces. Machine opcodes
profile nothing: the morph overlays MachineSDNode's memory references on
the fields the MemSDNode checks read.
[2 lines not shown]
[NFC][HLSL] Refactor texture type declaration (#219561)
Fixes https://github.com/llvm/llvm-project/issues/219542
Refactors texture type declaration in `HLSLExternalSemaSource.cpp` so
that new
texture types can more easily be added without adding a bunch of new
helper
functions.
This is accomplished with the introduction of a new `TextureTypeInfo`
struct
to record the properties of each texture type, as well as its
capabilities
indicated by the `TexCap` bitmask enum.
Adding a new texture type to be declared should, in most cases, only
require
appending a new entry to the static `TextureTypes` array of
[12 lines not shown]
[CSSPGO] Don't let pseudo probes block early-exit vectorization (#219872)
llvm.pseudoprobe is modeled as accessing inaccessible memory, so
mayReadFromMemory()/mayWriteToMemory() return true even though the
intrinsic
carries no real memory dependence. An otherwise vectorizable early-exit
loop
is therefore rejected as soon as it contains a pseudo probe.
This patch skips pseudo probes in isVectorizableEarlyExitLoop(),
isReadOnlyLoop() and
areAllLoadsDereferenceable() so the three checks agree and such loops
vectorize as they would without pseudo probe instrumentation.
Discussion:
https://discourse.llvm.org/t/csspgo-unblocking-pseudo-probe-safe-optimizations/90946
[ORC] Add SymbolStringPtr overloads for recordAddr/recordProxy (#220125)
Allow clients to pass symbol names as SymbolStringPtrs (in addition to
StringRefs).
[AMDGPU] Allow combining uniform OR/AND to V_PERM (#220048)
Allow the OR/AND -> V_PERM DAG combine for values, even if they are
uniform.
Co-authored by Brendon Cahoon and Cursor
[Clang] Honor -Xarch_gfx* when linking the UBSan offload runtime
Empty bound architecture misses per-GPU sanitizer flags, so inspect each
offload arch when deciding whether the host interceptor is required.
[Clang] Enable UBSan for AMDGPU device offload
Summary:
This enables the device UBSan runtime for AMDGPU decides. Primarily this
required modifications to the `addSanitizerRuntime` interface so we can
query the compilation's offload status. Also need to forward it through
the linker wrapper interface. Works on all AMDGPU offload, slight hacks
around the other targets as they do not advertise sanitizer
runtimes properly.
This is linked in via a new `-u __ubsan_device_initialize` hook to pull
in the side library. This is standard behavior and keeps the core logic
mostly unchanged and re-used.
[LoongArch] Add omitted LASX patterns for vector extend (#219351)
Adds omitted 128-bit to 256-bit patterns for `sign_extend_vector_inreg`,
including `v16i8 -> v4i64` and `v8i16 -> v4i64`, which will generate by
the combination of `icmp + or/and/xor + zext/sext`, all related tests
are added.
Fix: https://github.com/llvm/llvm-project/issues/219224