[DropAssumes] Print drop-deref in pipeline where appropriate (#211905)
drop-deref was added as an option in #166947 and is parsed correctly,
but is not serialized. This can make it difficult to reduce test cases,
especially automatically with llvm/utils/reduce_pipeline.py.
This patch implements printPipeline() to serialize drop-deref where
appropriate.
[OpenACC][NFC] Minor clean up in ACCRoutineLowering. (#213333)
Minor NFC clean up after recent changes to remove nohost handling from
ACCRoutineLowering.
Assisted-by: Codex
[IPO] Remove IR Outliner (#211971)
The IR Outliner has major bugs and no active maintainer, and is disabled
by default. The new LLVM Policy states the pass should be removed.
This commit removes:
- The IROutliner pass
- The IRSimilarity analysis
- The `llvm-sim` executable, used for understanding the latter
- All tests of the above
Related discussion:
https://discourse.llvm.org/t/ir-outliner-status-interest/89672
[flang-rt][cuda] Skip scope-exit cleanup when the context has a sticky error (#213184)
A sticky CUDA error (e.g. an illegal memory access in a kernel) leaves
the primary context active but unusable, so CUFDeviceIsActive() reports
it as fine and the compiler-generated scope-exit frees abort a program
that ran to completion: 'cudaFree(p)' failed with
'cudaErrorIllegalAddress'.
Detect this by freeing a null pointer, a no-op that still reports the
sticky error. It runs only once the primary context is known active, so
it cannot lazily create one.
[SLP]Keep scheduled order when moving body in alias-check versioning
The scheduler physically reorders the block before versioning, and
emission insertion points depend on that order. Moving the body by the
pre-scheduling snapshot scrambled it and produced use-before-def IR.
Move in the block's current order instead, filtered to the original
body so the emitted check instructions stay in the header block.
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/213358
[NVPTX] Pass `-Xcuda-ptxas` to the nvlink wrapper for LTO (#213351)
Summary:
This allows `-foffload-lto -fgpu-rdc -Xcuda-ptxas` to work properly by
forwarding it.
[lldb] Avoid returning a stale AddressOf (#212915)
The 'ValueObject::AddressOf()' method assumes that the address of a
value object cannot change, so when it is calculated once, it does not
need to be updated afterwards. However, this is not the case if the
'ValueObject' is a dependent object obtained by calling 'Dereference()'
of another 'ValueObject'. If the latter object is changed, the dependent
value object should return a new address from the 'AddressOf()' method
to reflect the change.
clang: Replace Is*OffloadArch free functions with OffloadArch methods
Drop the IsNVIDIAOffloadArch/IsAMDOffloadArch/IsIntel*OffloadArch free
functions in favor of the OffloadArch member predicate functions.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
clang: Store vendor GPU kinds in OffloadArch instead of re-listing GPUs
OffloadArch was a flat enum that hand-duplicated every AMDGPU and NVPTX
targets, plus a few edge cases. This was yet another place that needed
updating every time a new target is added, which should now be avoided.
Replace with a tagged union-like scheme.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
[AArch64] Fix zero-call-used-regs crash on targets without NEON (#211603)
## Summary
`-fzero-call-used-regs=all` crashes on AArch64 targets without NEON
support, such as `-mgeneral-regs-only` or `-march=armv8-a+nosimd`. The
former is a common configuration used by the Linux kernel.
```c
int p(int a) { return a + 1; }
```
```sh
$ clang --target=aarch64-linux-gnu -O2 -S -mgeneral-regs-only \
-fzero-call-used-regs=all test.c -o -
Assertion failed: (STI.hasNEON() && "Expected to have NEON."),
buildClearRegister
```
[17 lines not shown]
[Clang][Sema] Synthesize a memcpy body for defaulted union assignment (#206579)
A defaulted copy or move assignment operator for a union is synthesized
with an empty body. `DefineImplicitCopyAssignment` /
`DefineImplicitMoveAssignment` skip union members in the memberwise
loop, and the implied copy of the object representation has no AST
representation, a FIXME that has sat at that skip for a long time. The
operator ends up copying nothing.
Classic CodeGen hides this at ordinary call sites by lowering a trivial
assignment to a memcpy at the call site, so `u1 = u2` works even though
the operator body is a no-op. But when the operator is genuinely called,
through a pointer-to-member for instance, it silently copies nothing.
ClangIR calls the assignment operator at the call site rather than
eliding it, so it hits the empty body directly and drops every union
assignment. The no-op is then deleted at `-O3`. That is the MultiSource
`kc` miscompile, where a `YYSTYPE` union assignment (`*++yyvsp =
yylval`) becomes a no-op.
[37 lines not shown]
[libc][__support] Implement exact linear binning for TLSFFreeStoreImpl
Previously, size_to_bit_index used size >> UNIT_SIZE_LOG2 for linear bins,
which mapped size 24 to bin 1 and size 17 to bin 1, causing a mismatch
between allocation request sizes and physical block bucket sizes.
This change introduces exact mapping for linear bins:
- Introduces LINEAR_BINS to compute the exact number of linear bins needed
to reach the exponential table boundary (29 on MSVC Windows where UNIT_SIZE is 8,
31 on Linux where UNIT_SIZE is 16).
- Sizes <= MIN_INNER_SIZE map to bin 0.
- Larger linear sizes map to ((size - MIN_INNER_SIZE - 1) >> UNIT_SIZE_LOG2) + 1.
- Uses LINEAR_BINS as the base index for exponential bins in size_to_bit_index,
index_to_min_size, and find_and_remove_fit, ensuring strict monotonicity
across all indices without runtime clamping.
- Adds NegativeTestForFullHeap unit test in freelist_heap_test.cpp and updates
freestore_test.cpp.
TAG=agy
CONV=cff84e8c-ee22-4f39-af3c-344d1e6f417b
[libc][__support] Clean up index_to_min_size clamping by introducing LINEAR_BINS
Instead of forcing EXP_BASE (32) linear bins and clamping min_size with cpp::min,
this change introduces LINEAR_BINS to calculate the exact number of linear bins
needed to reach the exponential table boundary (29 on MSVC Windows where UNIT_SIZE is 8,
31 on Linux where UNIT_SIZE is 16).
Using LINEAR_BINS as the transition threshold in size_to_bit_index, index_to_min_size,
and find_and_remove_fit eliminates all runtime clamping, padding bins, and overshooting
while preserving strict monotonicity across all indices.
TAG=agy
CONV=cff84e8c-ee22-4f39-af3c-344d1e6f417b
[clang] Follow-up fixes for offsetof unsigned index change (#207808)
- Rename CastNoOverflow to CastAPToOffsetIndex for clarity
- Add parentheses in overflow guard condition in InterpBuiltin.cpp
- Improve comment explaining why 0x8000000000000000 is rejected
AI Tool Use: GitHub Copilot (Claude Sonnet 4.6) was used to assist in
identifying and drafting the fix. The fix was reviewed, tested, and
validated manually.
---------
Co-authored-by: Cadanus da Costa <maccosta at beenox.com>
workflows/release-binaries: Move environment declaration to upload job (#212687)
This is the only job that actually needs to use the environment secrets,
so the environment must be declared. We were using secrets in the
prepare job to do a permissions check, but this is unnecessary, because
that job does not do anything that is security sensitive.
Only the upload job needs to have these permission checks and these are
already included in the upload-release-artifact composite action.
[LoopFusion] Allow loop fusion for idempotent output dependency (#206401)
Loop Fusion incorrectly blocks fusion of loops that write the same value
to two arrays. Writing the same value twice produces identical observable results
regardless of execution order, so fusion is safe (whether the two arrays are aliased or not).
Fixes #94676
Co-authored-by: AntonyCJ30 <cj6186609 at gmail@gmail.com>
[AMDGPU] Do not treat bitcast across FP types as canonicality-preserving (#203560)
isCanonicalized recursed through ISD::BITCAST ignoring the type change,
so value canonical as v2bf16 was wrongly treated as canonical when
bitcast to v2f16 (that has different exponent width), dropping a
required fcanonicalize