[Bazel] Fixes for commit b8afba677502 (#227808)
Add missing //llvm:CodeGen, //llvm:Core, and //llvm:Demangle
dependencies to clang:codegen_utils, and add :codegen_utils to
clang:cir_frontend_action.
Flang improve fidelity of unformatted I/O endianness with FORT_CONVERT_UNIT (#223831)
Add further fidelity to specifying desired endianness of unformatted I/O
file on a per unit basis with the `FORT_CONVERT_UNIT` environment
variable.
To reviewers (friends) I'm not sure about:
1. The correct function prefixing for the FLANG runtime environment.
2. How to test/validate on Windows
3. How to test/validate with CUDA runtime support enabled
Thank you.
https://discourse.llvm.org/t/runtime-unformatted-i-o-conversion-between-big-and-little-endian-formats/91751
Documentation for new environment variable:
[85 lines not shown]
[Clang] Retain constructor/destructor variants when symbol must be kept (#226572)
With -mconstructor-aliases, complete constructor and destructor variants
with discardable-if-unused linkage can be silently replaced in the IR
(RAUW) rather than emitted as distinct symbols. This prevents
-fkeep-inline-functions and `__attribute__((used))` from retaining the
complete (C1/D1) variants.
Skip RAUW when the declaration requires its symbol to be kept by
introducing structorSymbolMustBeRetained(), which returns true when`
__attribute__((used)) `is present or -fkeep-inline-functions is active
for an inline definition that is not available_externally.
Assisted-by: IBM Bob
[RISC-V][MC] Reject x0 as the address/temporary register for load/store pseudos (#227512)
Using the zero register as the destination for integer load pseudos or as
the temporary register for floating-point load and store pseudos is
illegal since it is used to synthesize the target address, and using x0
would mean the auipc/qc.e.li result is ignored and we access an invalid
location.
This commit was created with the help of AI tools
Pull-Request: https://github.com/llvm/llvm-project/pull/227512
[InstCombine] Fold exp2(uitofp x) to ldexp(1.0, x) with ninf
`exp2(uitofp iN x) -> ldexp(1.0, zext x)` currently requires N to be
narrower than `int`, because `ldexp` takes a signed `int` exponent.
When N equals the width of `int`, the fold is still correct under `ninf`.
Every LLVM FP type has $E_{\max} \le 16383 < 2^{15} \le 2^{N-1}$, so:
- $x < 2^{N-1}$: signed and unsigned $x$ agree, so $\operatorname{ldexp}(1.0, x) = 2^{x} = \operatorname{exp2}(x)$.
- $x \ge 2^{N-1}$: $\operatorname{exp2}(x) \ge 2^{2^{N-1}} > 2^{E_{\max}}$ overflows to $+\infty$, which is poison under `ninf`.
This only applies to the intrinsic, since the `exp2` libcall may set
`errno` on overflow.
This also enables `pow(2.0, uitofp x) -> ldexp` via the existing
`pow(2^n, x) -> exp2(n * x)` fold.
[InstCombine] Pre-commit tests for exp2(uitofp) -> ldexp with ninf. NFC
Add baseline tests for folding `exp2(uitofp iN x)` to `ldexp(1.0, x)`
when N equals the width of C int, including negative tests (no ninf,
source wider than int, libcall).
[clang][CodeGen] Fix stack-use-after-return in deferred annotations (#226942)
CodeGenModule::DeferredAnnotations was keyed by StringRef, but not every
mangled name passed to GetOrCreateLLVMFunction outlives the call.
CodeGenVTables::maybeEmitThunk mangles the thunk name into a stack-local
SmallString and hands it to GetAddrOfThunk, so for an annotated virtual
function the map retained a reference into a frame that was gone by the
time EmitGlobalAnnotations looked the key up.
ASan reports this as a stack-use-after-return in EmitGlobalAnnotations,
with the freed frame being maybeEmitThunk's 'Name'. The user-visible
effect is that annotations silently disappear from the this-adjusting
thunks once the dead frame has been reused.
Make the key own its storage, using StringMap<unsigned> as the MapVector
map type so lookups still hash a StringRef without allocating.
The clang/test/CodeGenCXX/attr-annotate-member-functions.cpp test has
been created with help of an AI.
[SelectionDAG] Take phi op divergence from the op not the phi (#220853)
When creating registers for phi operands, take the type and the
IsDivergent bit from the operand not from the result of the phi node.
In practice this only affects AMDGPU and only affects constants used as
operands to a divergent phi. The effect is to put the constant into an
SGPR even if the result of the phi is a VGPR.
As well as fixing a correctness problem that arose when enabling
-structurizecfg-skip-uniform-regions, this has a small overall effect in
increasing SALU usage but reducing VALU usage. Overall register usage is
generally unaffected, except in some regalloc stress tests.
Co-authored-by: Claude Opus 5 (1M context) <noreply at anthropic.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply at anthropic.com>
[CIR][NFC] Add missing enum elements to TypeInfoIsInStandardLibrary (#227781)
Fix warning for missing enum element in the TypeInfoIsInStandardLibrary
helper function after #224277 is merged
[Bazel] Add Analysis dependency to LLVMToLLVMIRTranslation (#227785)
Fixes LLVMToLLVMIRTranslation bazel layering failure introduced in
commit 233a25d068a9, where LLVMToLLVMIRTranslation.cpp includes
llvm/Analysis/ConstantFolding.h.
[Matrix] Implement matrix support for the `abs` intrinsic (#227131)
Closes #184491.
This PR implements the matrix api for `abs` in `HLSLintrinsics.td`, adds
matrix codegen tests, matrix sema tests, and SPIRV matrix backend tests.
DirectX matrix backend tests were not added because no DirectX backend
changes were made.
Assisted-by: Claude Opus 4.8
[InstCombine] Fold icmp eq/ne X, select(icmp pred X, P, C1, C2) to set membership (#226756)
## Summary
- Add InstCombine fold for `icmp eq/ne X, select(icmp pred X, P, C1,
C2)` pattern
- Transforms self-referential select comparisons into simple set
membership tests
- Handles all icmp predicates (signed, unsigned) and commuted operands
## Details
When we have code like:
```llvm
%cond = icmp sgt i32 %x, 0
%s = select i1 %cond, i32 2, i32 0
%r = icmp eq i32 %x, %s
The comparison can only be true when X equals one of the select's constant arms. This patch recognizes when:
1. The select condition compares X with a constant P
[19 lines not shown]
[InstCombine] Fix select combine with null_pointer_is_valid. (#227168)
When null_pointer_is_valid is set, make sure we don't treat an inbounds
gep of "ptr null" as poison.
While I'm here, also write out the justification of
simplifyNonNullOperand, since it's pretty unintuitive.
Fixes #227002
[CIR] Lower variadic arguments through an indirect call (#227477)
On x86_64, CallConvLowering now classifies an indirect call with
ellipsis arguments from its own operands, as it does a direct variadic
call. Those arguments stay out of the retyped callee's function type, as
in classic CodeGen.
The cir.call verifier now checks indirect calls against the callee
pointer's function type, and CIRGen's asm-label redirect keeps a
variadic declaration's ellipsis instead of dropping it.
Assisted-by: Claude Code / claude-opus-5-5
[libc++] Inline the build-at-commit composite action into the benchmark workflows (#227750)
Yet another twist in the endless saga of setting up performance
benchmarking infrastructure for libc++. Since the recent switch to
Kubernetes runners, there are now two layers of "pods". First, there's
the runner pod which executes the Github action itself. It's the one
that loads the Docker image specified in the Github workflow and then
launches it.
Then, there's the workflow/job pod, which executes the actual Docker
image and the commands described in `steps` in the workflow file.
This causes problems because the workflow pod is the only one that has
access to the checked out monorepo (since it's checked out within the
running Docker container). This means the runner pod does not have
access to composite actions stored in the repository, which leads to
errors like:
Can't find 'action.yml', 'action.yaml' or 'Dockerfile' under '/home/runner/_work/<...>/libcxx/build-at-commit'.
[3 lines not shown]
[Option] Shrink Info from 24 to 16 bytes (#227582)
Move Flags, Visibility, and Param (getNumArgs() for `MultiArg`) into
`InfoExtra` and narrow `PrefixesOffset` to 8 bits. Outside clang, Flags
and Visibility take a handful of values per table; clang has 436
distinct InfoExtra rows for 3871 options.
clang's Info and InfoExtra tables shrink from 96,968 to 72,400 bytes.
LLM-aided
[flang][cuda] Lower on_device() declared by an interface to cuf.on_device (#227498)
on_device() was lowered to cuf.on_device only when it came from the
cudadevice module. A program that declares on_device() itself, through
an
interface body or an EXTERNAL statement (for example with !$acc routine
seq),
got a plain call to on_device_ that nothing folded, and linking failed.
Move on_device into the regular CUDA intrinsic handler table and drop
the
separate BIND(C) table. Remove bind(c) from the on_device interface in
the
cudadevice module, since the handler now does the lowering and the
binding
label is no longer used.
In a CUDA Fortran or OpenACC compilation, look up the CUDA intrinsic
handlers
[5 lines not shown]
[msan] handle llvm.masked.{udiv,sdiv,urem,srem} intrinsics (#225363)
Check the divisor and propagate the dividend's shadow on active lanes,
poisoning disabled result lanes. This poisoning is consistent with
the definition of these intrinsics; see #189705 (9ba774566).
This gets rid of false positives from inactive input lanes since
visitInstruction() was checking every operand across all lanes,
active or not.
Depends on: https://github.com/llvm/llvm-project/pull/227462
[CIR] Preserve padding bytes across a byval argument (#224672)
Fill the parameter's own slot in the callee with a copy from the
object's address rather than a store of a loaded record value. At the
call site, pass the object's own address, or fill the byval slot the
same way when it cannot be passed. A store writes only the fields of the
record's type and leaves its padding unwritten, which for a union can be
another member's live data. A _BitInt keeps its load and store, since
those are its conversion to and from memory.
Assisted-by: Cursor / claude-opus-5
[SeparateConstOffsetFromGEP] Fix offsets across trunc and ext (#225195)
Stop extracting offsets from arithmetic through a truncation followed by
sign or zero extension.
[Compiler Explorer reproducer](https://godbolt.org/z/xK1WTT7nY) and
[Alive2 counterexample](https://alive2.llvm.org/ce/z/KWcMNZ).
[CIR] Don't 'simplify' switch-case if there is something in the way. (#227437)
The simplify case code was assuming that cases were mergeable, and
didn't realize that code in the middle (like the goto in the test case)
meant they couldn't be re-joined. This patch adds that to the
CIRSimplify logic to make sure we con't try to combine cases that have
something 'in between'.
[CodeGen][CIR] Split out Backend Diags to utils, impl CIR error-attr (#225966)
So first, the motivation for this is the `error` attribute, which
applies to a function and diagnoses if it 'survives' into the backend as
a call. In order to do this, we need to add the
dontcall-error/dontcall-warn attributes to functions.
The call operation ALSO needs the srcloc to be present, so this threads
that through as well. Note this is pretty fragile in classic-codegen, so
we inherit some fragility from that source location as well.
HOWEVER, in order to get diagnostics to work properly in the backend, we
need to consume them. This patch ALSO extracts the handling for that
from CodeGenAction.cpp into CodeGenUtils, plus uses it from both sides.
Disclaimer: I ended up using Claude to do a lot of the refactoring. It
was mostly copy/paste with some minor changes to generalize it, but
Cladue helped.
[CIR] Fix failing tests after #226904 (#227523)
A number of CIR tests were checking LLVM IR output for GEP statements
that contained `inbounds` and `nuw` flags that had been added by LLVM's
constant folder. Following
https://github.com/llvm/llvm-project/pull/226904 these are no longer
produced.
This change updates the tests to match the current output.
Assisted-by: Cursor / claude-opus-5.5
[CodeGen] Remove unused InstructionMoveBefore (NFC) (#227586)
The last use was removed on December 15, 2023 in commit
163aeca33d4adb97e8599584409457ca14b1419b.
Assisted-by: Antigravity
[Clang][HLSL] Emit constant matrix values in column-major order (#227356)
Emit matrix APValues in the canonical column-major register layout,
independent of the selected matrix memory layout.
When a constant is used as a row-major memory initializer, reorder its
elements at the register-to-memory boundary. This preserves the physical
layout of global and static row-major matrices while ensuring constexpr
matrix expressions produce canonical register values.
Assisted by Cop-pilot GPT 5.6-Sol
[REPL][Windows] Don't redraw the prompt into redirected REPL output (#227755)
On Windows without `libedit`, `IOHandlerEditline::PrintAsync` re-printed
the prompt after every async output chunk, even when stdout is not a
console. When REPL output is redirected, as in lit tests, `"> "` lands
in the middle of the program's output, which made `SwiftREPL` tests like
`ExclusivityREPL` flaky.
The prompt is now redrawn only when the handler is interactive and
stdout is a console that can erase it.
[lldb] Fix logic bugs in BytecodeSyntheticChildren (#227483)
Fix a pair of bugs in `BytecodeSyntheticChildren`.
1. Fix conditions that should compare with `>= 0` (not `> 0`)
2. Fix an inverted conditional when checking if a formatter implements
`@get_child_index`
Assisted-by: claude