[LoongArch] Use [X]VSETNEZ.V/[X]VSETEQZ.V for whole-vector zero check (#226814)
Introduce `loongarch_vanynonzero` and `loongarch_vallzero` for checking
wether a vector has any non zero or all of them are zero, which could
simply some vector comparison such as the one below.
Before
```
xvseq.d $xr0, $xr0, $xr1
xvmskltz.d $xr0, $xr0
xvpickve2gr.wu $a0, $xr0, 0
xvpickve2gr.wu $a1, $xr0, 4
bstrins.d $a0, $a1, 3, 2
beqz $a0, .LBB2_2
```
After
```
xvseq.d $xr0, $xr0, $xr1
xvseteqz.v $fcc0, $xr0
[4 lines not shown]
[LoongArch] Use [X]VSETNEZ.V/[X]VSETEQZ.V for whole-vector zero check (#226814)
Introduce `loongarch_vanynonzero` and `loongarch_vallzero` for checking
wether a vector has any non zero or all of them are zero, which could
simply some vector comparison such as the one below.
Before
```
xvseq.d $xr0, $xr0, $xr1
xvmskltz.d $xr0, $xr0
xvpickve2gr.wu $a0, $xr0, 0
xvpickve2gr.wu $a1, $xr0, 4
bstrins.d $a0, $a1, 3, 2
beqz $a0, .LBB2_2
```
After
```
xvseq.d $xr0, $xr0, $xr1
xvseteqz.v $fcc0, $xr0
[4 lines not shown]
[CodeGenPrepare] Prefix option names with cgp-
Rename the CodeGenPrepare options so that they start with cgp-, and turn
the -disable-* options into positive options that default to true:
```
-addr-sink-* -cgp-addr-sink-*
-cgpp-huge-func -cgp-huge-func
-disable-cgp-X -cgp-X=0
-disable-complex-addr-modes -cgp-complex-addr-modes=0
-disable-preheader-prot -cgp-preheader-prot=0
-enable-andcmp-sinking -cgp-andcmp-sinking
-force-split-store -cgp-force-split-store
-stress-cgp-X -cgp-stress-X
```
The section prefix options (-profile-guided-section-prefix, etc.) are
kept, as they are not specific to CodeGenPrepare.
Aided by Opus 5.5
[CodeGen] Declare command line options in TableGen (#228911)
Move the cl::opts of TargetPassConfig.cpp and CodeGenPrepare.cpp into
CodeGenOptions.td, private to lib/CodeGen; other files will follow.
-enable-machine-outliner (cl::ValueOptional) and -regalloc
(RegisterPassParser) stay cl::opt.
CodeGenPrepare and its addressing-mode helpers hold
`const CodeGenOptions &Opts`; TargetPassConfig functions read
CodeGenOptions::Global. getCGPassBuilderOption() converts BoolOrDefault
members to the cl::boolOrDefault and std::optional<bool> fields of the
public CGPassBuilderOption. -basic-block-section-match-infer, which was
not cl::Hidden, is now listed by -help-hidden only.
Aided by Opus 5.5
[Mips] Add microMIPS CP0 move encodings (#228884)
Define the microMIPS MFC0 and MTC0 formats and instruction records for
pre-R6 targets, including the select field. Accept the two-operand forms
with an implicit select value of zero and add generic scheduling
entries.
Extend the control-instruction tests with explicit and implicit select
operands, checking both byte orders.
Assisted-by: OpenAI Codex # Testcases
[CodeGenPrepare] Prefix option names with cgp-
Rename the CodeGenPrepare options so that they start with cgp-, and turn
the -disable-* options into positive options that default to true:
```
-addr-sink-* -cgp-addr-sink-*
-cgpp-huge-func -cgp-huge-func
-disable-cgp-X -cgp-X=0
-disable-complex-addr-modes -cgp-complex-addr-modes=0
-disable-preheader-prot -cgp-preheader-prot=0
-enable-andcmp-sinking -cgp-andcmp-sinking
-force-split-store -cgp-force-split-store
-stress-cgp-X -cgp-stress-X
```
The section prefix options (-profile-guided-section-prefix, etc.) are
kept, as they are not specific to CodeGenPrepare.
Aided by Opus 5.5
[CodeGen] Declare command line options in TableGen
Move the cl::opts of TargetPassConfig.cpp and CodeGenPrepare.cpp into
CodeGenOptions.td, private to lib/CodeGen; other files will follow.
-enable-machine-outliner (cl::ValueOptional) and -regalloc
(RegisterPassParser) stay cl::opt.
CodeGenPrepare and its addressing-mode helpers hold
`const CodeGenOptions &Opts`; TargetPassConfig functions read
CodeGenOptions::Global. getCGPassBuilderOption() converts BoolOrDefault
members to the cl::boolOrDefault and std::optional<bool> fields of the
public CGPassBuilderOption. -basic-block-section-match-infer, which was
not cl::Hidden, is now listed by -help-hidden only.
Aided by Opus 5.5
[VPlan] Support pointer min/max bounds in VPlan memory runtime checks. (#225828)
Expand pointer-typed SCEV min/max expressions in VPSCEVExpander as
icmp + select (including profile metadata), matching SCEVExpander.
This allows modeling memory runtime checks with pointer min/max bounds
in VPlan, e.g. for accesses with a runtime stride of unknown sign.
Depends on https://github.com/llvm/llvm-project/pull/221483
PR: https://github.com/llvm/llvm-project/pull/225828
[Mips] Add microMIPS hazard-barrier jump encodings (#228885)
Add microMIPS encodings for jr.hb and jalr.hb. Use BaseOpcode and MMRel
to map the standard opcodes to their microMIPS encodings.
Assisted-by: OpenAI Codex # Testcases
[mlir][acc] Declare the memory effects of acc.atomic.read
The operation read `x` and wrote `v` but declared no memory effects at
all, so every memory analysis had to treat both locations as unknown and
stay maximally conservative. In particular it could not be seen to
overwrite `v`.
Declare the effects per operand so the read on `x` and the write on `v`
are visible.
[OCaml] Use DataLayout instead of string in (set_)data_layout (#228445)
Now that DataLayout is part of llvm rather than llvm_target, make use of
it in the `data_layout` and `set_data_layout` APIs.
[Mips] Restrict microMIPS GP load folding to word loads (#228881)
Only LW has a GP-relative variant, so restrict the folding to LW_MM and
LW_MMR6.
Assisted-by: OpenAI Codex # Testcases
[lld][Mips] Remove signed range check for R_MICROMIPS_26_S1 (#228880)
R_MICROMIPS_26_S1 encodes a shifted index within the current PC region,
not a signed absolute address. Applying the PC-relative relocation's
signed 27-bit check rejects valid destinations in higher address
regions.
Write the jump index without that check, as for R_MIPS_26, while
retaining
the signed range check for R_MICROMIPS_PC26_S1. Cover both jumps and
calls
at a high address in little- and big-endian objects.
Assisted-by: OpenAI Codex # Testcases
[Mips] Emit 16-bit microMIPS NOPs for alignment (#228879)
Zero-filled halfword padding is not a complete microMIPS NOP. Emit
move $zero, $zero when alignment requires a two-byte instruction, then
use zero-encoded 32-bit NOPs for the remaining padding.
Handle an odd leading byte as data padding and preserve zero filling in
standard MIPS mode. Extend the unaligned-NOP test to cover both byte
orders, odd-sized data, and mode changes.
Assisted-by: OpenAI Codex # Testcases
[lldb] Replay the cached error from GetLoadImageUtilityFunction (#226986)
`Process::LoadImage` needs a live process to load a library: on POSIX it
makes the C library's `dlopen` run inside the debuggee, and on Windows
it makes `LoadLibraryExW` run there, both through a JIT-compiled
function call that needs a process to run the call in. On a core file
(or, on Windows, a saved minidump) there is no such process, so the load
fails every time. POSIX and Windows each lose the resulting error
message, but for two different reasons.
On POSIX, only the first failure is reported correctly. Load a corefile,
and load two libraries that do not exist shows this:
```
(lldb) process load /tmp/does-not-exist-1.dylib
error: failed to load '/tmp/does-not-exist-1.dylib': dlopen error: could not make function caller: Can't make a function caller without a process.
(lldb) process load /tmp/does-not-exist-2.dylib
error: failed to load '/tmp/does-not-exist-2.dylib': (null)
```
[57 lines not shown]
[flang][OpenMP] Remove unnecessary check for extension clauses (#228878)
Remove the check for extension clauses. These now do have descriptors,
so the check guarding the descriptor retrieval is no longer necessary.
[clang][www] Fix typos in OpenProjects page (#228822)
Fix various typos in Clang's [OpenProjects
page](https://clang.llvm.org/OpenProjects.html).
Assisted-by: Claude Sonnet 5.5
[Clang][OpenMP][NFC] Add test for `class-type` data members as loop counters (#228831)
Fixes #140243
An OpenMP loop whose init assigns to a data member of class type, like
`for (a = x; ...)` with `I<int> a`, crashed in
`OpenMPIterationSpaceChecker::checkAndSetInit`. Such an assignment is an
`operator=` call, so it goes through the `CXXOperatorCallExpr` branch,
and two of the `setLCDeclAndLB` calls there took the bound from `BO`,
the `BinaryOperator` cast that had already failed and is null at that
point. The invalid code in the report is not needed: a valid loop over
an iterator-typed data member crashed the same way.
#203252 replaced those `BO->getRHS()` uses with `CE->getArg(1)` as part
of another fix, so the crash is gone on trunk and no source change is
needed. This PR only adds a regression test so the issue can be closed.
[SLP] Cancel the phantom load saving on widened loads
On a target whose consecutive scalar loads coalesce, cancel the load
saving of a bundle feeding a cast that widens its lanes, in store
chains, value lists and integer reductions. Without it a load of
<4 x i8> through sitofp into a store chain vectorizes into longer code.
A bundle keeps its saving when the loaded or the widened lanes pack two
to a register, since the scalar form unpacks every lane. Floating point
reductions keep their own rule.
[SLP] Cancel the phantom load saving on fadd reductions that lose an fma
Add TargetTransformInfo::consecutiveLoadsCoalesce, implemented by
AMDGPU, whose consecutive scalar loads already coalesce into one wide
access. On such a target cancel the load saving of a contract fadd
reduction over contract fmuls, which otherwise trades every scalar fma
for a saving that never materializes. A reassociable reduction keeps it,
as its vector fmuls fuse into the reduction.
[SLP] Drop the context of a vector fmul whose user is not vectorized
Pass the context of a vector fmul only when its user is vectorized or
the target emits it lane by lane. AMDGPU prices an fmul with a contract
fadd user as free, which does not hold for the extracted lanes of a
root. Other operations keep their context for the fast math flags and
the function attributes.
[clang][X86] bring `regparm` in line with GCC (#227130)
Removes some divergences between GCC and Clang using `regparm`. In
particular
- All floats count as floats (previously `f16`, `f16b`, `long double`
and `f128` did not)
- Complex numbers are always passed via the stack
- Unions are always passed like integers
- Some changes to how wrapping structs (with a single non-ZST field) are
handled
I've tested this empirically (with abi-cafe)
---------
Co-authored-by: Reid Kleckner <rkleckner at nvidia.com>