[lldb][test][NFC] Set multithreaded_queue pop timeout in one place (#226988)
All callers of pop were passing 5 seconds. I don't forsee us needing to
change it per caller either, so just hardcode it in pop().
AArch64: Mark the LR def of GlobalISel calls dead (#226256)
In SelectionDAG InstrEmitter would have marked this dead.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[SystemZ][z/OS] Allocate a stack frame when callee-saved registers are saved (#226681)
On z/OS (XPLINK64), a function without calls and without stack objects
that uses callee-saved registers got a stack size of 0. No stack frame
was allocated, so `STMG` stored the registers relative to the unchanged
stack pointer, i.e. into the register save area of the caller's DSA,
overwriting the caller's saved return address. Only skip the frame when
there is nothing to save.
Two existing tests (`zos-ret-addr.ll`, `zos-frameaddr.ll`) encoded this
store into the caller's frame and are updated; the values returned by
`llvm.returnaddress`/`llvm.frameaddress` are unchanged (the old stack
pointer is still read from offset 2048 of the current frame). New test
`zos-frame-callee-saved.ll`.
Found with GMP 6.3.0 on z/OS 3.1 (`mpn_mul_1`): code after the call ran
twice, then S0C4/S0C6. With this change GMP works.
Tests: `llvm-lit test/CodeGen/SystemZ test/MC/SystemZ test/MC/GOFF`
[6 lines not shown]
[libc] Add a few missing fcntl.h macro. (#226810)
* Add a few POSIX-specified symbolic constants to `<fcntl.h>`:
`O_RSYNC`,
`AT_SYMLINK_FOLLOW`, and `F_DUPFD_CLOEXEC`.
* Don't add `_CLOFORK` variants, since those are not supported by Linux.
* List the constants in the `fcntl.yaml` file for posterity.
[InstCombine] Fix miscompilation in sinkNotIntoOtherHandOfLogicalOp (#226783)
Consider the following: `or X, (not (not X))`
sinkNotIntoOtherHandOfLogicalOp calls `freelyInvert(X)` but as X is used
in the not this call result in the new instruction is `and (not X), (not
(not X)) -> and (not X), X` instead of the expected `and (not X), (not
X)`
fixes https://github.com/llvm/llvm-project/issues/226504
[lldb][test] Disable TestEvents.py (#226486)
This test is flaky on Linux, macOS and Windows. The usual failure on
macOS is as follows:
```
FAIL: test_shadow_listener (TestEvents.EventAPITestCase)
----------------------------------------------------------------------
Traceback (most recent call last):
llvm-project/lldb/packages/Python/lldbsuite/test/decorators.py", line 801, in wrapper
func(*args, **kwargs)
llvm-project/lldb/test/API/python_api/event/TestEvents.py", line 485, in test_shadow_listener
state, restarted = self.wait_for_next_event(None, False)
llvm-project/lldb/test/API/python_api/event/TestEvents.py", line 346, in wait_for_next_event
self.assertEqual(
>>> AssertionError: 2 != 3 : matching stop hook
Config=arm64-/Users/ec2-user/jenkins/workspace/llvm.org/lldb-cmake-matrix/lldb-build/bin/clang
----------------------------------------------------------------------
```
[2 lines not shown]
[AMDGPU] Fold 24 bit multiply with zero low bits
Fold `MUL_I24` and `MUL_U24` to zero when either operand has known zero
low 24 bits.
For example:
```
llvm.amdgcn.mul.i24(x, 0x01000000) -> 0
```
[SelectionDAG][AMDGPU] Fold mul24 with an AND operand whose low bits are zero
Use SimplifyMultipleUseDemandedBits to simplify AND operands based on the
low 24 bits consumed by mul24.
Fold the multiply to zero when the simplified operand is zero.
This folds cases such as:
mul24(x & 0xff000000, y) -> 0
[SelectionDAG] Handle constants in SimplifyMultipleUseDemandedBits
Replace a non-zero constant with zero when none of its set bits are
demanded.
This allows users of `SimplifyMultipleUseDemandedBits` to eliminate
irrelevant constant bits while preserving the convention that a null
SDValue indicates no simplification.
[AMDGPU] Fold mul24 with an operand whose low 24 bits are zero
mul24 only reads the low 24 bits of each operand. Try SimplifyDemandedBits
before SimplifyMultipleUseDemandedBits so that a single use operand such as
`(x & 0xff000000)` is simplified to zero, then fold the multiply to zero when
either operand is a zero constant.
This folds cases such as:
```
mul24(x & 0xff000000, y) -> 0
mul24(x, 0) -> 0
```
[lldb][test] Simplify multithreaded_queue (#226979)
This makes it easier to see how the timeout is applied in pop.
push now releases the mutex before notifying, where before the watchers
would have to wait for push to return. Likely not going to make anything
much faster, but it makes a bit more sense. Finish adding the new item,
then notify.
[offload][l0] Execute dynamic linking for all formats (#226980)
Also do dynamic linking when necessary for Offload Binaries and SPIR-V
images and not just for the old ELF format.
NVPTX: Drop LiveVariables from the register allocation pipeline
The optimized RegAlloc pipeline ran LiveVariables only to satisfy PHIElimination
and TwoAddressInstruction, both of which no longer need it. Remove the
LiveVariables run (and, in the new pass manager, the UnreachableMachineBlockElim
that was there only as a LiveVariables prerequisite).
Co-authored-by: Claude (Claude-Opus-4.8)
CodeGen: Drop the LiveVariables parameter from convertToThreeAddress
This was used for analysis updates, but now the analysis is being
removed.
Co-authored-by: Claude (Claude-Opus-4.8)
AMDGPU: Remove update-only LiveVariables maintenance from SILowerControlFlow
This was only maintained, never relied on. Part of staged LiveVariables
removal.
Co-authored-by: Claude (Claude-Opus-4.8)
CodeGen: Remove LiveVariables use from TwoAddressInstructionPass
Now that LiveIntervals is computed unconditionally before TwoAddressInstructions
in the pipeline, the pass no longer needs LiveVariables.
Co-authored-by: Claude (Claude-Opus-4.8) <noreply at anthropic.com>
CodeGen: Compute LiveIntervals before TwoAddressInstructions
TwoAddressInstructions is the traditional primary use of LiveVariables,
but it has gained a LiveIntervals path. By moving LiveIntervals earlier,
the default flips to rely on it instead of LiveVariables. The overall
test churn is mostly neutral, with more net wins than losses.
This should move before phi elimination. This is a staging move to
incrementally remove the LiveVariables support from TwoAddressInstructions,
and because the move to running LiveIntervals on SSA is a bigger leap.
Co-authored-by: Claude (Claude-Opus-4.8)
[AArch64] Add custom v2i32->v2i16 store lowering. (#223659)
This extends the existing v4i16->v4i8 lowering to handle v2i32->v2i16
stores too, which in general helps keep the number of instructions down.
[flang] Lower loops whose branching is confined to their body structurally
Such a loop was classified separately by a previous change but still
lowered as unstructured, so its structured form was lost.
Lower it structurally instead, with its body folded into a region that can
hold the branching. The loop keeps its bounds on the op, so it remains
available to whatever transforms or parallelizes it. Only the body is
folded: the loop control statements are emitted as they are for any
structured loop, since a branch from outside may target either of them.
Loops an OpenACC or OpenMP directive owns are lowered the same way, so
they keep their form too.
[flang][Test] Cover the lowering of loops with a non-terminating body
A previous change leaves such a loop unstructured. Check what that produces:
the cycle survives as a block branching to itself, no fir.do_loop is emitted
for the loop control, and a loop that only needs a block of its own still
gets the structured form with its body in a region.
[flang] Keep loops with a non-terminating body unstructured
A loop body that cannot run to completion must not be folded into an
scf.execute_region: the region carries no memory effects, so DCE deletes
it outright, dropping the non-termination and letting execution fall past
the loop. Branches survive that, being terminators.
An infinite DO was already rejected. Follow chains of unconditional GO TOs
as well and reject a body whose chain closes on itself, which is the same
bound the cf.br canonicalization applies to cyclic branches.
[flang] Honor the execute-region wrap flag when detecting loop internals
A loop whose branching is confined to its body is lowered with that body
in an scf.execute_region, since the CFG needs more than the single block
fir.do_loop's region admits. With the wrap disabled there is nowhere to
put those blocks, so skip the reclassification and leave the loop
unstructured.
[flang] Detect loops whose branching is confined to their body
A DO loop is classified as either structured or unstructured, and a single
raw branch anywhere in its body forces the loop -- and every construct
enclosing it -- onto the unstructured path.
That is stronger than necessary. A loop keeps its structured control flow
as long as its branching neither leaves its body nor enters it from
outside. Classify such a loop separately from a fully unstructured one.
This only classifies: lowering is unchanged. PFT dumps mark the new
classification with '~', which is what the tests key on.
[flang] Resolve an assigned GO TO's targets from the completed assign map
An assigned GO TO reaches any label ASSIGNed to its variable, and a label
list does not bound that: lowering allows a branch to any ASSIGNed label
whether or not the list names it. Branch analysis only sees the ASSIGNs
preceding the GO TO in program order, so the successors it records, and the
incoming branches derived from them, can be incomplete.
The symbol-to-labels map is complete once branch analysis has finished,
which is when the classification runs. Ask it for the full target set
instead of trusting the recorded successors, so a loop whose assigned GO TO
stays within its body is still recognised.