[Clang] Fix pack vs non-pack tie breaker. (#229024)
The tentative resolution for CWG1432 was not consistent for the
resolution of CWG1395.
This fixes the crash reported in #228870.
Additionally, this update the status of related core issues and papers
touching the same wording. Clang was never affected because we never
fully implemented CWG1395.
Fixes #228870
Fixes #27357
Assisted-by: Opus 5.5
[AArch64][CostModel] Consider some nxv1 casts as illegal (#230153)
Codegen does not support all kinds of casts for these types.
This patch adds missing costmodel and codedgen tests.
[flang-rt][libclc] Install multilib variants in their own directory (#230251)
A multilib runtimes build sets <PROJECT>_LIBDIR_SUBDIR. flang-rt
libraries and libclc bitcode were installed with the same name as the
base library. This doesn't really work for `libclc` because the clang
driver doesn't look through multilibs currently, but it should still go
somewhere else so it doesn't clobber.
[offload][nfc] Pull OpenMP's InteropTbl out of PluginManager
PluginManager is shared with OpenACC, so the OpenMP interop table moves
to OmpPluginManager in libomptarget. The OpenMP PM is defined there, and
interop cleanup is registered from initRuntime.
[offload][nfc] Use the PluginManager instance instead of the global PM
PluginManager methods referred to the global PM pointer instead of the
object they were called on. DeviceTy now carries the PluginManager that
created it. Move the PM definition and the RTL atomics to OffloadRTL.
[CIR][Matrix] Implement vector splat/add for CIR lowering (#230185)
This patch extends cir.add/cir.fadd to work on matrixes of int/float,
and lower correctly to add/fadd in LLVMIR.
It also extends cir.vec.splat to work with a matrix as well, so we could
get mixed-matrix-int/float operations to work. I considered making this
its own operation, however it is so nearly identical to cir.vec.splat
(and will become more so as we extend matrix) that it didn't seem
valuable to consider it separately, particularly as they lower to the
same things in LLVM.
One limitation: Splat is sometimes constant-folded during simplify.
However, we don't yet have a constant attribute type for a matrix, so
this is left for future work.
[RISCV] Custom lower fixed length mask_beforefirst
Rather than add a new _VL node, just put it in a scalable container with VLMAX so it gets the scalable vmsbf.m pattern. RISCVVLOptimizer can optimize the VL of it based on its users.
[X86] Declare memory effects for the AMX intrinsics (#229025)
The AMX intrinsics that name tile registers (`llvm.x86.tileloadd64`,
`llvm.x86.tdpbssd` and the rest of the non-`_internal` forms) have no
memory attributes, so to the optimizer every call may read and write any
memory. A loop-invariant value that lives in memory is reloaded after
every tile instruction, a store isn't forwarded past a tile load, and
two reads of the same tile row aren't merged.
This PR models the tile registers and the tile configuration as
`target_mem0`, the way AArch64 models SME's ZT0 and ZA
(`target_mem0`/`target_mem1`; `SME_Load_Intrinsic` is
`[IntrRead<[ArgMem, ZA]>, IntrWrite<[ZA]>]`), and declares what each
intrinsic actually does:
| Intrinsics | Reads | Writes |
|---|---|---|
| `tileloadd64`, `tileloaddt164`, `tileloaddrs64`, `tileloaddrst164` |
pointer argument, tile state | tile state |
[63 lines not shown]
CodeGen: Only visit live-out vregs when splitting a critical edge with LIS (#230441)
To update LiveIntervals, SplitCriticalEdge checked every virtual
register in the function for liveness at the end of the split block.
That made PHIElimination's edge splitting O(splits * vregs).
PHIElimination now computes the set of virtual registers live out of
each block before the first split and passes it through
SplitCriticalEdgeAnalyses. Only those registers are visited, and the set
for the new block is added once the intervals are updated.
This essentially resurrects the per-block sparse register sets used to
update LiveVariables, which were removed in #228618 and #230147, applied
to the LiveIntervals update instead.
Instructions retired in phi-node-elimination, x86_64 -O3, on a generated
chain of N compare blocks branching to a shared PHI block:
N before after after/before
[11 lines not shown]
[AArch64][llvm] Fix incorrect diagnostic (index/immediate in simm9)
Improve the error message in the simm9 diagnostics and use the word
'immediate' instead of 'index'. This was noticed when creating the
CFLT instructions for Armv9.8-A, but would have caused a lot of
unrelated churn if merged as part of that change.
Co-authored-by: Martin Wehking <martin.wehking at arm.com>
[AArch64][llvm] Fix constant expressions in CMPBR immediate aliases
Use the standard immediate parser for CMPBR aliases so that parenthesized
expressions and explicit unary plus are accepted. This also simplifies
the amount of code required, since we can remove `tryParseAdjImm0_63`,
and reuse code in `AdjImmAsmOperand` which does the same job.
Check immediates in their original ranges and apply the adjustment
when rendering MC operands. Add coverage for constant expressions in
`cbge`, `cbhs`, `cble` and `cbls`.
[AArch64][llvm] Improve diagnostics for invalid operands
Return `DiagnosticPredicate` from `isSImm()`, `isImmInRange()`, and
`isUImm12Offset()` so invalid operand types do not produce misleading
immediate range errors, while out-of-range constants retain their
range diagnostics.
Keep the load/store fallback predicate boolean and remove the
CFLT-specific diagnostic workaround. Update affected diagnostic
tests and add prefetch and CFLT regression testcases.
[libc] Implement tcsetwinsize in termios (#228498)
Implement the standard POSIX.1-2024 function `tcsetwinsize` in
`<termios.h>`.
Fixes #228380
Part of #228378
Implementation was assisted by Antigravity by analysing other functions
in header and reviewed by Aman Maurya.
[MLIR][Remark] Report why the LLVM remark streamer could not be created
LLVMRemarkStreamer::createToFile dropped the reason it failed: a file that
could not be opened and a serializer that could not be created both turned
into a bare failure, and enableOptimizationRemarksWithLLVMStreamer failed
without a diagnostic.
- createToFile returns llvm::Expected. It opens the file with
mlir::openOutputFile and returns its message; a serializer error is
returned as a FileError naming the file.
- enableOptimizationRemarksWithLLVMStreamer emits the error on the context.
Assisted-by: Claude Code (Claude Opus 5.5)
[IR] Name unnamed values after each pass (#221519)
Unnamed IR values are hard to follow in pass-by-pass output. Their
printed
numbers can shift when a pass adds or removes values.
Add the hidden `-instnamer-after-each-pass` option to name unnamed
arguments,
basic blocks, and non-void instructions before the first snapshot and
after
each new pass manager pass. It reuses InstructionNamer and keeps a
counter
across the pipeline, so a generated name is not reused after its value
is
removed. Existing names are preserved.
This makes text snapshots easier to compare, but names do not identify
instruction objects: a pass can transfer or reuse an existing name. The
ordinary `instnamer` pass and `-print-changed=diff` output are
unchanged.
[ARM][Disassembler] Add missing predicates when decoding Thumb barriers (#223540)
`DecodeThumb2BCCInstruction` omits predicate operands for Thumb barriers
when the data-barrier feature is disabled, causing llvm-objdump
--arch-name=thumb to abort.
Add the missing operands using the current IT state, with disassembler
and llvm-objdump regression tests.
Fixes #193877
[RISCV] Generalize combineBinOpOfZExt to allow sext
The combine currently only handles cases where both operands are zext, but we can also allow sext.
If any operand is sexted then we need to sext the result as the sign bit may not be zero.
We can't handle udiv + urem if either of the operands are sext, since sign extending a narrower udiv/urem gives incorrect results. E.g. `udiv (sext (i8 -1) to i32), (zext (i8 2) to i32)`:
- At i32: `udiv 0xffffffff, 2 = 0x7fffffff`
- Narrowed to i16: `sext (udiv 0xffff, 2) to i32 = sext 0x7fff to i32 = 0x00007ffff`
[CIR] Find conditional cleanups in implicit code (#229829)
`ConditionalEvaluationFinder` in `CIRGenCleanup` skipped implicit code
because that's the default for `RecursiveASTVisitor`. This lead to
default arguments and default member initializers being skipped.
This patch enables the traversal of implicit code with the exception of
the implicit call to `await_resume()` in `co_await` and `co_yield`
expressions. This requires cleanup scopes for await full-expressions
which don't exist yet.
---------
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[LLVM][CodeGen][SVE] Add lowering for bfloat strict-fp cast operators. (#223709)
The majority of the changes are just a case of ensuring the chain is
routed correctly and the matching STRICT passthrough node is used.
NOTE: At present full strict-fp support has a minimum requirement of
+sve2+bf16, otherwise we lack the necessary cast instructions. Of these
+bf16 is fundamental whereas +sve2 is only required for double->bfloat.
[SLP]Fix poison shift after MinBW demotion
Lanes converted from mul-by-power-of-2 are emitted as shl by the
exponent; check the node shift amounts, not the scalar operands.
Fixes #230392
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/230464