[BOLT,test] Add missing pipe in dwarf5-debug-names-skip-forward-decl.s (#226338)
With cl::Grouping, --check-prefix=POSTCHECK parses as -c -h..., so
llvm-dwarfdump prints --help and exits 0, and the checks never run.
[libc++] Run the PR benchmarking tooling from main (#226278)
Previously, we'd use the workflow file and machines.json from the main
branch, but the rest of the tools (e.g. build-at-commit) would be taken
from the PR head. This patch switches to using the tools from main and
only using the PR head's content for the actual code and benchmarks.
This should make it easier to evolve the tools and the workflow files
without breaking PR benchmarking for people who submit PRs from slightly
outdated bases.
[RISCV] Support frame pointers in SiFive CLIC preemptible handlers (#221318)
This commit adds frame-pointer support for SiFive CLIC preemptible
handlers by using t0 to save mcause and mepc to dedicated stack slots in
both cases.
[CIR] Update the fastmath contract lowering test
cir-to-llvm now requires a module triple, and the LLVM dialect prints
fastmath flags as fastmath<contract> rather than an attribute dictionary.
Co-authored-by: Cursor <cursoragent at cursor.com>
[AMDGPU][InstCombine] Fold zero dot operands to accumulator
Fold AMDGPU dot intrinsics when either operand is zero.
`dot(a, 0) = 0` and `dot(0, b) = 0`, so replace the intrinsic with its accumulator.
This avoids unrelated clamp and add/sub reassociation cases.
[libc++] Take the benchmark testing configuration from the test suite (#226267)
Previously, we'd use the testing configuration (the .cfg.in file) from
the checked out test suite for the PR benchmarking, but we'd use the one
from `main` for the historical benchmarking. Since the config file is
logically part of the test suite (just like the rest of the Lit setup),
it makes more sense to take it from the checked-out test suite version.
[SelectionDAG][AMDGPU] Fold mul24 with an AND operand whose low bits are zero
Use SimplifyMultipleUseDemandedBits to simplify AND operands based on the
low 24 bits consumed by mul24.
Fold the multiply to zero when the simplified operand is zero.
This folds cases such as:
mul24(x & 0xff000000, y) -> 0
[AMDGPU] Fold 24 bit multiply with zero low bits
Fold `MUL_I24` and `MUL_U24` to zero when either operand has known zero
low 24 bits.
For example:
```
llvm.amdgcn.mul.i24(x, 0x01000000) -> 0
```
[SelectionDAG] Handle constants in SimplifyMultipleUseDemandedBits
Replace a non-zero constant with zero when none of its set bits are
demanded.
This allows users of `SimplifyMultipleUseDemandedBits` to eliminate
irrelevant constant bits while preserving the convention that a null
SDValue indicates no simplification.
[AMDGPU] Fold a constant add/sub into the sudot4/sudot8 accumulator (#225322)
Fold a constant add into the accumulator operand of sudot4 and sudot8
when
clamping is disabled:
```
sudot(a, b, C1, false) + C2 -> sudot(a, b, C1 + C2, false)
```
Subtraction by a constant is canonicalized to addition of its negation.
[HLSL] Add support for dynamic resources (#221103)
Adds support for dynamic resources, also known as bindless or directly
indexed resources. The design is described in
https://github.com/llvm/wg-hlsl/issues/204.
The change adds:
- The `ResourceDescriptorHeap` and `SamplerDescriptorHeap` global
symbols, which expose the Shader Model 6.6 descriptor heaps.
- Internal `hlsl::__detail::heap_resource_info` and
`hlsl::__detail::heap_sampler_info` types, which are produced by
indexing the corresponding descriptor heap. HLSL resource and sampler
types can be implicitly constructed from these values.
- Constructors that accept the corresponding internal heap info type for
all currently implemented resource classes.
- Clang builtins `__builtin_hlsl_resource_handlefromheap` and
`__builtin_hlsl_resource_counterhandlefromheap` for creating resource
and counter handles from descriptor heap indices.
- Sema validation and return-type handling for the new builtins.
[8 lines not shown]
[MIR][AMDGPU] Serialize register allocation anti-hints (#218073)
This PR adds MIR print and parse support for the register allocation
anti-hints through
a per vreg anti-hints: [ .. ] field in the registers: section. The
printer only emits
the anti-hint registers only when the it is non-empty to keep existing
tests unchanged.
## Stack
PR **2/4**. Depends on #218071. Next: Apply Occupancy-Aware allocation
anti-hints #218074.
[AMDGPU] Track allocator-reserved register pressure separately
GCNRegPressure counts a tuple at its live lane count, but the excess and
critical limits the scheduler compares against are expressed in
allocatable registers. A tuple is reserved at its full class width for
its whole live range, so lanes that are dead at a given point are still
unavailable to any other value, and the live lane count underestimates
what the allocator has to set aside.
Track a third set of per-kind values recording the reserved width, and
use it for the pressure the GCN trackers hand to the scheduling
heuristics.
Also add -amdgpu-trackers-compare-rp to print the GCN tracker and the
generic tracker pressure side by side for every candidate, which is what
the discrepancy was found with.
WIP
AI-Assisted
[WebKit Checkers] Add alpha.webkit.UnborrowedLambdaCapturesChecker (#226288)
Like alpha.webkit.UnborrowedLocalVarsChecker, but for lambda captures.
This checker is pretty restrictive because, unlike a refcounted object,
a lifetime-dependent pointer/reference/view cannot be captured in an
escaping closure at all. Still, we permit capturing in a NOESCAPE
closure.
[CIR] Record -ffp-contract=fast as a per-op contract flag
Classic CodeGen stamps contract on floating-point instructions so a later
Standard-fusion backend can still form an FMA. CIR only fused within a
statement via cir.fmuladd, which dropped FFMA on the CUDA device default.
Co-authored-by: Cursor <cursoragent at cursor.com>
[clang][DepScan] Disable relocation checks (#225563)
This internally broke an incremental build. Disable while it gets
investigated.
This is partial revert of cf8597bd3b87
resolves: rdar://188026923
[CIR] Record -ffp-contract=fast as a per-op contract flag
Classic CodeGen stamps contract on floating-point instructions so a later
Standard-fusion backend can still form an FMA. CIR only fused within a
statement via cir.fmuladd, which dropped FFMA on the CUDA device default.
Co-authored-by: Cursor <cursoragent at cursor.com>
[mlir][bufferization] Support matching permutation maps in Linalg analysis (#226079)
Generalize `LinalgOpInterface::bufferizesToElementwiseAccess` to
recognize participating tensor operands with identical permutation
indexing maps, instead of requiring identity maps.
This enables in-place bufferization after loop interchange. Projected
and mismatching maps, reduction iterators, and sparse operands remain
conservatively rejected.
The tests cover the analysis decision as well as end-to-end
bufferization without an allocation or copy.
[mlir][bufferization] Handle disjoint subsets in One-Shot Analysis (#224184)
Teach One-Shot Analysis to ignore RaW conflicts between provably
disjoint subset reads and writes. Trace aliasing reads back to subset
extraction ops to handle alias-only operations such as
tensor.extract_slice.
Add tensor and vector tests for disjoint, overlapping, and unknown
subsets.
[RISCV] Make the getJumpTableIndex search more structured. NFC (#226266)
Instead of recursively searching through instructions with certain
opcodes, look for the specific patterns we use for jump tables.
I've added a few TODOs for Zilx and SHXADD_UW.
[RISCV][llvm-readobj] Treat RISC-V float ABI e_flags field as an enum (#226313)
EF_RISCV_FLOAT_ABI is a 2-bit field, not independent flag bits, so
printFlags needs to be told about the mask. Previously a quad-float ABI
object printed "single-float ABI, double-float ABI, quad-float ABI"
instead of just "quad-float ABI".
Co-Authored-By: Claude Sonnet 5 <noreply at anthropic.com>
[mlir][tosa] Use optnone for compliance initialization under HWAsan (#226310)
Similar to #223586, compiling the file is very slow when we enable
HWAsan internally. Mark the constructor with optnone under
LLVM_HWADDRESS_SANITIZER_BUILD to skip expensive optimization and
register allocation passes, bringing compilation time from 10+m to ~30s.
[LiveDebugVariables] Repair stale SlotIndexes
The analysis keeps its indexes from before the first register allocator
until DBG_VALUEs are emitted, by which point passes in between have
erased some of the instructions they point at. Resolve them at the
start of each allocator run and before emitting.
SlotIndexes can then reclaim the entries of erased instructions without
sparing the ones held here, which would have made generated code depend
on -g. Emitted locations are unchanged, except that intervals resolving
to one position now emit a single DBG_VALUE rather than identical
consecutive ones.
[SlotIndexes] Add queries for stale indexes
An erased instruction leaves its index list entry in place, making the
index indistinguishable from a block boundary entry. Add
isBlockBoundaryIndex() and isStaleIndex() to tell the two apart, and
canonicalizeIndex() to resolve a stale index to the closest preceding
instruction's register slot, or the block start if none survives.
NFC. No caller yet. LiveDebugVariables is next.
Improve llvm-gsymutil to set the end address correctly. (#225593)
When converting DWARF in llvm-gsymutil, if the last function info had no
size, it would set it to the end of the valid text ranges. This meant a
symbol could get a larger than necessary size in mach-o files. This
patch will set the address to the minimum of the last TextRanges or the
end of the section that contains the symbol.